CAS9 PROTEINS WITH ENHANCED SPECIFICITY AND USES THEREOF
Patent Information
- Application Number
- JP2024542033
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-01-12
- Filing Date
- 2023-01-11
- Publication Date
- 2026-01-20
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This PCT application claims priority benefit of U.S. Provisional Patent Application No. 63 / 298,822, filed January 12, 2022, which is incorporated by reference in its entirety.
[0002] Reference to Electronically Submitted Sequence Listing The contents of the electronically submitted sequence listing submitted with this application (Name: 4603_001PC02_Seqlisting_ST26.xml, Size: 339,518 bytes, Creation Date: January 10, 2023) are incorporated herein by reference in their entirety.
[0003] The present disclosure provides Cas9 proteins that have been modified to exhibit enhanced specificity (or fidelity), as well as compositions, polynucleotides, vectors, cells, and kits related to such Cas9 proteins. The present disclosure also provides methods for making and using the modified Cas9 proteins in a wide range of clinical settings (e.g., both therapeutic and diagnostic). [Background technology]
[0004] Gene editing technology has emerged as a powerful and versatile technique with the potential for a wide range of clinical applications. In particular, the CRISPR-Cas system has been used by researchers to successfully disable or repair genes in various cell types and species. Despite such advances, off-target events remain a major problem, preventing the broad application of the CRISPR-Cas system to various targets and disease conditions. See, for example, Jinek et al., Science 337:816-821 (2012); and Fu et al., Nature Biotechnology 31:822-826 (2013).
[0005] Protein engineering has been used to improve the specificity of conventional Cas proteins (e.g., Streptococcus pyogenes Cas9 (SpCas9)). For example, compared with wild-type SpCas9 proteins, eSpCas and HF-SpCas9 proteins have much higher specificity. See, for example, Slaymaker et al., Science 351:84-88 (2016); and Kleinstiver et al., Nature 529:490-495 (2016). However, off-target effects are still observed quite frequently, especially when a single base mismatch is present during the gene editing process. In particular, when CRISPR is used for in vitro cleavage, once damaged DNA is not repaired because there is no repair mechanism like in cells, and this is no exception even in mismatch cleavage. Therefore, under in vitro cleavage conditions, mismatch cleavage becomes more prominent, making the limitations of the current CRISPR-Cas system more apparent. Therefore, there remains a need for improved Cas9 proteins that can increase specificity and reduce off-target effects during gene editing. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] Jinek et al.,Science 337:816-821(2012) [Non-Patent Document 2] Fu et al.,Nature Biotechnology 31:822-826(2013) [Non-Patent Document 3] Slaymaker et al.,Science 351:84-88(2016) [Non-Patent Document 4] Kleinstiver et al.,Nature 529:490-495(2016) Summary of the Invention
[0007] Provided herein is a Cas9 protein that includes a cavity domain that includes a plurality of positively charged amino acids, at least one of which is modified compared to a corresponding wild-type Cas9 protein (an "amino acid modification"), which can increase the specificity of the Cas9 protein.
[0008] In some embodiments, the plurality of positively charged amino acids of the cavity domain comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or more amino acid modifications. In some embodiments, the Cas9 protein comprises the amino acid sequence of SEQ ID NO: 1, and the amino acid modifications are made at one or more of the following residues of SEQ ID NO: 1: R785, K789, R455, R721, R919, R1241, R939, K1189, K941, R1226, K1228, or a combination thereof. In some embodiments, the amino acid modifications are made at residues R785, K1189, R1241, or a combination thereof. In some embodiments, the amino acid modifications are made at residues K1189 and R1241. In some embodiments, the amino acid modifications are made at residues R785, K1189, and R1241.
[0009] Provided herein is a Cas9 protein comprising an amino acid sequence as set forth in SEQ ID NO:1 with at least one amino acid modification, the at least one amino acid modification being at residues K405, R455, K546, K561, K562, K564, K566, K578, K579, R618, R622, K664, R721, R785, K786, K788, K789, R807, K808, R809, R810, R811, R812, R813, R814, R815, R816, R817, R818, R819, R820, R821, R822, R823, R824, R825, R826, R827, R828, R829, R830, R831, R832, R833, R834, R835, R836, R837, R838, R839, R840, R841, R842, R843, R844, R845, R846, R850, R851, R852, R853, R854, R855, R856, R857, R858, R859, R860, R861, R862, R863, R864, R865, R866, R870, R871, R872, R873, R874, R875, R876, R877, R878, R879, R880, R881, R882, R883, R884, R885, R886, R887, R888, R889 49, R856, K914, K917, R919, R920, K921, K922, R926, K934, K936, R939, K941, K945, R1047, R1131, R1137, K1142, K1152, K1155, R1178, K1189, K1198, K1206, K1213, K1223, R1226, K1227, K1228, R1241, or combinations thereof.
[0010] In some embodiments, at least one amino acid modification is made at residue K405, R455, K566, K578, K664, R721, R785, K786, K789, K914, K917, R919, K921, K922, R926, K934, K936, R939, K941, K945, R1137, K1142, K1152, K1189, K1198, K1206, K1223, R1226, K1227, K1228, R1241, or a combination thereof of SEQ ID NO:1. In some embodiments, at least one amino acid modification is made at residues R455, R785, R721, K789, R919, R1241, R939, K941, K1189, R1226, K1228, or a combination thereof, of SEQ ID NO: 1. In some embodiments, the amino acid modification is made at residues K1189 or R1241 of SEQ ID NO: 1. In some embodiments, the amino acid modification is made at residues (i) K1189 and R1241 of SEQ ID NO: 1, (ii) R721 and R1241 of SEQ ID NO: 1, or (iii) R785 and R1241 of SEQ ID NO: 1. In some aspects, the amino acid modifications are made at residues (i) R785, K1189, and R1241 of SEQ ID NO:1, (ii) R721, K1189, and R1241 of SEQ ID NO:1, or (iii) K1189, K1228, and R1241 of SEQ ID NO:1.
[0011] In any of the Cas9 proteins described herein that comprise at least one amino acid modification, in some embodiments, the amino acid modification comprises an alanine substitution.
[0012] The present disclosure provides a Cas9 protein comprising, consisting of, or essentially consisting of the amino acid sequence set forth in SEQ ID NO:2. Also provided herein is a Cas9 protein comprising, consisting of, or essentially consisting of the amino acid sequence set forth in SEQ ID NO:3. Also provided herein is a Cas9 protein comprising, consisting of, or essentially consisting of the amino acid sequence set forth in SEQ ID NO:4. Also provided herein is a Cas9 protein comprising, consisting of, or essentially consisting of the amino acid sequence set forth in SEQ ID NO:5. Provided herein is a Cas9 protein comprising, consisting of, or essentially consisting of the amino acid sequence set forth in SEQ ID NO:6. Provided herein is a Cas9 protein comprising, consisting of, or essentially consisting of the amino acid sequence set forth in SEQ ID NO:7. Provided herein is a Cas9 protein comprising, consisting of, or essentially consisting of the amino acid sequence set forth in SEQ ID NO:8. Provided herein is a Cas9 protein comprising, consisting of, or essentially consisting of the amino acid sequence set forth in SEQ ID NO:9.
[0013] The present disclosure further provides a composition comprising any of the Cas9 proteins of the present disclosure. In some embodiments, the composition further comprises a guide polynucleotide. In some embodiments, the guide polynucleotide comprises a single guide RNA (sgRNA).
[0014] Provided herein is an isolated polynucleotide encoding any of the Cas9 proteins of the present disclosure.Also provided herein is a vector comprising the isolated polynucleotide.Also provided herein is a cell comprising the vector.
[0015] Also disclosed herein are kits comprising any of the Cas9 proteins of the present disclosure and instructions for use. In some embodiments, the kits further comprise a guide polynucleotide. In some embodiments, the guide polynucleotide comprises a single guide RNA (sgRNA).
[0016] The present disclosure provides a method of enriching a first nucleotide sequence in a biological sample comprising a first nucleotide sequence and a second nucleotide sequence, the method comprising contacting the biological sample with any of the Cas9 proteins of the present disclosure, wherein the first nucleotide sequence comprises a mutation and the second nucleotide sequence does not comprise the mutation, and wherein the Cas9 protein recognizes the mutation and is thereby capable of cleaving the second nucleotide sequence but not the first nucleotide sequence.
[0017] In some embodiments, after contacting, the percentage of the first nucleotide sequence present in the biological sample is increased by at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 40-fold, or at least about 50-fold compared to the percentage of the first nucleotide sequence present in the reference sample (e.g., the biological sample before contacting). In some embodiments, after contacting, the amount of the second nucleotide sequence present in the biological sample is decreased by at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 100% compared to the amount of the second nucleotide sequence present in the reference sample (e.g., the biological sample before contacting). In some embodiments, after contacting, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or about 100% of the nucleotide molecules present in the biological sample comprise the first nucleotide sequence.
[0018] Also provided herein is a method of measuring the amount of a first nucleotide sequence that comprises a mutation in a biological sample, the method comprising contacting the biological sample with any of the Cas9 proteins of the present disclosure, whereby the amount of a second nucleotide sequence present in the biological sample is reduced, and the second nucleotide sequence does not comprise the mutation.
[0019] In some embodiments, the first and second nucleotide sequences are identical except for mutations, hi some embodiments, the mutations include substitutions, insertions, deletions, indels, duplications, inversions, large genomic rearrangements, or combinations thereof.
[0020] In some embodiments, the first nucleotide sequence comprises a single mutation. In some embodiments, the first nucleotide sequence comprises multiple mutations. When multiple mutations are present, in some embodiments, each of the multiple mutations is the same. In some embodiments, two or more of the multiple mutations are different.
[0021] In some embodiments, the mutation is within (i) the target site to which the guide polynucleotide binds, (ii) the protospacer adjacent motif (PAM), or (iii) both (i) and (ii).
[0022] In some embodiments, the biological sample is obtained from a subject suffering from or at high risk of developing a disease. In some embodiments, the mutation is associated with a disease.
[0023] In some embodiments, the first nucleotide sequence comprises circulating tumor DNA (ctDNA) and the second nucleotide sequence comprises non-ctDNA.
[0024] The disclosure further provides a method of diagnosing a disease in a subject in need thereof, the method comprising detecting whether the amount of a nucleotide sequence comprising a mutation associated with the disease is increased in a biological sample obtained from the subject compared to the corresponding amount present in a reference sample (e.g., a biological sample taken from a subject not suffering from the disease), where prior to detection, the biological sample has been contacted with any of the Cas9 proteins described herein.
[0025] In some embodiments, if the amount of nucleotide sequence containing mutation is increased in biological sample compared to the corresponding amount present in reference sample, the subject develops or is at risk of developing disease.In some embodiments, the amount of nucleotide sequence containing mutation is increased at least about 1-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 40-fold, or at least about 50-fold in biological sample compared to the corresponding amount present in reference sample.
[0026] In some embodiments, diagnosis is performed ex vivo.
[0027] In some embodiments, the disease comprises cancer, hematological disease, neurodegenerative / neurological disease, infectious disease, rheumatic disease, allergic disease, psychiatric disease, visual disease, endocrine disease, congenital disease, cardiovascular disease, pulmonary disease, renal disease, gastrointestinal disease, liver disease, or a combination thereof. In some embodiments, the cancer comprises lung cancer (e.g., non-small cell lung cancer), breast cancer, pancreatic cancer, bile duct cancer, gallbladder cancer, liver cancer, colorectal cancer, renal cancer, prostate cancer, gastric cancer, ovarian cancer, uterine cancer, cervical cancer, musculoskeletal cancer, or a combination thereof. In some embodiments, the neurodegenerative / neurological disease comprises Alzheimer's disease, Parkinson's disease, amyotrophic lateral sclerosis, Friedreich's ataxia, Huntington's disease, Lew body disease, spinal muscular atrophy, stroke, or a combination thereof.
[0028] Provided herein are methods of reducing the occurrence of off-target cleavage of a nucleic acid sequence during CRISPR-based gene editing, the methods comprising contacting the nucleic acid sequence with a complex comprising a Cas9 protein and a guide polynucleotide, where the Cas9 protein comprises an amino acid modification that can increase the specificity of the Cas9 protein and thereby reduce the occurrence of off-target cleavage. In some aspects, the Cas9 protein comprises any of the Cas9 proteins disclosed herein.
[0029] In some embodiments, the occurrence of off-target cleavage is at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 100% lower than the occurrence of off-target cleavage by a reference Cas9 protein. In some embodiments, the reference Cas9 protein comprises a corresponding Cas9 protein that does not contain an amino acid modification. In some embodiments, the reference Cas9 protein comprises an amino acid sequence set forth in any one of SEQ ID NO:244, SEQ ID NO:1, SEQ ID NO:245, SEQ ID NO:246, SEQ ID NO:247, or SEQ ID NO:248.
[0030] Disclosed herein are methods of increasing the specificity of a Cas9 protein, comprising modifying at least one amino acid residue of the Cas9 protein, wherein the at least one amino acid residue is capable of interacting with a backbone phosphate of a DNA sequence.
[0031] In some embodiments, the at least one amino acid residue to be modified is selected from the group consisting of residues K405, R455, K546, K561, K562, K564, K566, K578, K579, R618, R622, K664, R721, R785, K786, K788, K789, R807, K808, R849, R856, K914, K9 17, R919, R920, K921, K922, R926, K934, K936, R939, K941, K945, R1047, R1131, R1137, K1142, K1152, K1155, R1178, K1189, K1198, K1206, K1213, K1223, R1226, K1227, K1228, R1241, or combinations thereof. In some embodiments, the at least one amino acid residue to be modified comprises residue K405, R455, K566, K578, K664, R721, R785, K786, K789, K914, K917, R919, K921, K922, R926, K934, K936, R939, K941, K945, R1137, K1142, K1152, K1189, K1198, K1206, K1223, R1226, K1227, K1228, R1241, or a combination thereof, corresponding to the amino acid sequence set forth in SEQ ID NO:1. In some embodiments, the at least one amino acid residue to be modified comprises residue R455, R785, R721, K789, R919, R1241, R939, K941, K1189, R1226, K1228, or a combination thereof, corresponding to the amino acid sequence set forth in SEQ ID NO: 1. In some embodiments, the at least one amino acid residue to be modified comprises K1189 or R1241, corresponding to SEQ ID NO: 1. In some embodiments, the at least one amino acid residue to be modified comprises (i) K1189 and R1241, (ii) R721 and R1241, (iii) R785 and R1241, or (iv) K1189 and R1241, corresponding to SEQ ID NO: 1. In some embodiments, the at least one amino acid residue to be modified is (i) R785, K1189, and R1241, (ii) R721, K1189, and R1241, or (iii) K1189, K1228, and R1241, corresponding to SEQ ID NO:1. Includes.
[0032] In some aspects, the amino acid modification comprises an alanine substitution.
[0033] In some embodiments, after modification, the selectivity of the Cas9 protein is increased by at least about 1-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 40-fold, or at least about 50-fold compared to the specificity of the reference Cas9 protein. In some embodiments, the reference Cas9 protein comprises a corresponding Cas9 protein that does not contain an amino acid modification. In some embodiments, the reference Cas9 protein comprises an amino acid sequence set forth in any one of SEQ ID NO:244, SEQ ID NO:1, SEQ ID NO:245, SEQ ID NO:246, SEQ ID NO:247, or SEQ ID NO:248.
[0034] In some embodiments, after modification, the Cas9 protein is able to distinguish between a first nucleotide sequence that contains the mutation and a second nucleotide sequence that does not contain the mutation, such that the Cas9 protein cleaves the second nucleotide sequence but not the first nucleotide sequence. In some embodiments, the mutation is within (i) the target site to which the guide polynucleotide binds, (ii) the protospacer adjacent motif (PAM), or (iii) both.
[0035] Provided herein is a method for genetically modifying a cell, comprising contacting the cell with any of the Cas9 proteins of the present disclosure, whereby one or more DNA sequences of the cell are modified. In some embodiments, the cell comprises a eukaryotic cell, a yeast cell, a plant cell, a mammalian cell, or a combination thereof. [Brief description of the drawings]
[0036] [Figure 1]A comparison of the specificity of wild-type SpCas9 and FnCas9 proteins with single-base mismatch sgRNAs for the KRAS target sequence is shown. As further described in Example 1, sgRNAs with single-base mismatches at different positions in the target sequence were constructed and the specificity of Cas9 proteins was determined by measuring the ability to cleave the target KRAS sequence with the different sgRNAs using an in vitro cleavage assay. Table 5 shows the sequences of the various KRAS targeting sgRNAs that were tested (see experiment labeled "Comparison of specificity_SpCas9 vs. FnCas9"). A shows the results for wild-type SpCas9 proteins with a control KRAS sgRNA (no mismatch with the target; "T") (i.e., "KRAS-T" in Table 5) or any one of the KRAS sgRNAs KRAS-1 to KRAS-10 identified in Table 5. B shows results for wild-type SpCas9 protein with any of the KRAS sgRNAs KRAS-11 to KRAS-20 identified in Table 5. C shows results for wild-type FnCas9 protein with a control KRAS sgRNA (no mismatch to target; "T") ("KRAS-T" in Table 5) or any of the KRAS-1 to KRAS-10 sgRNAs. D shows results for wild-type FnCas9 protein with any of the KRAS-11 to KRAS-20 sgRNAs.
[0037] [Figure 2A] A comparison of the specificity of the following SpCas9 protein variants with single-base mismatched sgRNAs for the KRAS target sequence is shown. Results for SpCas9-HF1 protein are shown. Results shown on the left are for the use of a control KRAS sgRNA (no mismatch with target; "T") ("KRAS-T" in Table 5) or either KRAS-1 through KRAS-10 sgRNAs (as identified in Table 5), and results shown on the right are for the use of either KRAS-11 through KRAS-20 sgRNAs (as identified in Table 5). The KRAS sgRNAs are the same as those described in Figures 1A-1D. [Figure 2B]Figure 1 shows a comparison of the specificity of the following SpCas9 protein variants with single-base mismatched sgRNAs for the KRAS target sequence. Results for SpCas9-HF4 protein are shown. Results shown on the left are for the use of a control KRAS sgRNA (no mismatch with target; "T") ("KRAS-T" in Table 5) or either KRAS-1 through KRAS-10 sgRNAs (as identified in Table 5), and results shown on the right are for the use of either KRAS-11 through KRAS-20 sgRNAs (as identified in Table 5). KRAS sgRNAs are the same as those described in Figures 1A-1D. [Figure 2C] A comparison of the specificity of the following SpCas9 protein variants with single-base mismatched sgRNAs for the KRAS target sequence is shown. Results for eSpCas9(1.0) protein are shown. Results shown on the left are for the use of a control KRAS sgRNA (no mismatch with target; "T") ("KRAS-T" in Table 5) or either KRAS-1 through KRAS-10 sgRNAs (as identified in Table 5), and results shown on the right are for the use of either KRAS-11 through KRAS-20 sgRNAs (as identified in Table 5). The KRAS sgRNAs are the same as those described in Figures 1A-1D. [Figure 2D] A comparison of the specificity of the following SpCas9 protein variants with single-base mismatched sgRNAs for the KRAS target sequence is shown. Results for eSpCas9(1.1) protein are shown. Results shown on the left are for the use of a control KRAS sgRNA (no mismatch with target; "T") ("KRAS-T" in Table 5) or either KRAS-1 through KRAS-10 sgRNAs (as identified in Table 5), and results shown on the right are for the use of either KRAS-11 through KRAS-20 sgRNAs (as identified in Table 5). The KRAS sgRNAs are the same as those described in Figures 1A-1D.
[0038] [Diagram 3]FIG. 1 shows heat maps illustrating a quantitative comparison of the cleavage efficiency data shown in FIGS. 1A-1D and 2A-2D.
[0039] [Figure 4A] Figure 1 shows the specificity of different FnCas9 protein variants with single-base mismatched sgRNAs for the KRAS target sequence. Figure 2 shows a heat map comparing the cleavage efficiency of different FnCas9 protein variants. To generate the different FnCas9 protein variants shown, alanine substitutions were made at one of the following residues in the wild-type FnCas9 protein: K405, R455, K546, K561, K562, K564, K566, K578, K579, R618, R622, K664, R721, R785, K786, K788, K789, R807, K808, R849, R851, R852, R853, R854, R855, R856, R857, R858, R859, R860, R861, R862, R863, R864, R865, R866, R867, R868, R869, R870, R871, R872, R873, R874, R875, R876, R877, R878, R879, R879, R878, R879, R879, R880, R881, R882, R883, R884, R885, R886, R885, R886, R889, R882, R881, R882, R883, R884, R885, R885, R886 ... 56, K914, K917, R919, R920, K921, K922, R926, K934, K936, R939, K941, K945, R1047, R1131, R1137, K1142, K1152, K1155, R1178, K1189, K1198, K1206, K1213, K1223, R1226, K1227, K1228, or R1241 (see top of heatmap). Wild-type FnCas9 protein was used as a control (WT). The KRAS sgRNAs used to generate the data are indicated to the left of the heatmap and correspond to those described in Figure 1A (i.e., control KRAS-T and KRAS-1 to KRAS-10 sgRNAs). [Figure 4B] Figure 1 shows the specificity of different FnCas9 protein variants with single-base mismatched sgRNAs for the KRAS target sequence. A bar graph comparing the specificity scores of different FnCas9 protein variants is shown. Specificity scores were calculated as the mean squared difference between on-target and off-target cleavage rates. The specificity scores shown were normalized to the specificity score of the wild-type FnCas9 protein (i.e., specificity score = 1). The horizontal line represents a specificity score of 1. The specific amino acid substitutions are indicated at the bottom of the bar graph.
[0040] [Diagram 5] 1 shows the crystal structure of the FnCas9 protein, and identifies exemplary amino acid residues that can be modified to increase the specificity of the FnCas9 protein.
[0041] [Figure 6] The following FnCas9 protein variants have demonstrated the ability to cleave target NRAS gene sequences using single-base mismatch sgRNAs: (i) an FnCas9 protein with a single modification at residue K1189 (FnCas9-K1189A); (ii) an FnCas9 protein with a single modification at residue R1241 (FnCas9-R1241A); (iii) an FnCas9 protein with dual modifications at residues K1189 and R1241 (FnCas9-K1189A,R1241A, also referred to herein as "FnCas9-AF1"); and (iv) an FnCas9 protein with triple modifications at residues R785, K1189, and R1241 (FnCas9-R785A, K1189A, R1241A, also referred to herein as "FnCas9-AF2"). Table 5 shows the sequences of the various NRAS-targeting sgRNAs tested. The nucleotides (i.e., G, C, A, U) shown on the left side of each heatmap correspond to the specific single-base mismatches made to the sgRNA tested.
[0042] [Figure 7A]Heat map analysis of the cleavage efficiency of additional FnCas9 protein variants containing single-, double-, or triple-base mutations by single-base mismatch sgRNAs against the KRAS target sequence. The table above the figure heat map shows the various amino acid modifications made to generate the various FnCas9 protein variants tested. The sequences of the various sgRNAs used are shown to the left of the heat map. The KRAS sgRNAs are as follows (from top to bottom): KRAS-wild type (SEQ ID NO: 48), KRAS-1 (SEQ ID NO: 49), KRAS-2 (SEQ ID NO: 50), KRAS-3 (SEQ ID NO: 51), KRAS-4 (SEQ ID NO: 52), KRAS-5 (SEQ ID NO: 53), KRAS-6 (SEQ ID NO: 54), KRAS-7 (SEQ ID NO: 55), KRAS-8 (SEQ ID NO: 56), KRAS-9 (SEQ ID NO: 57), KRAS-1 (SEQ ID NO: 58), KRAS-2 (SEQ ID NO: 59), KRAS-3 (SEQ ID NO: 60), KRAS-4 (SEQ ID NO: 61), KRAS-5 (SEQ ID NO: 62), KRAS-6 (SEQ ID NO: 63), KRAS-7 (SEQ ID NO: 64), KRAS-8 (SEQ ID NO: 65), KRAS-9 (SEQ ID NO: 66), KRAS-10 (SEQ ID NO: 67), KRAS-11 (SEQ ID NO: 68), KRAS-12 (SEQ ID NO: 69), KRAS-13 (SEQ ID NO: 70), KRAS-14 (SEQ ID NO: 71), KRAS-15 (SEQ ID NO: 72), KRAS-16 (SEQ RAS-10 (SEQ ID NO:58), KRAS-11 (SEQ ID NO:59), KRAS-12 (SEQ ID NO:60), KRAS-13 (SEQ ID NO:61), KRAS-14 (SEQ ID NO:62), KRAS-15 (SEQ ID NO:63), KRAS-16 (SEQ ID NO:64), KRAS-17 (SEQ ID NO:65), KRAS-18 (SEQ ID NO:66), KRAS-19 (SEQ ID NO:67), and KRAS-20 (SEQ ID NO:68). [Figure 7B]A heat map analysis of the cleavage efficiency of additional FnCas9 protein variants containing single-, double-, or triple-base mutations by single-base mismatch sgRNAs against the NRAS target sequence is shown. The table above the heat map in the figure shows the various amino acid modifications made to generate the various FnCas9 protein variants tested. The sequences of the various sgRNAs used are shown to the left of the heat map. The NRAS sgRNAs are as follows (from top to bottom): NRAS-wild type (SEQ ID NO: 105), NRAS-1-G (SEQ ID NO: 106), NRAS-2-C (SEQ ID NO: 109), NRAS-3-T (SEQ ID NO: 112), NRAS-4-G (SEQ ID NO: 115), NRAS-5-T (SEQ ID NO: 118), NRAS-6-A (SEQ ID NO: 121), NRAS-7-T (SEQ ID NO: 124), NRAS-8-C (SEQ ID NO: 127), NRAS-9-C (SEQ ID NO: 130), NRAS -10-A (SEQ ID NO: 133), NRAS-11-G (SEQ ID NO: 136), NRAS-12-T (SEQ ID NO: 139), NRAS-13-A (SEQ ID NO: 142), NRAS-14-T (SEQ ID NO: 145), NRAS-15-G (SEQ ID NO: 148), NRAS-16-T (SEQ ID NO: 151), NRAS-17-C (SEQ ID NO: 154), NRAS-18-C (SEQ ID NO: 157), NRAS-19-A (SEQ ID NO: 160), and NRAS-20-A (SEQ ID NO: 163).
[0043] [Figure 8]Figure 5 shows a comparison of the specificity of wild-type FnCas9 and FnCas9-AF2 proteins with single-base mismatched sgRNAs for KRAS and EGFR target sequences. A shows the KRAS (top) (GTAGTTGGAGCTGGTGGCGT; SEQ ID NO: 249) and EGFR (bottom) (CAGATTTTGGGCTGGCCAAA; SEQ ID NO: 250) target sequences. B shows the cleavage efficiency data of wild-type FnCas9 for KRAS (top heat map) and EGFR (bottom heat map) sequences. C shows the cleavage efficiency data of FnCas9-AF2 for KRAS (top heat map) and EGFR (bottom heat map) sequences. Table 5 shows the sequences of the various KRAS- and EGFR-targeted sgRNAs tested. The nucleotides (i.e., G, C, A, U) shown to the left of each heat map correspond to the specific single-base mismatches made to the sgRNAs tested.
[0044] [Figure 9A] Digenome-seq analysis comparing genome-wide unbiased off-target occurrence observed after digestion of genomic DNA of HEK293T cells with several Cas9 protein variants. Manhattan plots showing DNA cleavage positions generated by various Cas9 proteins: (i) wild-type SpCas9 protein (top left plot), (ii) wild-type FnCas9 protein (top right plot), (iii) eSpCas9(1.1) (center left plot), (iv) FnCas9-AF1 (center right plot), (v) SpCas9-H4 (bottom left plot), and (vi) FnCas9-AF2 (bottom right plot). [Figure 9B]Digenome-seq analysis comparing genome-wide unbiased off-target occurrence observed after digestion of genomic DNA of HEK293T cells with several Cas9 protein variants. Venn diagram showing the number of off-targets generated by SpCas9-wild type, SpCas9-HF4, FnCas9-wild type, FnCas9-AF2 (left panel) and eSpCas9(1.1), SpCas9-HF4, FnCas9-AF1, FnCas9-AF2 (right panel).
[0045] [Figure 10] FIG. 1 is a schematic diagram showing the use of highly specific Cas9 proteins (e.g., as described herein) to enrich for mutant DNA present in a sample containing cell-free DNA (cfDNA). Specifically, the sgRNA is designed to cleave wild-type DNA (major allele DNA in cfDNA; "wtDNA"), while mutant DNA (minor allele DNA in cfDNA; "mtDNA") remains uncleaved, resulting in enrichment of mutant DNA. As shown, mutant DNA can be classified into one of two types based on the location of the mutation present: (1) type I-mutation(s) within the PAM site, and (2) type II-mutation(s) within the sgRNA target sequence but outside the PAM site. CRISPR-Cas proteins known in the art are generally capable of recognizing mutations within the PAM site. As a result, such Cas proteins can effectively distinguish type I mtDNA from wtDNA, but cannot distinguish type II mtDNA, resulting in cleavage of both wtDNA and type II mtDNA (see diagram below dashed line). In contrast, the Cas9 protein described herein (which has a high degree of specificity) is able to effectively distinguish between both type I and type II mtDNA and cleave only wtDNA.
[0046] [Figure 11]The proportion of applicable target variants that are observed to occur within (type I variants, light grey bars) or outside (type II variants, dark grey bars) the PAM region and can be targeted by CRISPR enrichment is shown for various cancer types. Cancer types shown include lung, breast, liver, pancreatic, and thyroid cancers.
[0047] [Figure 12] Figure 1 shows the enrichment of type II EGFR or KRAS mutant DNA sequences after digestion of a mixture containing mutant and wild-type DNA sequences with FnCas9-AF2 protein. Enrichment is shown as the ratio of mutant to wild-type DNA sequences observed after digestion (enriched mutant ratio). Mutated EGFR sequences had one of the following mutations within the sgRNA target sequence, but outside the PAM site: T790M (A), L858R (B), or exon 19 deletion (C). Mutated KRAS sequences had a G12D mutation (D), which was located outside the PAM site but present within the sgRNA target sequence. The mixtures before digestion contained 5%, 1%, 0.1%, or 0% mutant DNA. Samples that had not been digested with Cas9 protein were used as controls ("Cas9(-)").
[0048] [Figure 13] Heatmap analysis showing CRISPR-Cas9-mediated enrichment of mutant DNA sequences in human cancer patient samples. Blood and tissue samples from 10 patients with non-small cell lung cancer, 9 with stage I and 1 with stage II (identified as P03, P07-P12, P18, P22, and P26), were obtained and analyzed for the presence of 1,056 genomic variants. Results are shown as follows: (1) non-enriched tissue sample (i.e., original tissue; "OT"), (2) tissue sample enriched with highly specific Cas9 protein (i.e., CRISPR-enriched tissue; "CT"), (3) non-enriched blood sample (i.e., original cfDNA; "Oc"), and (4) blood sample enriched with highly specific Cas9 protein (i.e., CRISPR-enriched cfDNA; "Cc").
[0049] [Figure 14] Comparison of the specificity of wild-type FnCas9 and FnCas9-AF2 proteins with single-base mismatch sgRNAs for NRAS target sequences. Table 5 shows the sequences of the various NRAS-targeting sgRNAs (NRAS-sgRNAs #1 to #20) tested. A and C show the results of wild-type SpCas9 and FnCas9-AF2 proteins with the following NRAS sgRNAs (identified in Table 5), respectively: (T) NRAS-wild type (no mismatch with target), (1) NRAS-1-G, (2) NRAS-2-C, (3) NRAS-3-T, (4) NRAS-4-G, (5) NRAS-5-T, (6) NRAS-6-A, (7) NRAS-7-T, (8) NRAS-8-C, (9) NRAS-9-C, and (10) NRAS-10-A. Panels B and D show the results for wild-type SpCas9 and FnCas9-AF2 proteins, respectively, with the following NRAS sgRNAs (identified in Table 5): (11) NRAS-11-G, (12) NRAS-12-T, (13) NRAS-13-A, (14) NRAS-14-T, (15) NRAS-15-G, (16) NRAS-16-T, (17) NRAS-17-C, (18) NRAS-18-C, (19) NRAS-19-A, and (20) NRAS-20-A.
[0050] [Figure 15] 1 shows a heatmap analysis showing CRISPR-Cas9-mediated enrichment of 392 genomic variants that meet the criteria defined in categories 1, 3, 4, and 6 (listed in Table 6) in non-small cell lung cancer (NSCLC) patient samples (listed in FIG. 13). Subjects were divided into four groups according to patient number, with 10 subjects in each group. The groups were arranged in order: original tissue (OT), CRISPR-enriched tissue (CT), original cfDNA (Oc), and CRISPR-enriched cfDNA (Cc). Statistical significance was shown between the OT and CT groups (tissue correlation; Tr) and between the Oc and Cc groups (cfDNA correlation; cr). The higher the fold change (FC) value, the higher the value for the CRISPR-enriched (CE) group, and the higher the -log10 p-value (PV) value, the higher the statistical significance.
[0051] [Figure 16A] Heatmap analysis showing CRISPR-Cas9-mediated enrichment of a specific subset of the 392 genomic variants listed in Figure 15 based on statistical significance. Higher fold change (FC) values indicate higher CRISPR enriched (CE) populations, and higher -log10 p-values (PV) indicate higher statistical significance. Results are shown for 11 variants with p-values < 0.05 and FC > 0.1 compared between OT and CT (tissue correlation; Tr). Different variants are shown immediately to the right of the heatmap. [Figure 16B] Heatmap analysis showing CRISPR-Cas9-mediated enrichment of a specific subset of the 392 genomic variants listed in Figure 15 based on statistical significance. Higher fold change (FC) values indicate higher CRISPR enriched (CE) populations, and higher -log10 p-values (PV) indicate higher statistical significance. Results are shown for 17 variants with p-values < 0.05 and FC > 0.1 compared between Oc and Cc (cfDNA correlation; cr). Different variants are shown immediately to the right of the heatmap.
[0052] [Figure 17A] The box plots of the heat map analysis shown in Figure 13 are shown. Specifically, the box plots show the statistical significance in the detection of genomic variants among various NSCLC patient samples. The subjects were divided into four groups in the order of patient numbers, with 10 subjects in each group. The original tissue (OT), original cfDNA (Oc), CRISPR enriched (CE) tissue (CT), and CE cfDNA (Cc) groups were arranged in order. In each group, the statistical groups and their p values are shown. The Kruskal wallis test was performed in the four groups, and statistically significant p values (<2.2e-16) were detected. Overall view of 1,056 variants. In each group, 1,056*10=10,560 plots were listed. [Figure 17B]The box plots of the heat map analysis shown in Figure 15 are shown. Specifically, the box plots show the statistical significance in the detection of genomic mutations among various NSCLC patient samples. The subjects were divided into four groups in the order of patient numbers, with 10 subjects in each group. The original tissue (OT), original cfDNA (Oc), CRISPR enriched (CE) tissue (CT), and CE cfDNA (Cc) groups were arranged in order. In each group, the statistical groups and their p values are shown. The Kruskal wallis test was performed in the four groups, and statistically significant p values (<2.2e-16) were detected. After filtering by conditions, a total of 392 variants and 392*10=3,920 plots were used to create the box plots. [Figure 17C] The box plots of the heat map analysis shown in Figure 16A are shown. Specifically, the box plots show the statistical significance in the detection of genomic mutations between various NSCLC patient samples. The subjects were divided into four groups in the order of patient numbers, with 10 subjects in each group. The original tissue (OT), original cfDNA (Oc), CRISPR enriched (CE) tissue (CT), and CE cfDNA (Cc) groups were arranged in order. In each group, the statistical groups and their p-values are shown. The Kruskal wallis test was performed in the four groups, and statistically significant p-values (<2.2e-16) were detected. For the tissue comparison, a total of 11 variants from 180 plots were used for the box plots. [Figure 17D] The box plot of the heat map analysis shown in Figure 16B is shown. Specifically, the box plot shows the statistical significance in the detection of genomic mutations between various NSCLC patient samples. The subjects were divided into four groups in the order of patient numbers, with 10 subjects in each group. The original tissue (OT), original cfDNA (Oc), CRISPR enriched (CE) tissue (CT), and CE cfDNA (Cc) groups were arranged in order. In each group, the statistical group and its p value are shown. The Kruskal wallis test was performed in the four groups, and a statistically significant p value (<2.2e-16) was detected. In the cfDNA comparison, a total of 17 variants of 276 plots were used for the box plot.
[0053] [Figure 18A] Correlation plots of all genomic variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment with post-enrichment. OT samples before enrichment are compared with Oc samples. Additionally, correlation plots of all observed genomic variants with filtered variants are shown. [Figure 18B] Correlation plots of all genomic variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment with post-enrichment. OT and CT samples after enrichment are compared. Additionally, correlation plots of all observed genomic variants with filtered variants are shown. [Figure 18C] Correlation plots of all genomic variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment with post-enrichment. Oc and Cc groups within cfDNA samples are compared. Additionally, correlation plots of all observed genomic variants with filtered variants are shown. [Figure 18D]Correlation plots of all genomic variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (tissue of origin (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)) are shown. In all figures shown, black dots represent higher ratios when comparing pre-enrichment with post-enrichment. Comparing the CT and Cc groups confirms the CE patterns observed in the cfDNA samples. Additionally, correlation plots of all observed genomic variants with filtered variants are shown. [Figure 18E] Correlation plots of all genomic variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment with post-enrichment. OT samples before enrichment are compared with Oc samples. Additionally, correlation plots of all observed genomic variants with filtered variants are shown. [Figure 18F] Correlation plots of all genomic variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment with post-enrichment. OT and CT samples after enrichment are compared. Additionally, correlation plots of all observed genomic variants with filtered variants are shown. [Figure 18G]Correlation plots of all genomic variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment with post-enrichment. Oc and Cc groups within cfDNA samples are compared. Additionally, correlation plots of all observed genomic variants with filtered variants are shown. [Figure 18H] Correlation plots of all genomic variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (tissue of origin (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)) are shown. In all figures shown, black dots represent higher ratios when comparing pre-enrichment with post-enrichment. Comparing the CT and Cc groups confirms the CE patterns observed in the cfDNA samples. Additionally, correlation plots of all observed genomic variants with filtered variants are shown. [Figure 18I] Correlation plots of genome-wide variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment to post-enrichment. Pre-enrichment OT and Oc samples are compared. Correlation plots filtered by groups for tissue and cfDNA before and after CE are shown. [Figure 18J] Correlation plots of genome-wide variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment to post-enrichment. OT and CT samples after enrichment are compared. Correlation plots filtered by groups for tissue and cfDNA before and after CE are shown. [Figure 18K] Correlation plots of genome-wide variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment to post-enrichment. Oc and Cc groups within cfDNA samples are compared. Correlation plots filtered by groups before and after CE for tissue and cfDNA are shown. [Figure 18L] Correlation plots of genome-wide variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment with post-enrichment. Comparing CT and Cc groups confirms the CE patterns observed in cfDNA samples. Correlation plots filtered by groups for tissue and cfDNA before and after CE are shown. [Figure 18M] Correlation plots of genome-wide variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment to post-enrichment. Pre-enrichment OT and Oc samples are compared. Correlation plots filtered by groups for tissue and cfDNA before and after CE are shown. [Figure 18N] Correlation plots of genome-wide variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment to post-enrichment. OT and CT samples after enrichment are compared. Correlation plots filtered by groups for tissue and cfDNA before and after CE are shown. [Figure 18O] Correlation plots of genome-wide variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment to post-enrichment. Oc and Cc groups within cfDNA samples are compared. Correlation plots filtered by groups before and after CE for tissue and cfDNA are shown. [Figure 18P] Correlation plots of genome-wide variants detected before and after CRISPR-Cas9-mediated enrichment (CE) in various NSCLC patient samples (original tissue (OT), original cfDNA (Oc), CE tissue (CT), and CE cfDNA (Cc)). In all figures shown, black dots represent higher ratios when comparing pre-enrichment with post-enrichment. Comparing CT and Cc groups confirms the CE patterns observed in cfDNA samples. Correlation plots filtered by groups for tissue and cfDNA before and after CE are shown.
[0054] [Figure 19] The crystal structure of the FnCas9 protein shows the amino acids that can interact with the phosphate backbone of either the target or non-target DNA strand. Specific amino acids (49 in total) are identified in FIG. 4B. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0055] The present disclosure relates to Cas9 proteins that have been modified to include one or more features that differ (e.g., structurally and / or functionally) from reference Cas9 proteins known in the art (e.g., wild-type S. pyogenes Cas9 protein). For example, as further described herein, the Cas9 proteins of the present disclosure include one or more amino acid modifications that increase the specificity of the Cas9 protein. Thus, compared to reference Cas9 proteins, the Cas9 proteins described herein can more accurately recognize and discriminate between base mismatches within a target gene sequence. As shown herein, due to such enhanced specificity, in some embodiments, the Cas9 proteins described herein are associated with significantly reduced off-target effects (e.g., off-target binding, editing, and / or cleavage activity). In some embodiments, the Cas9 proteins described herein are associated with increased on-target effects (e.g., on-target binding, editing, and / or cleavage activity). In some embodiments, the Cas9 proteins described herein are associated with both reduced off-target effects and increased on-target effects. Further embodiments of the disclosure are provided throughout the application.
[0056] Before describing the present disclosure in more detail, it is to be understood that the present disclosure is not limited to the specific compositions or process steps described, which may, of course, vary. As will be apparent to one of ordinary skill in the art upon reading this disclosure, each of the individual aspects described and illustrated herein has separate components and features that can be readily separated from or combined with the features of any of the other several aspects without departing from the scope or spirit of the present disclosure. Any method described can be carried out in the order of events described, or in any other order that is logically possible.
[0057] The headings provided herein are not limitations of the various aspects of the disclosure, which can be defined by reference to the specification as a whole. It is also to be understood that the terminology used herein is for the purpose of describing particular aspects only, and is not intended to be limiting, since the scope of the disclosure will be limited only by the appended claims.
[0058] I. Definition In order that this disclosure may be more readily understood, certain terms are first defined. As used in this application, unless otherwise stated herein, each of the following terms shall have the meaning set forth below. Additional definitions are set forth throughout this application.
[0059] The term "a" or "an" entity refers to one or more of that entity, for example, "a Cas9 protein" is understood to refer to one or more Cas9 proteins. Thus, the terms "a" (or "an"), "one or more," and "at least one" can be used interchangeably herein.
[0060] Furthermore, as used herein, "and / or" should be understood as a specific disclosure of each of two particular features or elements with or without the other. Thus, the term "and / or" used herein in phrases such as "A and / or B" is intended to include "A and B," "A or B," "A" (single), and "B" (single). Similarly, the term "and / or" used in phrases such as "A, B, and / or C" is intended to encompass each of the following aspects: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (single); B (single); and C (single).
[0061] Where an embodiment is described herein using the language "comprising," it is understood that similar embodiments are also provided that are otherwise described with respect to "consisting of" and / or "consisting essentially of." As used herein, "comprising" is synonymous with "including," "containing," or "characterized by" and is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. As used herein, "consisting of" excludes elements, steps, or ingredients not specified in the claim element. As used herein, "consisting essentially of" does not exclude materials or steps that do not materially affect the basic and novel characteristics of the claim.
[0062] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this disclosure pertains. For example, Concise Dictionary of Biomedicine and Molecular Biology, Juo, Pei-Show, 2nd ed., 2002, CRC Press, The Dictionary of Cell and Molecular Biology, 5th ed., 2013, Academic Press, and Oxford Dictionary Of Biochemistry and Molecular Biology, 2nd ed., 2008, Oxford University Press provide those of ordinary skill in the art with a general dictionary of many of the terms used in this disclosure.
[0063] Units, prefixes, and symbols are written in the format accepted by the International System of Units (SI). Numeric ranges include the numbers that define the range. When a range of values is listed, it is understood that each intervening integer value between the upper and lower limit values listed for that range, and each fractional part thereof, is also expressly disclosed, along with each subrange between such values. The upper and lower limits of any range may be independently included or excluded from the range, and each range that includes either the upper or lower limit, each range that does not include either, or each range that includes both, is also encompassed by the present disclosure. Thus, ranges listed herein are understood to be shorthand for all values within the range, including the recited endpoints. For example, a range of 1 to 10 is understood to include any number, combination of numbers, or subrange from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10.
[0064] When a value is explicitly stated, it is to be understood that values that are about the same number or amount as the stated value (e.g., 10) (e.g., ±10%) are also included within the scope of the disclosure. When a combination is disclosed, each subcombination of the elements of the combination is also specifically disclosed and is included within the scope of the disclosure. Conversely, when different elements or groups of elements are individually disclosed, combinations thereof are also disclosed. When any element of the disclosure is disclosed as having multiple alternatives, examples of the disclosure in which each alternative is excluded, alone or in any combination with other alternatives, are also disclosed herein, and two or more elements of the disclosure may have such an exclusion, and all combinations of elements with such exclusions are disclosed herein.
[0065] Nucleotides are represented by their commonly accepted single letter codes. Unless otherwise stated, nucleotide sequences are written from left to right in the 5' to 3' direction. Nucleotides are represented herein by the commonly known single letter symbols of nucleotides recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Thus, "a" represents adenine, "c" represents cytosine, "g" represents guanine, "t" represents thymine, and "u" represents uracil. It should be understood that T and U in the disclosed sequences are interchangeable depending on whether the sequence is DNA or RNA. For example, target sequences are presented in this disclosure as DNA (A / T / C / G), while guide RNAs are presented as RNA (A / U / C / G).
[0066] Amino acid sequences are written from left to right in the amino to carboxyl terminus direction. Amino acids are referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission.
[0067] As used herein, the term "Cas9 protein" (including any variants thereof) refers to a polypeptide that can interact with a guide RNA (gRNA) molecule and, in coordination with the gRNA molecule, localize to a site that includes a target sequence and, in some embodiments, a PAM. As further described elsewhere in this disclosure, the Cas9 proteins described herein have been modified, altered, or engineered to provide one or more properties (e.g., enhanced specificity). Thus, the terms "Cas9 protein as described herein," "Cas9 protein provided herein," and "Cas9 protein of the present disclosure" (including variants thereof) refer to a Cas9 protein that includes one or more amino acid modifications as described herein and has improved specificity compared to a reference Cas9 protein (e.g., a wild-type Cas9 protein). Furthermore, unless otherwise indicated, the terms "modified," "engineered," and "modified" are used interchangeably and, when used in this context, refer only to differences from a reference or native sequence and do not impose limitations of a particular process or origin. Additional disclosure regarding the Cas9 proteins of the present disclosure is provided elsewhere herein.
[0068] As described herein, in some embodiments, the Cas9 protein that can be modified using the disclosure provided herein comprises a Cas9 protein from Francisella novicida ("FnCas9"). The amino acid sequence of the wild-type FnCas9 protein is shown in Table 1 (below) (i.e., SEQ ID NO:1). Thus, in some embodiments, the Cas9 protein provided herein comprises an amino acid sequence that differs from SEQ ID NO:1 by one or more amino acids. In some embodiments, the amino acid sequences of the Cas9 proteins provided herein have less than about 99.999%, less than about 99.998%, less than about 99.997%, less than about 99.996%, less than about 99.995%, less than about 99.994%, less than about 99.993%, less than about 99.992%, less than about 99.991%, less than about 99.99%, less than about 99.8%, less than about 99.7%, less than about 99.6%, less than about 99.5%, less than about 99.4%, less than about 99.3%, less than about 99.2%, less than about 99.1%, less than about 99%, less than about 98%, less than about 97%, less than about 96%, or less than about 95% sequence identity to the amino acid sequence set forth in SEQ ID NO:1. It will be apparent to one of skill in the art that the disclosure provided herein can be applied to any suitable Cas9 protein (wild type and variants thereof) known in the art. Non-limiting examples of such Cas9 proteins are described, for example, in U.S. Publication No. 2018 / 0051281A1, which is incorporated herein by reference in its entirety.
[0069] In some embodiments, the modifiable Cas9 protein includes the Cas9 protein from Streptococcus pyogenes ("SpCas9"). The amino acid sequence of the wild-type SpCas9 protein is shown in Table 1 (below) (i.e., SEQ ID NO:244). Thus, in some embodiments, the Cas9 proteins provided herein include an amino acid sequence that differs from SEQ ID NO:244 by one or more amino acids. In some embodiments, the amino acid sequences of the Cas9 proteins provided herein have less than about 99.999%, less than about 99.998%, less than about 99.997%, less than about 99.996%, less than about 99.995%, less than about 99.994%, less than about 99.993%, less than about 99.992%, less than about 99.991%, less than about 99.99%, less than about 99.8%, less than about 99.7%, less than about 99.6%, less than about 99.5%, less than about 99.4%, less than about 99.3%, less than about 99.2%, less than about 99.1%, less than about 99%, less than about 98%, less than about 97%, less than about 96%, or less than about 95% sequence identity to the amino acid sequence set forth in SEQ ID NO:244.
[0070] Additional examples of Cas9 proteins suitable for this disclosure are described elsewhere in this disclosure. [Table 1-1] [Table 1-2]
[0071] As used herein, the term "eSpCas9(1.1)" refers to a modified SpCas9 protein having the following amino acid mutations: K848A, K1003A, and R1060A. "eSpCas9(1.0)" refers to a modified SpCas9 protein having the following amino acid mutations: K810A, K1003A, and R1060A. The amino acid sequences of eSpCas9(1.1) and eSpCas9(1.0) are shown in Table 8 (i.e., SEQ ID NO:245 and SEQ ID NO:246, respectively). Additional details regarding eSpCas9(1.1) and eSpCas9(1.0) are described, for example, in Slaymaker et al., Science 351(6268):84-88 (Jan. 2016).
[0072] As used herein, the term "SpCas9-HF1" refers to a modified SpCas9 protein having the following amino acid mutations: N497A, R661A, Q695A, and Q926A. The term "SpCas9-HF4" refers to a modified SpCas9 protein that includes the amino acid mutations of SpCas9-HF1 and further has a Y450A amino acid mutation (i.e., has the following five mutations: N497A, Y450A, R661A, Q695A, and Q926A). The amino acid sequences of SpCas9-HF1 and Sp-Cas9-HF2 are shown in Table 8 (i.e., SEQ ID NOs: 247 and 248, respectively). See also Kleinstiver et al., Nature 529:490-495 (2016).
[0073] The term "guide RNA" refers to an RNA molecule (or an entire group of RNA molecules) that binds to a Cas protein and supports targeting the Cas protein to a specific location within a target polynucleotide (e.g., a target sequence). A guide RNA may include a crRNA segment and a tracrRNA segment. As used herein, the term "crRNA" or "crRNA segment" refers to an RNA molecule or a portion thereof that includes a polynucleotide-targeting guide sequence, a stem sequence, and optionally a 5-overhang sequence. As used herein, the term "tracrRNA" or "tracrRNA segment" refers to an RNA molecule or a portion thereof that includes a protein-binding segment (e.g., the protein-binding segment can interact with a CRISPR-associated protein, such as Cas9). The term "guide RNA" encompasses single guide RNAs (sgRNAs) in which the crRNA segment and the tracrRNA segment are located in the same RNA molecule. The term "guide RNA" also collectively encompasses a group of two or more RNA molecules in which the crRNA segment and the tracrRNA segment are located in separate RNA molecules.
[0074] As is evident from the present disclosure, in combination with Cas9 nuclease (such as those described herein), guide RNAs promote target specificity of the CRISPR / Cas9 system. Some embodiments, such as promoter selection, can provide additional mechanisms for achieving target specificity, for example, selecting a promoter for the polynucleotide encoding the guide RNA that promotes expression in a particular organ or tissue. Thus, the selection of a gRNA suitable for a particular disease, disorder, or condition is also considered and further described herein. As provided herein, in some embodiments, gRNAs useful in the present disclosure can be chemically synthesized to include a specific guide sequence (i.e., a "synthetic gRNA"). For example, in some embodiments, a synthetic gRNA can include one or more base modifications (e.g., nucleotide substitutions) such that the gRNA differs in sequence compared to the corresponding wild-type gRNA. Methods for constructing such synthetic gRNAs are described elsewhere in this disclosure (see, e.g., Example 1) and also in, e.g., Doench, J., et al., Nature biotechnology 32(12):1262-7 (2014); Mohr, S. et al., FEBS Journal 283:3232-38 (2016); Graham, D., et al., Genome Biol. 16:260 (2015); Kelley, M. et al, J Biotechnology 233:74-83 (2016), each of which is incorporated herein by reference in its entirety. In some embodiments, the gRNAs described herein may contain one or more modifications that further enhance the specificity of the Cas9 proteins described herein, such as shortening the length of the target sequence (e.g., 18 nucleotides instead of 20 nucleotides) or adding a guanine to the 5' end of the gRNA. Further examples of such modifications are known in the art.
[0075] As used herein, the term "target polynucleotide" or "target gene" (including variants thereof) refers to a polynucleotide that comprises a target nucleic acid sequence. A target polynucleotide can be single-stranded or double-stranded, and in some embodiments is double-stranded DNA. In some embodiments, a target polynucleotide is single-stranded RNA. As used herein, "target nucleic acid sequence" or "target sequence" refers to a sequence that a gRNA is designed to bind to (e.g., complementary to the guide sequence of the gRNA), where hybridization (or binding) of the target sequence to the guide sequence promotes the formation of a CRISPR complex, which ultimately cleaves the sequence. Unless otherwise stated, a target sequence can include any polynucleotide, such as a DNA or RNA polynucleotide.
[0076] The term "hybridization" or "hybridizing" refers to the process in which fully or partially complementary polynucleotide strands come together under suitable hybridization conditions to form a double-stranded structure or region in which the two constituent strands are linked by hydrogen bonds. As used herein, the term "partial hybridization" includes cases in which the double-stranded structure or region contains one or more bulges or mismatches. Hydrogen bonds are usually formed between adenine and thymine, or adenine and uracil (A and T, or A and U), or cytosine and guanine (C and G), although other non-standard base pairs may also be formed (see, e.g., Adams et al., "The Biochemistry of the Nucleic Acids," 11th ed., 1992).
[0077] The term "CRISPR" refers to clustered regularly interspaced short palindromic repeats (CRISPR). Generally, CRISPR-Cas, CRISPR-Cas9, or CRISPR system, as used in the aforementioned documents, such as US2017 / 0152528, which are incorporated herein by reference in their entirety, refers to a sequence encoding a Cas gene (specifically the Cas9 gene in the case of CRISPR-Cas9), a tracr (transactivating CRISPR) sequence (e.g., tracrRNA or active portion tracrRNA), a tracr-mate sequence (including "direct repeat sequences" and tracrRNA processing portion direct repeat sequences in the context of endogenous CRISPR systems), a guide sequence (also referred to as "spacer" in the context of endogenous CRISPR systems), or "RNA(s)" as that term is used herein (e.g., RNA(s) that guide Cas9, e.g., CRISPR It collectively refers to the transcripts and other elements involved in directing the expression or activity of CRISPR-associated ("Cas") genes, including RNA and transactivating (tracr)RNA or single guide RNA (sgRNA) (chimeric RNA), or other sequences and transcripts derived from the CRISPR locus. In general, CRISPR systems are characterized by elements that promote the formation of CRISPR complexes at the site of the target sequence (also referred to as protospacers for endogenous CRISPR systems).
[0078] The term "protospacer adjacent motif" or "PAM" refers to a nucleotide sequence present in a target double-stranded polynucleotide (located adjacent to the protospacer) that can be recognized by the Cas9 protein. Upon recognizing the PAM, the Cas9 protein opens the double-stranded polynucleotide and determines whether the sequence adjacent to the PAM is complementary to the guide sequence of the gRNA. If the adjacent sequences are complementary, the Cas9 protein cleaves the target polynucleotide. Otherwise, the Cas9 protein proceeds along the target DNA strand searching for additional PAMs. The PAM sequence and its location on the target DNA strand vary depending on the type of CRISPR-Cas system. For example, in the S. pyogenes type II system, the PAM has a NGG consensus sequence that contains two G:C base pairs and is present one base pair downstream of the sequence derived from the protospacer in the target DNA. The PAM sequence is present on the non-complementary strand of the target DNA (protospacer), and the reverse complement of the PAM is located 5' of the target DNA sequence. The PAM sequence can be specific to the system, e.g., the system from which the site-directed engineered protein is derived.
[0079] As used herein, the term "specificity" (also referred to as "fidelity") refers to the ability of a Cas9 protein to specifically recognize and cleave a desired target sequence, but with little or no cleavage of polynucleotides that differ in sequence and / or location from the desired target sequence. Thus, specificity refers to minimizing off-target effects and / or enhancing on-target effects of a Cas9 protein. Activity (e.g., on-target activity and / or off-target activity) of a Cas9 protein as described herein can be assessed using methods provided herein (see, e.g., Example 1) and / or any suitable method known in the art. Non-limiting examples of such methods include in vitro cleavage assays (see, e.g., www.neb.com / protocols / 2014 / 05 / 01 / in-vitro-digestion-of-dna-with-cas9-nuclease-s-pyogenes-m0386, which is incorporated herein by reference in its entirety), Digenome-seq (see, e.g., Kim et al., Nature Methods 12:237-243(2015), which is incorporated herein by reference in its entirety), GUIDE-seq (see, e.g., Tsai et al., Nat Biotechnol 33:187-197(2015), which is incorporated herein by reference in its entirety), CIRCLE-seq (see, e.g., Tsai et al., Nature Methods 14:607-614(2017), which is incorporated herein by reference in its entirety), or ChIP-seq (see, e.g., O'Geen et al., Nucleic Acids Res 43:3389-3404(2015), which is incorporated herein by reference in its entirety). In some embodiments, the rate of off-target effects can be assessed by measuring the rate of indels at the off-target sites.
[0080] As used herein, the term "cleavage" refers to the cleavage of the covalent backbone of a DNA molecule, for example, caused by Cas9 nuclease.Cleavage can be initiated by a variety of methods, including, but not limited to, enzymatic or chemical hydrolysis of phosphodiester bonds.Both single-strand and double-strand cleavage are possible, and double-strand cleavage can occur as a result of two distinct single-strand cleavage events.DNA cleavage can result in the generation of either blunt ends or sticky ends.
[0081] As used herein, the term "target site" or "target sequence" refers to a region of a polynucleotide sequence to which a binding molecule can bind if sufficient conditions for binding (e.g., sufficient complementarity) exist. In some embodiments, a target sequence is a nucleic acid sequence to which a nuclease described herein (e.g., a modified Cas9 protein) binds and / or is cleaved by such a nuclease. In some embodiments, a target sequence is a nucleic acid sequence to which a guide RNA described herein binds. A target site can be single-stranded or double-stranded.
[0082] As will be apparent to one of skill in the art, the target sequence may vary depending on the nuclease utilized. For example, with respect to an RNA-guided nuclease (e.g., the Cas9 protein described herein), the target sequence typically includes a nucleotide sequence complementary to the guide sequence of the guide RNA of the RNA programmable nuclease, and a protospacer adjacent motif (PAM) at the 3' or 5' end adjacent to the guide RNA complementary sequence. More specifically, for the RNA-guided nuclease Cas9, in some embodiments, the target sequence may be about 16-24 base pairs in length plus 3-6 base pairs of PAM (e.g., NNN (N represents any nucleotide)). As shown herein, in some embodiments, the target sequence of the Cas9 protein described herein is 20 base pairs in length (excluding the PAM).
[0083] In some embodiments, a target sequence of a Cas9 protein described herein may comprise the structure [Nz]-[PAM], where each N is independently any nucleotide and z is an integer between 1 and 50. In some embodiments, z is at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, or at least about 50. In some embodiments, Z is 20.
[0084] The term "off-target" refers to binding and cleavage of a polynucleotide by a Cas9 nuclease to an unintended or unexpected region (i.e., a region that is not the target sequence). In some embodiments, a region of a polynucleotide is an off-target region if it differs from the target region / sequence by at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, or at least about 20 or more nucleotides. Due to the increased specificity, as further described elsewhere, in some embodiments, the Cas9 proteins provided herein are associated with reduced off-target events.
[0085] As used herein, the term "on-target" refers to the intended or predicted region of a polynucleotide (e.g., a target sequence) being bound to and cleaved by a Cas9 nuclease.
[0086] The term "biomarker" refers to a protein or nucleic acid (eg, including mutations) that causes and / or associates with the presence of a particular disease or disorder.
[0087] As used herein, the terms "disease," "disorder," and "pathology" (including variations thereof) are used interchangeably to refer to an abnormal condition that adversely affects the structure or function of all or part of a subject and is not necessarily due to direct external damage. Generally, the diseases and disorders described herein are associated with certain signs and symptoms. Furthermore, the diseases and disorders that can be diagnosed and / or treated using the present disclosure are not particularly limited. In some embodiments, the disease or disorder is associated with abnormal expression / activity of a gene (e.g., a variant DNA pattern that is not present in the corresponding gene in a healthy subject), and thus the disease or disorder can be diagnosed using the methods provided herein. In some embodiments, the disease or disorder is associated with abnormal expression / activity of a gene, and thus the disease or disorder can be treated using the methods provided herein (e.g., by deleting or repairing a mutated gene using the Cas9 protein described herein).
[0088] In some embodiments, diseases or disorders that can be diagnosed and / or treated using the present disclosure include cancer. Non-limiting examples of cancer include mesothelioma, cervical cancer, pancreatic cancer, ovarian cancer, squamous cell carcinoma (e.g., epithelial squamous cell carcinoma), lung cancer (e.g., small cell lung cancer (SCLC), non-small cell lung cancer, lung adenocarcinoma, lung squamous cell carcinoma), skin cancer (e.g., basal cell carcinoma (BCC), cutaneous squamous cell carcinoma (cSCC), melanoma, Merkel cell carcinoma (MCC)), peritoneal cancer, hepatocellular carcinoma, gastric cancer (e.g., gastrointestinal cancer), esophageal cancer (e.g., gastroesophageal junction cancer), brain cancer (e.g., For example, glioblastoma), liver cancer (e.g., hepatocellular carcinoma), bladder cancer, hepatocellular carcinoma, breast cancer (e.g., triple-negative breast cancer (TNBC)), colon cancer, rectal cancer, colorectal cancer, endometrial or uterine cancer, salivary gland cancer, renal or kidney cancer (e.g., renal cell carcinoma), prostate cancer, vulvar cancer, thyroid cancer, liver cancer, anal cancer, penile cancer, head and neck cancer (e.g., head and neck squamous cell carcinoma), bile duct cancer, gallbladder cancer, musculoskeletal cancer, or a combination thereof. In some embodiments, the cancer comprises breast cancer, pancreatic cancer, colon cancer, or a combination thereof.
[0089] In some embodiments, diseases or disorders that can be diagnosed and / or treated using the present disclosure include neurodegenerative or neurological disorders, non-limiting examples of which include Alzheimer's disease, Parkinson's disease, amyotrophic lateral sclerosis, Friedreich's ataxia, Huntington's disease, Lew body disease, spinal muscular atrophy, stroke, or a combination thereof.
[0090] As used herein, the term "associated with" refers to a close relationship between two or more entities or properties. For example, when used in this disclosure to describe a diagnosable disease or condition, the term "associated with" means that if a subject exhibits abnormal levels of a biomarker, the subject is more likely to suffer from (i.e., suffer from) the disease or condition. In some embodiments, the abnormal expression causes the disease or condition. In some embodiments, the abnormal expression does not necessarily cause but is correlated with the disease or condition. Non-limiting examples of suitable methods that can be used to determine whether a subject exhibits abnormal expression of a biomarker associated with a disease or condition are provided elsewhere in this disclosure.
[0091] The term "afflicted with" can be used interchangeably with the term "suffering from" and refers to the state of having a disease or condition. In some embodiments, a subject suffering from a disease or condition (e.g., cancer and / or a neurodegenerative disease) exhibits one or more symptoms associated with the disease or condition. However, as will be apparent to one of skill in the art, a subject need not exhibit one or more symptoms to suffer from a disease or disorder disclosed herein (e.g., may have a genetic predisposition to a disease or disorder).
[0092] As used herein, the term "abnormal level" refers to a level (expression and / or activity) that is different (e.g., increased or decreased) from a reference subject that is not afflicted with, for example, a disease or condition described herein (e.g., cancer and / or a neurodegenerative disease). In some embodiments, an abnormal level (e.g., of a biomarker) refers to a level that is increased by at least about 0.1 fold, at least about 0.2 fold, at least about 0.3 fold, at least about 0.4 fold, at least about 0.5 fold, at least about 0.6 fold, at least about 0.7 fold, at least about 0.8 fold, at least about 0.9 fold, at least about 1 fold, at least about 2 fold, at least about 3 fold, at least about 4 fold, at least about 5 fold, at least about 10 fold, at least about 20 fold, at least about 30 fold, at least about 40 fold, at least about 50 fold, at least about 75 fold, at least about 100 fold, at least about 200 fold, at least about 300 fold, at least about 400 fold, at least about 500 fold, at least about 750 fold, or at least about 1,000 fold or more as compared to the corresponding level in a reference subject (e.g., a subject not suffering from a disease or condition described herein). In some embodiments, an abnormal level (e.g., of a biomarker) refers to a level that is reduced by at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% as compared to the corresponding level in a reference subject (e.g., a subject not suffering from a disease or condition described herein).
[0093] As used herein, the term "diagnosis" (and its derivatives) refers to a method that can be used to determine or predict whether a subject is suffering from, suffering from, or at risk (e.g., genetically predisposed) for a given disease or condition, thereby identifying subjects suitable for treatment. In some embodiments, the treatment can be a therapeutic (e.g., administered to a subject exhibiting one or more symptoms associated with a disease or disorder). In some embodiments, the treatment can be a prophylactic (e.g., administered to a subject at risk to prevent and / or reduce the onset of a disease or disorder). As described herein, in some embodiments, one of skill in the art can make a diagnosis based on a biomarker, where the presence, absence, amount, or change in amount of the biomarker indicates the presence, severity, or absence of a condition. The term "diagnosis" does not refer to the ability to determine with 100% accuracy the presence or absence of a particular disease or disorder, or even the ability to determine the likelihood that a given course or outcome will occur. Instead, one of skill in the art will understand that the term "diagnosis" refers to the likelihood that a particular disease or disorder is present in a subject.
[0094] As used herein, the term "administration" (and grammatical variations thereof) refers to the physical introduction of a Therapeutic (e.g., a Cas9 protein described herein) or a composition comprising a Therapeutic into a subject using any of a variety of methods and delivery systems known to those of skill in the art. Various routes of administration include, but are not limited to, intravenous, intraperitoneal, intramuscular, subcutaneous, spinal, or other parenteral routes of administration (e.g., injection or infusion).
[0095] As used herein, "parenteral administration" refers to a method of administration other than enteral and topical administration, usually by injection, including, but not limited to, intravenous, intraperitoneal, intramuscular, intraarterial, intrathecal, intralymphatic, intralesional, intracapsular, intraorbital, intracardiac, intradermal, transtracheal, intratracheal, transpulmonary, subcutaneous, subcuticular, intraarticular, subcapsular, subarachnoid, intraventricular, intravitreal, epidural, and intrasternal injection and infusion, and in vivo electroporation. Alternatively, a therapeutic agent (e.g., a Cas9 protein as described herein) can be administered by a non-parenteral route, such as a topical, epidermal, or mucosal route of administration, e.g., intranasally, orally, intravaginally, rectally, sublingually, or topically. Administration can also be performed, e.g., once, multiple times, and / or over one or more extended periods of time. Administration also includes self-administration and administration by another.
[0096] A "polypeptide" refers to a chain comprising at least two consecutively linked amino acid residues, with no upper limit on the length of the chain. One or more amino acid residues in a protein may contain modifications, such as, but not limited to, glycosylation, phosphorylation, or disulfide bond formation. A "protein" may include one or more polypeptides. Unless otherwise specified, the terms "protein" and "polypeptide" may be used interchangeably.
[0097] The terms "nucleic acid", "nucleic acid molecule", "nucleotide", "nucleotide(s) sequence", and "polynucleotide" may be used interchangeably and refer to ribonucleosides (adenosine, guanosine, uridine, or cytidine; "RNA molecule") or deoxyribonucleosides (deoxyadenosine, deoxyguanosine, deoxythymidine, or deoxycytidine; "DNA molecule") in phosphate polymeric form, either in single-stranded form or double-stranded helices, or any phosphate analogs thereof, such as phosphorothioates and thioesters. A single-stranded nucleic acid sequence refers to single-stranded DNA (ssDNA) or single-stranded RNA (ssRNA). Double-stranded DNA-DNA, DNA-RNA, and RNA-RNA helices are possible. The terms nucleic acid molecule, and particularly DNA or RNA molecule, refer only to the primary and secondary structure of the molecule and are not limited to any particular tertiary form. Thus, this term includes double-stranded DNA found, inter alia, in linear or circular DNA molecules (e.g., restriction fragments), plasmids, supercoiled DNA, and chromosomes. In describing the structure of a particular double-stranded DNA molecule, the sequence may be described herein according to the conventional convention of providing only the sequence in the 5' to 3' direction along the non-transcribed strand of DNA (i.e., the strand having a sequence homologous to mRNA).
[0098] The term "identity" or "sequence identity" refers to the overall monomer conservation between polymer molecules, e.g., between polypeptides or polynucleotides. The term "identical" without any additional modifiers, e.g., "polypeptide A is identical to polypeptide B," means that the polypeptide sequences are 100% identical (100% sequence identity). For example, describing two sequences as "70% identical" is equivalent to describing them as having, e.g., "70% sequence identity."
[0099] Calculation of the percentage (%) of identity of two polypeptide or polynucleotide sequences can be performed, for example, by aligning the two sequences for optimal comparison (e.g., gaps can be introduced into one or both of the first and second polypeptide or polynucleotide sequences for optimal alignment, and non-identical sequences can be ignored for comparison). In some embodiments, the length of the sequence aligned for comparison purposes is at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or about 100% of the length of the reference sequence. The amino acids, or in the case of polynucleotides, bases at corresponding amino acid positions are then compared.
[0100] If a position in the first sequence is occupied by the same amino acid or nucleotide as the corresponding position in the second sequence, the molecules are identical at that position. The percentage of identity between two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps that need to be introduced for optimal alignment of the two sequences and the length of each gap. The comparison of sequences and the determination of the percentage of identity between two sequences can be accomplished using a mathematical algorithm.
[0101] Suitable software programs that can be used to align different sequences (e.g., polynucleotide sequences) are available from various sources. One suitable program for determining percent sequence identity is bl2seq, which is part of the BLAST package of programs available from the U.S. government's National Center for Biotechnology Information BLAST website (blast.ncbi.nlm.nih.gov). Bl2seq uses the BLASTN or BLASTP algorithm to perform a comparison between two sequences. BLASTN is used to compare nucleic acid sequences, whereas BLASTP is used to compare amino acid sequences. Other suitable programs are, for example, Needle, Stretcher, Water, or Matcher, which are part of the EMBOSS bioinformatics program suite and available from the European Bioinformatics Institute (EBI) at www.ebi.ac.uk / Tools / psa.
[0102] Sequence alignment can be performed using methods known in the art, such as MAFFT, Clustal (ClustalW, ClustalX, or Clustal Omega), MUSCLE, etc.
[0103] Different regions in a single polynucleotide or polypeptide target sequence that are aligned with a polynucleotide or polypeptide reference sequence can each have their own percentage of sequence identity. Note that the percentage sequence identity value is rounded to one decimal place. For example, 80.11, 80.12, 80.13, and 80.14 are rounded down to 80.1, and 80.15, 80.16, 80.17, 80.18, and 80.19 are rounded up to 80.2. Length values are always integers.
[0104] In some embodiments, the percentage of identity (%ID) between a first amino acid sequence (or nucleic acid sequence) and a second amino acid sequence (or nucleic acid sequence) is calculated as %ID=100×(Y / Z), where Y is the number of amino acid residues (or nucleic acid bases) that are evaluated as perfect matches in an alignment of the first and second sequences (aligned by visual inspection or by a specific sequence alignment program), and Z is the total number of residues in the second sequence. If the length of the first sequence is greater than the length of the second sequence, the percentage of identity of the first sequence to the second sequence will be higher than the percentage of identity of the second sequence to the first sequence.
[0105] Those skilled in the art will understand that the generation of sequence alignments for calculating the percentage of sequence identity is not limited to binary sequence comparisons performed only by primary sequence data. It will also be understood that sequence alignments can be generated by integrating sequence data with data from heterogeneous sources, such as structural data (e.g., protein crystal structures), functional data (e.g., mutation locations), or phylogenetic data. A suitable program for integrating heterogeneous data to generate multiple sequence alignments is T-Coffee, available at www.tcoffee.org or, for example, from EBI. It will also be understood that the final alignments used to calculate the percentage of sequence identity can be curated either automatically or manually.
[0106] The term "variant" or "mutant" refers to a polypeptide that includes an amino acid sequence that differs from a reference polypeptide (e.g., a corresponding unmodified Cas9 protein, e.g., a wild-type Cas9 protein) by one or more amino acids, e.g., by the substitution, deletion, or addition of one or more amino acids. For example, a modified or variant Cas9 polypeptide differs from a wild-type Cas9 (e.g., SEQ ID NO: 1) by the substitution, deletion, and / or addition, i.e., mutation, of one or more amino acids. Unless otherwise specified, such amino acid mutations are also referred to herein as "amino acid modifications."
[0107] As used herein, the terms "isolated," "purified," "extracted," and grammatical variations thereof are used interchangeably to refer to a preparation of a desired composition of the disclosure, e.g., a Cas9 protein modified to exhibit enhanced specificity, that has been subjected to one or more purification processes. In some embodiments, isolation or purification, as used herein, is a process of removing or partially removing (e.g., a fraction of) a composition of the disclosure from a sample containing contaminants.
[0108] In some embodiments, the isolated composition has no detectable undesirable activity, or alternatively, the level or amount of undesirable activity is below an acceptable level or amount. The isolated composition may have an amount and / or concentration of the desired composition of the present disclosure at an acceptable amount and / or concentration and / or activity or higher. In some embodiments, the isolated composition is concentrated compared to the starting material from which the composition is obtained. This concentration may be at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.9%, at least about 99.9%, at least about 99.99%, at least about 99.999%, at least about 99.9999%, or more than 99.9999% compared to the starting material.
[0109] In some embodiments, the isolated preparation is substantially free of residual biological products. In some embodiments, the isolated preparation is 100% free, at least about 99% free, at least about 98% free, at least about 97% free, at least about 96% free, at least about 95% free, at least about 94% free, at least about 93% free, at least about 92% free, at least about 91% free, or at least about 90% free of any contaminating biological material. Residual biological products may include non-biological materials (including chemicals), or unwanted nucleic acids, proteins, lipids, or metabolites.
[0110] As used herein, the term "vector" is intended to refer to a nucleic acid molecule capable of transporting and / or carrying another nucleic acid to which it is linked. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA segments can be ligated. Another type vector is a viral vector, in which additional DNA segments can be ligated into the viral genome. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) can be integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. In addition, certain vectors are capable of directing the expression of genes to which they are operably linked. Such vectors are referred to herein as "recombinant expression vectors" (or simply "expression vectors"). In general, expression vectors useful in recombinant DNA techniques are often in the form of plasmids. Since plasmids are the most commonly used form of vector, "plasmid" and "vector" may be used interchangeably herein. However, other forms of expression vectors, such as viral vectors (eg, replication defective retroviruses, adenoviruses and adeno-associated viruses), are also included, which serve equivalent functions.
[0111] "Cancer" refers to a broad group of various diseases characterized by the uncontrolled growth of abnormal cells in the body. Uncontrolled cell division and growth can lead to the formation of malignant tumors that invade nearby tissues and can metastasize to distant parts of the body via the lymphatic system or bloodstream. Cancers treatable by the present disclosure include those associated with solid tumors.
[0112] A "subject" includes any human or non-human animal. The term "non-human animal" includes, but is not limited to, vertebrates, such as non-human primates, sheep, dogs, and rodents, such as mice, rats, and guinea pigs. In some aspects, the subject is a human. The terms "subject" and "patient" are used interchangeably herein.
[0113] "Treatment" or "treatment" of a subject refers to any type of intervention or process performed on a subject, or the administration of an active agent to a subject, for the purpose of reversing, alleviating, ameliorating, inhibiting, slowing, or preventing the onset, progression, development, severity, or recurrence of a symptom, complication, pathology, or biochemical manifestation associated with a disease.
[0114] As used herein, the terms "ug" and "uM" are used interchangeably with "μg" and "μM," respectively. Various aspects described herein are described in further detail in the following subsections.
[0115] II. Cas9 Protein Variants The present disclosure provides Cas9 proteins having one or more improved properties compared to a reference Cas9 protein (e.g., a wild-type Cas9 protein). For example, as shown herein, the Cas9 proteins described herein include one or more amino acid modifications, such that the Cas9 protein exhibits enhanced (or increased) specificity compared to the reference Cas9 protein. Thus, in some embodiments, provided herein is a Cas9 protein that includes a cavity domain that includes a plurality of amino acids, at least one of which is modified ("amino acid modified") compared to a reference Cas9 protein (e.g., a corresponding wild-type Cas9 protein), where the amino acid modification enhances the specificity of the Cas9 protein. As used herein, the term "cavity domain" refers to a portion of a Cas9 protein that plays a role in the interaction of the Cas9 protein with a nucleic acid sequence.
[0116] The Cas9 protein as present in nature comprises two lobes, a recognition (REC) lobe and a nuclease (NUC) lobe, each of which comprises a specific structural and / or functional domain. The "REC lobe" comprises an arginine-rich bridge helix (BH) domain and at least one REC domain (e.g., a REC1 domain, and optionally a REC2 domain and a REC3 domain), which are involved in the recognition of the guide RNA scaffold and the guide RNA / DNA heteroduplex by the Cas9 protein. For example, without wishing to be bound by any one theory, in some embodiments, the BH domain plays a role in the recognition of the gRNA:DNA, and the REC domain interacts with the repeat:anti-repeat duplex of the gRNA to mediate the formation of the Cas9 / gRNA complex. The "NUC lobe" comprises a RuvC domain, an HNH domain, and a PAM-interacting (PI) domain. The RuvC domain is primarily responsible for cleaving the non-complementary (i.e., bottom or non-target) strand of the target nucleic acid. Meanwhile, the HNH domain is responsible for cleaving the complementary (i.e., top or target) strand of the target nucleic acid. The PI domain contributes to the specificity of the PAM. As used herein, the term "cavity domain" includes both the REC lobe and the NUC lobe. As further described elsewhere in this disclosure, Applicant has determined that modifying one or more amino acids in the cavity domain of the Cas9 protein can increase the specificity of the Cas9 protein.
[0117] In some embodiments, the specificity of the Cas9 protein is at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 45-fold, or at least about 50-fold or more greater than that of a reference Cas9 protein (e.g., a corresponding wild-type Cas9 protein).
[0118] As is evident from the present disclosure, the enhanced specificity allows the Cas9 protein of the present disclosure to more accurately recognize base mismatches within a nucleic acid sequence (e.g., a sequence of a target gene to be modified). Without being bound to any one theory, as a result, in some embodiments, the Cas9 protein described herein may not cleave sequences containing such base mismatches, and thus is associated with reduced off-target effects (e.g., off-target binding, editing, and / or cleavage activity). Similarly, in some embodiments, such Cas9 proteins may have increased on-target effects (e.g., on-target binding, editing, and / or cleavage activity). As further described elsewhere in this disclosure, such base mismatches may be present within the target sequence to which the gRNA binds. In some embodiments, the base mismatches may occur within the PAM. In some embodiments, the base mismatches may occur both within the target sequence and within the PAM.
[0119] In some embodiments, the Cas9 proteins of the present disclosure can accurately distinguish nucleic acid sequences that contain multiple base mismatches (e.g., in the target sequence and / or in the PAM). For example, in some embodiments, the Cas9 proteins described herein can recognize nucleic acid sequences that contain at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 or more base mismatches, thereby not cleaving such sequences. As described herein, when multiple base mismatches are present, in some embodiments, the multiple base mismatches can all be present in the target sequence. In some embodiments, the multiple base mismatches can all be present in the PAM. In some embodiments, some of the multiple base mismatches can be present in the target sequence and some of the multiple base mismatches can be present in the PAM. In some embodiments, as shown herein, the Cas9 proteins described herein can recognize (and therefore not cleave) nucleic acid sequences with a single base mismatch. In some embodiments, the single base mismatch can be present in the target sequence. In some embodiments, a single base mismatch can be present within the PAM. Thus, in some embodiments, the Cas9 protein of the present disclosure can cleave only nucleic acid sequences that are fully complementary (i.e., 100% complementary) to the guide sequence of the gRNA.
[0120] As described herein, in some embodiments, the Cas9 proteins of the disclosure comprise at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, or at least about 20 or more amino acid modifications (e.g., substitutions) within the cavity domain of the Cas9 protein.
[0121] In some embodiments, the Cas9 proteins described herein comprise about one amino acid modification in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about two amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about three amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about four amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about five amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about six amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about seven amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about eight amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about nine amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about ten amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about 11 amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about 12 amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about 13 amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about 14 amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about 15 amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about 16 amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about 17 amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about 18 amino acid modifications in the cavity domain.In some embodiments, the Cas9 proteins described herein comprise about 19 amino acid modifications in the cavity domain. In some embodiments, the Cas9 proteins described herein comprise about 20 amino acid modifications in the cavity domain.
[0122] As shown herein, applicants have identified that modifications (e.g., substitutions) to specific amino acid residues in the cavity domain of the Cas9 protein can enhance the specificity of the Cas9 protein. For example, in some embodiments, the modified amino acid residues can interact with the backbone phosphate of the target DNA strand (i.e., the strand of the DNA molecule to which the Cas9-gRNA complex binds and cleaves) (e.g., in the REC lobe). In some embodiments, the modified amino acid residues can interact with the backbone phosphate of the non-target DNA strand (e.g., in the NUC lobe). When modifying multiple amino acid residues in the cavity domain of the Cas9 protein, in some embodiments, all of the modified amino acid residues are amino acid residues that can interact with the backbone phosphate of the target DNA strand. In some embodiments, all of the modified amino acid residues are amino acid residues that can interact with the backbone phosphate of the non-target DNA strand. In some embodiments, the modified multiple amino acid residues include both amino acid residues that interact with the backbone phosphate of the target DNA strand and amino acid residues that interact with the backbone phosphate of the non-target DNA strand. Without being bound by any one theory, in some embodiments, modification of amino acid residues involved in the interaction of Cas9 protein with target DNA can help improve the specificity of Cas9 protein. In some embodiments, amino acid residues that can be modified include positively charged amino acids such as histidine (H), lysine (K), arginine (R), or combinations thereof. Further details regarding such modifications are described elsewhere in this disclosure.
[0123] Methods for identifying amino acid residues of Cas9 proteins involved in DNA interactions are known in the art. Non-limiting examples of such methods include crystallography, nuclear magnetic resonance (NMR), and sequence aligners. Thus, while the present disclosure focuses primarily on Cas9 proteins from Francisella novicida (FnCas9), it will be apparent to one of skill in the art that the disclosure provided herein applies to Cas9 proteins from other sources as well. For example, in some embodiments, Cas9 proteins that can be modified to enhance specificity using the present disclosure include Cas9 proteins from Streptococcus pyogenes (SpCas9). Non-limiting examples of other suitable Cas9 proteins are known in the art. See, for example, Gasiunas et al., Nat Commun 11(1):5512 (Nov. 2020), which is incorporated herein by reference in its entirety. In some embodiments, the Cas9 proteins described herein include split-Cas9 molecules or inducible Cas9 molecules, e.g., as described in WO2015 / 089427 and WO2014 / 018423, each of which is incorporated by reference in its entirety.
[0124] In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with one or more of the following amino acid residues modified (e.g., substituted) relative to SEQ ID NO:1: K405, R455, K546, K561, K562, K564, K566, K578, K579, R618, R622, K664, R721, R785, K786, K788, K789, R807, K808 , R849, R856, K914, K917, R919, R920, K921, K922, R926, K934, K936, R939, K941, K945, R1047, R1131, R1137, K1142, K1152, K1155, R1178, K1189, K1198, K1206, K1213, K1223, R1226, K1227, K1228, R1241, or any combination thereof.
[0125] In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, wherein one or more of the following amino acid residues are altered (e.g., substituted) relative to SEQ ID NO:1: K405, R455, K566, K578, K664, R721, R785, K786, K789, K914, K917, R919, K921, K922, R926, K934, K936, R939, K941, K945, R1137, K1142, K1152, K1189, K1198, K1206, K1223, R1226, K1227, K1228, R1241, or a combination thereof.
[0126] In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, wherein one or more of the following amino acid residues are altered (e.g., substituted) relative to SEQ ID NO:1: R785, K789, R455, R721, R919, R1241, R939, K1189, K941, R1226, K1228, or a combination thereof.
[0127] In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K405 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R455 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K546 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K561 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K562 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K564 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K566 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K578 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K579 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R618 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R622 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K664 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R721 compared to SEQ ID NO:1.In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R785 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K786 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K788 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K789 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R807 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K808 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R849 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R856 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K914 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K917 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R919 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R920 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K921 compared to SEQ ID NO:1.In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K922 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R926 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K934 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K936 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R939 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K941 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K945 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, and have a modification at amino acid residue R1047 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, and have a modification at amino acid residue R1131 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, and have a modification at amino acid residue R1137 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, and have a modification at amino acid residue K1142 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, and have a modification at amino acid residue K1152 compared to SEQ ID NO:1.In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K1155 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue R1178 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K1189 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K1198 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K1206 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K1213 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, with an alteration at amino acid residue K1223 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO: 1, and have an alteration at amino acid residue R1226 compared to SEQ ID NO: 1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO: 1, and have an alteration at amino acid residue K1227 compared to SEQ ID NO: 1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO: 1, and have an alteration at amino acid residue K1228 compared to SEQ ID NO: 1. In some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO: 1, and have an alteration at amino acid residue R1241 compared to SEQ ID NO: 1.
[0128] As described herein, in some embodiments, the Cas9 proteins of the disclosure comprise multiple amino acid modifications (e.g., substitutions). In some embodiments, the Cas9 proteins described herein comprise two amino acid modifications. For example, in some embodiments, the Cas9 proteins described herein comprise an amino acid sequence set forth in SEQ ID NO:1, with modifications at amino acid residues K1189 and R1241 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise an amino acid sequence set forth in SEQ ID NO:1, with modifications at amino acid residues R721 and R1241 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise an amino acid sequence set forth in SEQ ID NO:1, with modifications at amino acid residues R785 and R1241 compared to SEQ ID NO:1. In some embodiments, the Cas9 proteins described herein comprise three amino acid modifications. For example, in some embodiments, the Cas9 proteins described herein comprise an amino acid sequence set forth in SEQ ID NO:1, with modifications at amino acid residues R785, K1189, and R1241 compared to SEQ ID NO:1. In some embodiments, the Cas9 protein described herein comprises the amino acid sequence set forth in SEQ ID NO: 1, with modifications at amino acid residues R721, K1189, and R1241 compared to SEQ ID NO: 1. In some embodiments, the Cas9 protein described herein comprises the amino acid sequence set forth in SEQ ID NO: 1, with modifications at amino acid residues K1189, K1228, and R1241 compared to SEQ ID NO: 1.In some embodiments, the Cas9 proteins described herein comprise an amino acid sequence as set forth in SEQ ID NO:1, and are altered in at least two distinct amino acid residues compared to SEQ ID NO:1, wherein the at least two distinct amino acid residues are independently: K405, R455, K546, K561, K562, K564, K566, K578, K579, R618, R622, K664, R721, R785, K786, K788, K 789, R807, K808, R849, R856, K914, K917, R919, R920, K921, K922, R926, K934, K936, R939, K941, K945, R1047, R1131, R1137, K1142, K1152, K1155, R1178, K1198, K1206, K1213, K1223, R1226, K1227, K1228, R1241, or K1189.
[0129] As is evident from the present disclosure, the Cas9 proteins described herein can include any suitable amino acid modification, so long as one or more of the amino acid modifications can enhance the specificity of the Cas9 protein. Non-limiting examples of such modifications include substitutions, deletions, insertions, or combinations thereof. As provided herein, in some embodiments, one or more of the amino acid residues described herein are exchanged, i.e., substituted, with a different amino acid. In some embodiments, suitable modifications include conservative substitutions. As used herein, "conservative substitution" (also referred to as conservative replacement) refers to an amino acid replacement that changes a particular amino acid to a different amino acid with similar biochemical properties (e.g., charge, hydrophobicity, and size). Although there are many ways to classify amino acids, amino acids are often classified into six major groups based on the general chemical properties of the structure and R groups. In some embodiments, suitable modifications include radical substitutions. As used herein, the term "radical replacement" refers to an amino acid replacement that replaces an initial amino acid with a final amino acid with different physicochemical properties. [Table 2]
[0130] When the Cas9 protein comprises a plurality of amino acid modifications, in some embodiments, one or more of the plurality of amino acid modifications comprises a conservative substitution. In some embodiments, one or more of the plurality of amino acid modifications comprises a radical substitution. In some embodiments, the plurality of amino acid modifications comprises a conservative substitution and a radical substitution.
[0131] In some embodiments, one or more of the amino acid residues provided herein are replaced with aliphatic amino acids. For example, in some embodiments, one or more positively charged amino acids present in the cavity domain of Cas9 protein are replaced with aliphatic amino acids. In some embodiments, the aliphatic amino acid comprises alanine.
[0132] Thus, in some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO:1, where one or more of the following positively charged amino acid residues in SEQ ID NO:1 are substituted with an aliphatic amino acid: K405, R455, K566, K578, K664, R721, R785, K786, K789, K914, K917, R919, K921, K922, R926, K934, K936, R939, K941, K945, R1137, K1142, K1152, K1189, K1198, K1206, K1223, R1226, K1227, K1228, R1241, or any combination thereof.
[0133] In some embodiments, the Cas9 proteins described herein comprise a K405A substitution (e.g., SEQ ID NO: 251). In some embodiments, the Cas9 proteins described herein comprise a R455A substitution (e.g., SEQ ID NO: 252). In some embodiments, the Cas9 proteins described herein comprise a K566A substitution (e.g., SEQ ID NO: 253). In some embodiments, the Cas9 proteins described herein comprise a K578A substitution (e.g., SEQ ID NO: 254). In some embodiments, the Cas9 proteins described herein comprise a K664A substitution (e.g., SEQ ID NO: 255). In some embodiments, the Cas9 proteins described herein comprise a R721A substitution (e.g., SEQ ID NO: 256). In some embodiments, the Cas9 proteins described herein comprise a R785A substitution (e.g., SEQ ID NO: 257). In some embodiments, the Cas9 proteins described herein comprise a K786A substitution (e.g., SEQ ID NO: 258). In some embodiments, the Cas9 proteins described herein comprise a K789A substitution (e.g., SEQ ID NO: 259). In some embodiments, the Cas9 proteins described herein comprise a K914A substitution (e.g., SEQ ID NO: 260). In some embodiments, the Cas9 proteins described herein comprise a K917A substitution (e.g., SEQ ID NO: 261). In some embodiments, the Cas9 proteins described herein comprise a R919A substitution (e.g., SEQ ID NO: 262). In some embodiments, the Cas9 proteins described herein comprise a K921A substitution (e.g., SEQ ID NO: 263). In some embodiments, the Cas9 proteins described herein comprise a K922A substitution (e.g., SEQ ID NO: 264). In some embodiments, the Cas9 proteins described herein comprise a R926A substitution (e.g., SEQ ID NO: 265). In some embodiments, the Cas9 proteins described herein comprise a K934A substitution (e.g., SEQ ID NO: 266). In some embodiments, the Cas9 proteins described herein comprise a K936A substitution (e.g., SEQ ID NO: 267). In some embodiments, the Cas9 proteins described herein comprise a R939A substitution (e.g., SEQ ID NO: 268).In some embodiments, the Cas9 proteins described herein comprise a K941A substitution (e.g., SEQ ID NO: 269). In some embodiments, the Cas9 proteins described herein comprise a K945A substitution (e.g., SEQ ID NO: 270). In some embodiments, the Cas9 proteins described herein comprise a R1137A substitution (e.g., SEQ ID NO: 271). In some embodiments, the Cas9 proteins described herein comprise a K1142A substitution (e.g., SEQ ID NO: 272). In some embodiments, the Cas9 proteins described herein comprise a K1152A substitution (e.g., SEQ ID NO: 273). In some embodiments, the Cas9 proteins described herein comprise a K1189A substitution (e.g., SEQ ID NO: 2). In some embodiments, the Cas9 proteins described herein comprise a K1198A substitution (e.g., SEQ ID NO: 274). In some embodiments, the Cas9 proteins described herein comprise a K1206A substitution (e.g., SEQ ID NO: 275). In some embodiments, the Cas9 proteins described herein comprise a K1223A substitution (e.g., SEQ ID NO: 276). In some embodiments, the Cas9 proteins described herein comprise a R1226A substitution (e.g., SEQ ID NO: 277). In some embodiments, the Cas9 proteins described herein comprise a K1227A substitution (e.g., SEQ ID NO: 278). In some embodiments, the Cas9 proteins described herein comprise a K1228A substitution (e.g., SEQ ID NO: 279). In some embodiments, the Cas9 proteins described herein comprise a R1241A substitution (e.g., SEQ ID NO: 3).
[0134] In some embodiments, the Cas9 proteins described herein comprise multiple amino acid modifications, the multiple amino acid modifications comprising K1189A and R1241A substitutions. Thus, in some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO: 4. In some embodiments, the Cas9 proteins described herein consist of the amino acid sequence set forth in SEQ ID NO: 4. In some embodiments, the Cas9 proteins described herein consist essentially of the amino acid sequence set forth in SEQ ID NO: 4.
[0135] In some embodiments, the multiple modifications include R721A and R1241A substitutions. Thus, in some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO: 5. In some embodiments, the Cas9 proteins described herein consist of the amino acid sequence set forth in SEQ ID NO: 5. In some embodiments, the Cas9 proteins described herein consist essentially of the amino acid sequence set forth in SEQ ID NO: 5.
[0136] In some embodiments, the multiple modifications include R785A and R1241A substitutions. Thus, in some embodiments, the Cas9 proteins described herein comprise the amino acid sequence set forth in SEQ ID NO: 6. In some embodiments, the Cas9 proteins described herein consist of the amino acid sequence set forth in SEQ ID NO: 6. In some embodiments, the Cas9 proteins described herein consist essentially of the amino acid sequence set forth in SEQ ID NO: 6.
[0137] In some embodiments, the multiple modifications include R785A, K1189A, and R1241A substitutions. Thus, the Cas9 proteins described herein can comprise the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, the Cas9 proteins described herein consist of the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, the Cas9 proteins described herein consist essentially of the amino acid sequence set forth in SEQ ID NO: 7.
[0138] In some embodiments, the multiple modifications include R721A, K1189A, and R1241A substitutions. Thus, the Cas9 proteins described herein can comprise the amino acid sequence set forth in SEQ ID NO: 8. In some embodiments, the Cas9 proteins described herein consist of the amino acid sequence set forth in SEQ ID NO: 8. In some embodiments, the Cas9 proteins described herein consist essentially of the amino acid sequence set forth in SEQ ID NO: 8.
[0139] In some embodiments, the multiple modifications include K1189A, K1228A, and R1241A substitutions. Thus, the Cas9 proteins described herein can comprise the amino acid sequence set forth in SEQ ID NO: 9. In some embodiments, the Cas9 proteins described herein consist of the amino acid sequence set forth in SEQ ID NO: 9. In some embodiments, the Cas9 proteins described herein consist essentially of the amino acid sequence set forth in SEQ ID NO: 9.
[0140] As described herein, in some embodiments, the Cas9 proteins described herein can be fusion proteins. For example, in some embodiments, the Cas9 proteins can be conjugated (e.g., directly or via a linker) or fused to an agent (e.g., a heterologous peptide). Any suitable agent known in the art can be conjugated or fused to the Cas9 proteins described herein to generate a fusion protein. For example, in some embodiments, the Cas9 proteins described herein can be fused to a therapeutic agent and useful for treating a disease or disorder as described herein. In some embodiments, the Cas9 proteins described herein are conjugated to a guide RNA to form a Cas9:guide RNA complex. In some embodiments, the Cas9 proteins described herein are conjugated or fused to an agent that helps improve the activity of the Cas9 protein. For example, in some embodiments, the Cas9 protein is conjugated or fused to a nuclear localization signal and / or a cell-penetrating amino acid sequence, whereby the Cas9 protein can more effectively penetrate a cell (or the nucleus of a cell). In some embodiments, the Cas9 protein can be conjugated or fused to a tag, e.g., an affinity / purification tag or a detectable tag, which can be useful, for example, in generating the Cas9 protein or in determining whether a cell contains the Cas9 protein. Non-limiting examples of such tags include β-galactosidase, glutathione-S-transferase, green fluorescent protein (GFP), epitope tags such as FLAG, myc tags, polyhistidine, nucleases (exo-, endo-) transcription factors, zinc fingers, TALs, deaminases, transposases, methyltransferases, single-stranded DNA binding proteins (SSBs), and inteins.
[0141] III. Polynucleotides, Vectors, and Cells Some aspects of the disclosure relate to polynucleotides (e.g., isolated polynucleotides) encoding any of the Cas9 protein variants (or functional fragments thereof) described herein. The polynucleotides may be present in whole cells, in cell lysates, or in partially purified or substantially pure form. A polynucleotide is "isolated" or "substantially pure" if it has been purified by standard techniques, including alkaline / SDS treatment, CsCl banding, column chromatography, restriction enzymes, agarose gel electrophoresis, and other methods known in the art, to remove other cellular components or other contaminants, such as other cellular nucleic acids (e.g., other chromosomal DNA, e.g., chromosomal DNA linked to the isolated DNA in nature) or proteins. The nucleic acids described herein may be, for example, DNA or RNA, and may or may not include intron sequences. In some aspects, the nucleic acids are cDNA molecules. The nucleic acids described herein may be obtained using standard molecular biology techniques known in the art.
[0142] Exemplary polynucleotides encoding RNA-guided nucleases (e.g., wild-type Cas9 proteins) have been previously described (see, e.g., Cong et al., Science 339(6121):819-23 (Feb. 2013); Wang et al., PLoS One 8(12):e85650 (Dec. 2013), each of which is incorporated by reference in its entirety). As is evident from the present disclosure, polynucleotides useful in the present disclosure differ (e.g., in sequence) from such exemplary polynucleotides because they encode Cas9 proteins that include one or more of the amino acid modifications described herein. Thus, in some embodiments, the isolated polynucleotides provided herein comprise a nucleic acid sequence having less than about 99.999%, less than about 99.998%, less than about 99.997%, less than about 99.996%, less than about 99.995%, less than about 99.994%, less than about 99.993%, less than about 99.992%, less than about 99.991%, less than about 99.99%, less than about 99.8%, less than about 99.7%, less than about 99.6%, less than about 99.5%, less than about 99.4%, less than about 99.3%, less than about 99.2%, less than about 99.1%, less than about 99%, less than about 98%, less than about 97%, less than about 96%, or less than about 95% sequence identity to the nucleic acid set forth in SEQ ID NO:2.
[0143] In some embodiments, the polynucleotides described herein (encoding the Cas9 proteins of the present disclosure) may contain at least one chemically modified nucleobase, sugar, backbone, or any combination thereof. Thus, the polynucleotides encoding the Cas9 proteins of the present disclosure may contain one or more modifications.
[0144] In some aspects, the disclosure provides a vector comprising an isolated polynucleotide encoding a Cas9 protein with enhanced specificity, as described herein.
[0145] Vectors suitable for the present disclosure include expression vectors, viral vectors, and plasmid vectors. In some aspects, the vector is a viral vector.
[0146] Viral vectors include, but are not limited to, retroviruses, such as Moloney murine leukemia virus, Harvey murine sarcoma virus, mouse mammary tumor virus, and Rous sarcoma virus, lentiviruses, adenoviruses, adeno-associated viruses, SV40 type viruses, polyoma viruses, Epstein-Barr viruses, papilloma viruses, herpes viruses, vaccinia viruses, polio viruses, and nucleic acid sequences derived from RNA viruses, such as retroviruses. Other vectors known in the art can also be readily used. Certain viral vectors are based on non-cytopathic eukaryotic viruses in which a non-essential gene has been replaced with a gene of interest. Non-cytopathic viruses include retroviruses, whose life cycle includes reverse transcription of genomic viral RNA into DNA followed by integration of the provirus into the host cell DNA.
[0147] In some embodiments, the vector is derived from an adeno-associated virus (AAV). In some embodiments, the vector is derived from a lentivirus. Examples of lentivirus vectors are disclosed in WO9931251, WO9712622, WO9817815, WO9817816, and WO9818934, each of which is incorporated herein by reference in its entirety.
[0148] Other vectors include plasmid vectors. Plasmid vectors have been widely described in the art and are well known to those skilled in the art. See, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, 1989. In the past few years, plasmid vectors have been found to be particularly advantageous for delivering genes to cells in vivo, since they cannot replicate or integrate into the host genome. However, these plasmids have promoters compatible with the host cell and can express peptides from genes functionally encoded within the plasmid. Some commonly used plasmids available from commercial sources include pBR322, pUC18, pUC19, various pcDNA plasmids, pRC / CMV, various pCMV plasmids, pSV40, and pBlueScript. Further examples of specific plasmids include pcDNA3.1 (catalog number V79020), pcDNA3.1 / hygro (catalog number V87020), pcDNA4 / myc-His (catalog number V86320), and pBudCE4.1 (catalog number V53220), all available from Invitrogen (Carlsbad, Calif.). Other plasmids will be known to those of skill in the art. Additionally, plasmids can be custom designed to remove and / or add specific DNA fragments using standard molecular biology techniques.
[0149] In some embodiments, the present disclosure relates to cells comprising any of the Cas9 proteins, polynucleotides, or vectors described herein. As further described elsewhere in this disclosure, in some embodiments, cells that have been modified (e.g., transduced) to comprise an isolated polynucleotide encoding a Cas9 protein as described herein, or a vector comprising that polynucleotide, can be useful for producing the Cas9 protein as described herein. As further described elsewhere in this disclosure, in some embodiments, cells that have been modified to comprise any of the Cas9 proteins, polynucleotides, or vectors described herein can be useful for treating a disease or disorder (e.g., as part of gene therapy).
[0150] IV. Pharmaceutical Compositions Provided herein are compositions comprising a Cas9 protein (e.g., having enhanced specificity) described herein having a desired purity (or an isolated polynucleotide, vector, or cell associated with such a Cas9 protein) and a pharma- ceutically acceptable carrier or excipient in a form suitable for administration to a subject. In some embodiments, the composition further comprises a guide RNA, which can interact with the Cas9 protein and guide the Cas9 protein to a target sequence.
[0151] Pharmaceutically acceptable excipients or carriers can be determined, in part, by the particular composition to be administered, as well as by the particular method used to administer the composition. Accordingly, there are a wide variety of suitable formulations of pharmaceutical compositions (see, e.g., Remington, 23). rd (See, e.g., The Science and Practice of Pharmacy, ed., A. Adejare, 2020, Academic Press). Pharmaceutical compositions are generally formulated under sterile conditions and in full compliance with all applicable Food and Drug Administration Good Manufacturing Practice (GMP) regulations.
[0152] Acceptable carriers, excipients, or stabilizers are nontoxic to recipients at the dosages and concentrations employed and include buffers, e.g., phosphate, citric acid, and other organic acids, antioxidants including ascorbic acid and methionine, preservatives (e.g., octadecyldimethylbenzylammonium chloride, hexamethonium chloride, benzalkonium chloride, benzethonium chloride, phenol, butyl or benzyl alcohol, alkyl parabens, e.g., methyl or propyl paraben, catechol, resorcinol, cyclohexanol, 3-pentanol, and m-cresol), low molecular weight (less than about 10 residues) polypeptides, proteins, e.g. For example, serum albumin, gelatin, or immunoglobulins; hydrophilic polymers such as polyvinylpyrrolidone; amino acids such as glycine, glutamine, asparagine, histidine, arginine, or lysine; monosaccharides, disaccharides, and other carbohydrates, including glucose, mannose, or dextrins; chelating agents such as EDTA; sugars such as sucrose, mannitol, trehalose, or sorbitol; salt-forming counterions such as sodium; metal complexes (e.g., Zn-protein complexes); and / or non-ionic surfactants such as TWEEN®, PLURONICS®, or polyethylene glycol (PEG).
[0153] In some embodiments, the pharmaceutical compositions disclosed herein comprise one or more additional components selected from a bulking agent, a stabilizer, a surfactant, a buffering agent, or a combination thereof.
[0154] Buffers useful in the present disclosure may be weak acids or weak bases used to maintain the acidity (pH) of a solution near a selected value after the addition of another acid or base. A suitable buffer can maximize the stability of a pharmaceutical composition by maintaining pH control of the composition. A suitable buffer can also ensure physiological compatibility or optimize solubility. Rheology, viscosity, and other properties may also depend on the pH of the composition. Common buffers include Tris buffer, Tris-Cl buffer, histidine buffer, TAE buffer, HEPES buffer, TBE buffer, sodium phosphate buffer, MES buffer, ammonium sulfate buffer, potassium phosphate buffer, potassium thiocyanate buffer, succinate buffer, tartrate buffer, DIPSO buffer, HEPPSO buffer, POPSO buffer, PIPES buffer, PBS buffer, MOPS buffer, acetate buffer, phosphate buffer, cacodylate buffer, glycine buffer, sulfate buffer, imidazole buffer, These include, but are not limited to, guanidine hydrochloride buffer, phosphate-citrate buffer, borate buffer, malonate buffer, 3-picoline buffer, 2-picoline buffer, 4-picoline buffer, 3,5-lutidine buffer, 3,4-lutidine buffer, 2,4-lutidine buffer, Aces, diethyl malonate buffer, N-methylimidazole buffer, 1,2-dimethylimidazole buffer, TAPS buffer, Bis-Tris buffer, L-arginine buffer, lactate buffer, glycolate buffer, or combinations thereof.
[0155] In some aspects, the pharmaceutical compositions disclosed herein further comprise a bulking agent. Bulking agents can be added to pharmaceutical products to add volume and mass to the product, facilitating its accurate measurement and handling. Bulking agents that can be used in the present disclosure include, but are not limited to, sodium chloride (NaCl), mannitol, glycine, alanine, or combinations thereof.
[0156] In some aspects, the pharmaceutical compositions disclosed herein may also include a stabilizer. Non-limiting examples of stabilizers that can be used in the present disclosure include sucrose, trehalose, raffinose, arginine, or combinations thereof.
[0157] In some embodiments, the composition disclosed herein comprises surfactant.Surfactant can be selected from the following: alkyl ethoxylate, nonylphenol ethoxylate, amine ethoxylate, polyethylene oxide, polypropylene oxide, fatty alcohol such as cetyl alcohol or oleyl alcohol, cocamide MEA, cocamide DEA, polysorbate, dodecyl dimethylamine oxide, or combinations thereof.In some embodiments, surfactant is polysorbate 20 or polysorbate 80.
[0158] In some aspects, the pharmaceutical compositions disclosed herein further comprise an amino acid. The amino acid can be selected from arginine, glutamic acid, glycine, histidine, or a combination thereof. In some aspects, the compositions further comprise a sugar alcohol. Examples of sugar alcohols include sorbitol, xylitol, maltitol, mannitol, or a combination thereof.
[0159] The pharmaceutical compositions disclosed herein (e.g., comprising the Cas9 proteins described herein) can be formulated for any route of administration to a subject. Specific examples of routes of administration include intramuscular, cutaneous, subcutaneous, ocular, intravenous, intraperitoneal, intradermal, intraorbital, intracerebral, intracranial, intraspinal, intraventricular, intrathecal, intracapsular, oral, rectal, intravaginal, intratumoral or intratympanic injection. Parenteral administration, characterized by, for example, cutaneous, subcutaneous, intramuscular or intravenous injection, is also contemplated herein.
[0160] Injectables can be prepared in conventional forms, either as liquid solutions or suspensions, solid forms suitable for dissolving or suspending in liquid before injection, or as emulsions. Injectables, solutions, and emulsions also contain one or more excipients. Suitable excipients are, for example, water, saline, dextrose, glycerol, or ethanol. In addition, if desired, the pharmaceutical composition to be administered may also contain small amounts of non-toxic auxiliary substances, such as wetting agents or emulsifying agents, pH buffers, stabilizers, solubility enhancers, and other such agents, such as sodium acetate, sorbitan monolaurate, triethanolamine oleate, and cyclodextrin.
[0161] Pharmaceutically acceptable carriers used in parenteral formulations include aqueous vehicles, non-aqueous vehicles, antimicrobial agents, isotonic agents, buffers, antioxidants, local anesthetics, suspending and dispersing agents, emulsifying agents, sequestering or chelating agents, and other pharma- ceutically acceptable substances. Examples of aqueous vehicles include sodium chloride injection, Ringer's injection, isotonic dextrose injection, sterile water injection, dextrose, and lactated Ringer's injection. Non-aqueous parenteral vehicles include fixed oils of vegetable origin, cottonseed oil, corn oil, sesame oil, and peanut oil. Antimicrobial agents in bacteriostatic or fungistatic concentrations can be added to parenteral preparations packaged in multi-dose containers, including phenol or cresol, mercurials, benzyl alcohol, chlorobutanol, methyl and propyl p-hydroxybenzoic acid esters, thimerosal, benzalkonium chloride, and benzethonium chloride. Isotonic agents include sodium chloride and dextrose. Buffers include phosphate and citrate. Antioxidants include sodium bisulfate. Local anesthetics include procaine hydrochloride. Suspending and dispersing agents include sodium carboxymethylcellulose, hydroxypropyl methylcellulose, and polyvinylpyrrolidone. Emulsifying agents include Polysorbate 80 (TWEEN® 80). Sequestering or chelating agents include EDTA. Pharmaceutical carriers also include ethyl alcohol, polyethylene glycol, and propylene glycol for water-miscible vehicles, and sodium hydroxide, hydrochloric acid, citric acid, or lactic acid for pH adjustment.
[0162] Preparations for parenteral administration include sterile solutions for injection, sterile dry soluble products such as lyophilized powders to be mixed with a solvent immediately before use, including tablets for subcutaneous injection, sterile suspensions for injection, sterile dry insoluble products to be mixed with a vehicle immediately before use, and sterile emulsions. Solutions can be either aqueous or non-aqueous.
[0163] If administered intravenously, suitable carriers include physiological saline or phosphate buffered saline (PBS), as well as solutions containing thickening and solubilizing agents, such as glucose, polyethylene glycol, and polypropylene glycol, and mixtures thereof.
[0164] Topical mixtures containing the antibody are prepared as described for local and systemic administration. The resulting mixture may be a solution, suspension, emulsion, etc., and may be formulated as a cream, gel, ointment, emulsion, solution, elixir, lotion, suspension, tincture, paste, foam, aerosol, douche, spray, suppository, bandage, transdermal patch, or any other formulation suitable for topical administration.
[0165] The therapeutic agents described herein (e.g., Cas9 protein variants with enhanced specificity) can be formulated as aerosols for topical application, such as inhalation (see, e.g., U.S. Pat. Nos. 4,044,126, 4,414,209, and 4,364,923, which describe aerosols for delivering steroids useful in the treatment of inflammatory diseases, particularly asthma). These formulations for administration to the respiratory tract can be in the form of an aerosol or solution for nebulizers, or ultrafine powders for insufflation, alone or in combination with an inert carrier such as lactose. In such cases, the particles of the formulation can have a diameter of less than about 50 microns, e.g., less than about 10 microns.
[0166] The therapeutic agents disclosed herein may be formulated for local or topical application (e.g., for topical application to the skin and mucous membranes, such as in the eye) in the form of gels, creams, and lotions, as well as for application to the eye, or for intracisternal or intrathecal application. Topical administration is contemplated for transdermal delivery, as well as for administration to the eye or mucous membranes, or for inhalation therapy.
[0167] Transdermal patches, such as iontophoretic and electrophoretic devices, are known to those of skill in the art and can be used to administer therapeutic agents, such as those disclosed herein. For example, such patches are disclosed in U.S. Patent Nos. 6,267,983, 6,261,595, 6,256,533, 6,167,301, 6,024,975, 6,010715, 5,985,317, 5,983,134, 5,948,433, and 5,860,957.
[0168] In some aspects, the pharmaceutical composition described herein is a lyophilized powder and can be reconstituted for administration as a solution, emulsion, and other mixture. It can also be reconstituted and formulated as a solid or gel. The lyophilized powder is prepared by dissolving the antibody or antigen-binding portion thereof described herein, or a pharma- ceutically acceptable derivative thereof, in a suitable solvent. In some aspects, the lyophilized powder is sterile. The solvent can contain excipients that improve stability or other pharmacological components of the powder or the reconstituted solution prepared from the powder. Excipients that can be used include, but are not limited to, dextrose, sorbitol, fructose, corn syrup, xylitol, glycerin, glucose, sucrose, or other suitable agents. The solvent can also contain a buffer, such as citric acid, sodium or potassium phosphate, or other such buffers known to those of skill in the art, at approximately neutral pH. Subsequent sterile filtration of the solution followed by lyophilization under standard conditions known to those of skill in the art provides the desired formulation. In some aspects, the resulting solution can be apportioned into vials for lyophilization. Each vial may contain a single dose or multiple doses of the compound. The lyophilized powder may be stored under appropriate conditions, such as at about 4° C. to room temperature.
[0169] Reconstitution of this lyophilized powder with water for injection provides a formulation for parenteral administration.When reconstituting, lyophilized powder is added to sterile water or other suitable carrier.The exact amount depends on the compound selected.Such amount can be empirically determined.
[0170] The pharmaceutical compositions provided herein can also be formulated to target specific tissues, receptors, or other areas of the body of the subject to be treated.Many such targeting methods are known to those skilled in the art.All such targeting methods are contemplated herein for use in the compositions.For non-limiting examples of targeting methods, see, for example, U.S. Patent Nos. 6,316,652, 6,274,552, 6,271,359, 6,253,872, 6,139,865, 6,131,570, 6,120,751, 6,071,495, 6,132,102, 6,133,102, 6,134,102, 6,135,102, 6,136,102, 6,137,102, 6,138,102, 6,139,865, 6,131,570, 6,120,751, 6,071,495, 6,139,865, 6,131,570, 6,120,751, 6,137,102 ... See Nos. 060,082, 6,048,736, 6,039,975, 6,004,534, 5,985,307, 5,972,366, 5,900,252, 5,840,674, 5,759,542, and 5,709,874.
[0171] Pharmaceutical compositions to be used for in vivo administration can be sterilized, which can be accomplished, for example, by filtration through sterile filtration membranes.
[0172] V. Kits / Systems Also disclosed herein are kits comprising any of the Cas9 proteins, polynucleotides, vectors, compositions, or cells described herein. In some embodiments, the kits comprise one or more containers comprising any of the Cas9 proteins, polynucleotides, vectors, compositions, or cells described herein. In some embodiments, the kits further comprise instructions for use according to any of the methods provided herein.
[0173] One of skill in the art will readily appreciate that any of the Cas9 proteins, polynucleotides, vectors, compositions, or cells described herein can be readily incorporated into one of the established kit formats well known in the art. In some embodiments, the kit further comprises additional components such as buffers and interpretive information. In some embodiments, the kit comprises a container and a label or package insert(s) on or associated with the container. In some embodiments, the disclosure provides an article of manufacture that includes the contents of the kit described herein.
[0174] The present disclosure further provides a gene editing system comprising (i) any of the Cas9 proteins described herein, and (ii) a guide polynucleotide. In some embodiments, the guide polynucleotide is a guide RNA comprising a guide sequence complementary to a target sequence of the gene to be modified. As is evident from the present disclosure, the gene editing system of the present disclosure can reduce off-target effects and / or increase on-target effects during the gene editing process, compared to other systems available in the art.
[0175] As described herein, in some embodiments, the likelihood of off-target effects (e.g., cleavage at non-target sequences) is reduced by at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 40-fold, or at least about 50-fold or more, as compared to off-target effects observed in other gene editing systems in the art (e.g., using wild-type Cas9 proteins).
[0176] VI. Methods of the Disclosure VI.A. Method of Construction The present disclosure also encompasses methods of producing / manufacturing the Cas9 protein described herein. In some embodiments, such methods may include expressing the Cas9 protein in a cell that contains a nucleic acid molecule encoding the protein. Host cells that contain these nucleotide sequences are encompassed herein. Non-limiting examples of host cells that can be used include immortal hybridoma cells, NS / 0 myeloma cells, 293 cells, Chinese Hamster Ovary (CHO) cells, HeLa cells, human amniotic fluid-derived cells (CapT cells), COS cells, bacterial cells, insect cells, plant cells, yeast cells, or combinations thereof.
[0177] In relation to the above-mentioned method for producing Cas9 protein, the present disclosure also relates to a method for producing Cas9 protein described herein. Specifically, provided herein is a method for increasing the specificity of Cas9 protein, allowing Cas9 protein to more accurately recognize one or more base mismatches in a gene sequence (e.g., in a target sequence and / or PAM). Applicant has discovered that the specificity of Cas9 protein can be increased by modifying specific amino acid residues in the cavity domain of Cas9 protein. In some embodiments, one or more of the modified amino acid residues can interact with the backbone phosphate of a DNA sequence. Without being bound by any one theory, in some embodiments, such modifications modulate the interaction between Cas9 protein and nucleic acid sequence (e.g., bind less strongly), so that Cas9 protein does not cleave nucleic acid sequence containing one or more base mismatches.
[0178] Non-limiting examples of amino acid residues that can be modified are described elsewhere in this disclosure. Additionally, methods are available in the art for introducing amino acid modifications (e.g., substitutions) into the amino acid sequence of a polypeptide. Nucleic acids encoding variant nucleases can be introduced into viral or non-viral vectors for expression in host cells (e.g., human cells, animal cells, bacterial cells, yeast cells, insect cells). In some embodiments, the nucleic acids encoding variant nucleases are operably linked to one or more regulatory domains for expression of the nuclease. Suitable bacterial and eukaryotic promoters are well known in the art and are described, for example, in Sambrook et al., Molecular Cloning, A Laboratory Manual (3d ed.2001); Kriegler, Gene Transfer and Expression: A Laboratory Manual (1990); and Current Protocols in Molecular Biology (Ausubel et al., eds., 2010). Bacterial expression systems for expressing modified proteins are available in, for example, E. coli, Bacillus species, and Salmonella (Paiva et al., 1983, Gene 22:229-235).
[0179] VI.B. Diagnostic Uses The present disclosure also relates to methods of measuring the expression level of a biomarker (e.g., a nucleotide sequence) associated with a disease or disorder (e.g., cancer and / or a neurodegenerative disease). As will be apparent to one of skill in the art, in some embodiments, measuring the expression level of a biomarker allows for the diagnosis of a disease or disorder in a subject in need of such diagnosis. Thus, in some embodiments, the disclosure provided herein relates to a method of diagnosing a disease or disorder in a subject. In some embodiments, such a method comprises measuring the expression level of a biomarker associated with a disease or disorder in a subject, the method comprising contacting a biological sample obtained from the subject with a Cas9 protein (e.g., modified to exhibit enhanced selectivity) as described herein. As described herein, in some embodiments, the Cas9 protein is contacted with the biological sample in combination with a guide RNA. In some embodiments, the diagnostic methods provided herein further comprise measuring the expression level of a biomarker in the biological sample.
[0180] In some embodiments, an abnormal level (increase or decrease) of a biomarker in a biological sample indicates that the subject suffers from or is at risk of developing a disease or disorder. In some embodiments, the expression level of the biomarker in the biological sample is higher than the corresponding expression level in a reference sample (e.g., a biological sample obtained from a subject determined not to have or at risk of developing a disease or disorder). In some such embodiments, the expression level of the biomarker in the biological sample is at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 45-fold, or at least about 50-fold or more higher than the corresponding expression level in the reference sample. In some embodiments, the expression level of the biomarker in the biological sample is lower than the corresponding expression level in a reference sample (e.g., a biological sample obtained from a subject determined not to have or at risk of developing a disease or disorder). In some such embodiments, the expression level of the biomarker in the biological sample is at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 45-fold, or at least about 50-fold or more lower compared to the corresponding expression level in a reference sample.
[0181] As described herein and as will be appreciated by those skilled in the art, for many diseases and disorders, the amount of biomarkers present in a subject may be extremely low. For example, the amount of circulating tumor DNA (ctDNA) present in the blood of cancer patients may be as low as 0.01% of the total cell-free DNA, making detection of such biomarkers very difficult. See, e.g., Elazezy et al., Comput Struct Biotechnol J 16:370-378 (Oct. 2018); Schwarzenbach et al., Ann NY Acad Sci 1137:190-6 (Aug. 2008); Forshew et al., Sci Transl Med 4(136):136ra68 (May 2012); and Kennedy et al., Nat Protoc 9(11):2586-606 (Nov. 2014), each of which is incorporated herein by reference in its entirety.
[0182] Compared to methods available in the art, the diagnostic methods provided herein allow for a more accurate and cost-effective approach to measure biomarkers, including biomarkers present at very low frequencies in subjects suffering from a disease or disorder. Without being bound to any one theory, in some embodiments, the diagnostic methods provided herein are superior to those available in the art because contacting a biological sample with the Cas9 protein described herein reduces the amount of one or more polynucleotides that differ (e.g., in sequence) from the biomarker, allowing the biomarker to be enriched in the biological sample.
[0183] Thus, provided herein is a method for measuring a first nucleotide sequence (i.e., a biomarker) in a biological sample comprising a first nucleotide sequence and a second nucleotide sequence, the first nucleotide sequence and the second nucleotide sequence being not the same, the method comprising contacting the biological sample with any of the Cas9 proteins described herein, whereby the amount of the second nucleotide sequence present in the biological sample is reduced. In some embodiments, the first nucleotide sequence is a biomarker for a disease or disorder (e.g., comprising a mutation associated with a disease or disorder). In some embodiments, the Cas9 protein is contacted with the biological sample in combination with a guide RNA. Without being bound by any one theory, by reducing the amount of the second nucleotide sequence present in the biological sample, in some embodiments, the presence of the biomarker (i.e., the first nucleotide sequence) can be more accurately measured.
[0184] In some embodiments, the first nucleotide sequence and the second nucleotide sequence differ only with respect to a particular mutation present in the first nucleotide sequence (i.e., a base mismatch compared to the guide sequence of the gRNA). In some embodiments, the mutation present in the first nucleotide sequence comprises a substitution, an insertion, a deletion, a deletion-insertion (indel), a duplication, an inversion, a large genome rearrangement, or a combination thereof. In some embodiments, the mutation comprises a single nucleotide. In some embodiments, the mutation comprises multiple nucleotides (e.g., at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, or at least about 10 or more nucleotides). In some embodiments, the mutation is present in the target sequence to which the Cas9 protein (e.g., Cas9:gRNA complex) binds. In some embodiments, the mutation is present in the PAM. In some embodiments, the mutation is present in both the target sequence and the PAM.
[0185] As is evident from the present disclosure, when a biological sample is contacted with Cas9 protein, Cas9 protein can recognize the mutation present in the first nucleotide sequence. Thus, Cas9 protein does not cleave the first nucleotide sequence, but cleaves the second nucleotide sequence, which comprises a target sequence complementary to the guide sequence of gRNA. As a result, after contact, the amount of the second nucleotide sequence present in the biological sample is reduced by at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or about 100%.
[0186] In some embodiments, it is possible to enrich the first nucleotide sequence (i.e., biomarker) in the biological sample, comprising a mutation, by reducing the amount of the second nucleotide sequence present in the biological sample. Thus, in some embodiments, a method is provided herein for enriching the first nucleotide sequence in a biological sample comprising a first nucleotide sequence and a second nucleotide sequence, the first nucleotide sequence and the second nucleotide sequence being not the same, the method comprising contacting the biological sample with any of the Cas9 proteins described herein (e.g., modified to exhibit enhanced specificity), where after contacting, the biological sample comprises a greater proportion of the first nucleotide sequence. In some embodiments, the Cas9 protein is contacted with the biological sample in combination with a guide RNA. Without being bound by any one theory, by enriching the first nucleotide sequence in the biological sample (i.e., the first nucleotide sequence comprises a greater proportion of the biological sample), in some embodiments, the presence of the biomarker (i.e., the first nucleotide sequence) can be more accurately measured.
[0187] In some embodiments, after contacting, the proportion of the first nucleotide sequence present in the biomarker is increased by at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 40-fold, or at least about 50-fold or more compared to the corresponding proportion in a reference sample (e.g., the biological sample before contacting).
[0188] In some embodiments, the diagnosis is performed ex vivo. For example, the contacting of the biological sample with the Cas9 protein described herein can be performed in vitro. In some embodiments, both the contacting and the measurement of the expression level of the biomarker are performed ex vivo.
[0189] As is apparent from the present disclosure, the expression level of a biomarker (e.g., a first nucleotide sequence comprising a mutation associated with a disease or disorder) can be measured using any suitable method known in the art. For example, the expression level of a biomarker can be measured using any sequencing-based method described herein (see, e.g., Example 1) and / or known in the art (e.g., PCR, real-time PCR, microarray, next generation sequencing (NGS), Sanger sequencing, LAMP, RFLP).
[0190] As described herein, in some embodiments, the diagnostic methods provided herein include contacting a biological sample obtained from a subject with a Cas9 protein of the present disclosure (e.g., in combination with a guide RNA). As used herein, the term "biological sample" refers to any sample containing a substance that can be derived from a subject (e.g., a human). Non-limiting examples of biological samples useful for the present disclosure include tissue, blood, cerebrospinal fluid (CSF), amniotic fluid, semen, vaginal fluid, urine, saliva, sputum, nasal discharge, tears, sweat, feces, keratin, hair, bile, pancreatic juice, gastric juice, serous fluid, transudate, synovial fluid, exudate, abscess, interstitial fluid (ISF), serum, plasma, cell culture medium, or any combination thereof. In some embodiments, the biological sample comprises blood. In some embodiments, the biological sample comprises CSF. In some embodiments, the biological sample comprises serum. In some embodiments, the biological sample comprises plasma. In some embodiments, the biological sample comprises cell culture medium. In some embodiments, the biological sample comprises both blood and CSF, hi some embodiments, the biological sample comprises any combination of blood, CSF, serum, plasma, and culture medium.
[0191] In some embodiments, once a subject has been diagnosed as suffering from a disease or disorder or at high risk of developing a disease or disorder, the subject can be treated with a therapy that, for example, helps to reduce or eliminate one or more symptoms of the disease ("therapeutic treatment") or prevents or delays the onset of the disease ("prophylactic treatment"). Thus, in some embodiments, the diagnostic methods provided herein further include administering a treatment / therapy to a subject identified as having or at risk of developing a disease using the methods provided herein. Further disclosure regarding such treatments is provided elsewhere in this disclosure.
[0192] Moreover, as described herein, the diagnostic methods provided herein may be useful for the diagnosis of a wide range of diseases and conditions. To the extent that a particular disease or disorder is associated with a particular biomarker (e.g., a unique DNA pattern resulting from a particular mutation present in a gene), it will be apparent to one of skill in the art that the Cas9 protein and its associated gRNA can be modified to identify the particular biomarker. Non-limiting examples of such diseases or conditions are described elsewhere in this disclosure. For example, in some embodiments, diseases or conditions to which the present disclosure is applicable include neoplastic diseases (e.g., malignant tumors / benign tumors), hematological diseases (e.g., leukemia / lymphoma), neurodegenerative diseases (e.g., Alzheimer's disease / Parkinson's disease), infectious diseases (e.g., viral infections / bacterial infections), rheumatic diseases (e.g., rheumatoid arthritis / ankylosing spondylitis), neurological diseases (e.g., stroke / amyotrophic lateral sclerosis), allergic diseases (e.g., dermatitis / asthma), psychiatric diseases (e.g., rheumatoid arthritis / ankylosing spondylitis), and the like. schizophrenia / depression), visual disorders (e.g., keratitis / retinitis), endocrine disorders (e.g., diabetes / thyroid dysfunction), congenital disorders (e.g., Down's syndrome / neurofibromatosis), obstetric diagnoses (e.g., prenatal diagnosis / pregnancy diagnosis), cardiovascular disorders (e.g., myocardial infarction / cardiac arrhythmia), pulmonary disorders (e.g., pulmonary embolism / bronchitis), renal disorders (e.g., nephritis / kidney damage), digestive disorders (e.g., gastritis / reflux disease), liver disorders (e.g., hepatitis / cirrhosis), and combinations thereof.
[0193] VI.C. Therapeutic Use As is evident from the present disclosure, the Cas9 proteins described herein (e.g., modified to enhance specificity) may be useful in a wide range of clinical settings in addition to the above diagnostic methods. With the advancement of gene editing technology (e.g., CRISPR-Cas9 system), it is becoming possible to treat various diseases and disorders by gene modification (e.g., gene therapy). For example, by regulating gene expression (e.g., deleting mutated genes and / or introducing healthy genes), cell function can be restored and / or improved, thereby treating a disease or disorder.
[0194] Thus, in some embodiments, provided herein are methods for genetically modifying a cell, comprising contacting the cell with any of the Cas9 proteins provided herein (e.g., modified to enhance specificity), whereupon the expression and / or activity of one or more genes in the cell is modified. In some embodiments, the Cas9 protein is contacted with the cell in combination with a guide RNA, where the guide RNA comprises a guide sequence that is complementary (e.g., fully complementary) to a target sequence in the gene(s) to be modified. Once modified, such cells can be administered to a subject, where the administered cells can provide an ectopic, required function in the subject. In some embodiments, the modified cells are derived from the subject to be treated. In some embodiments, the cells are isolated from the subject prior to contacting, and the modified cells are reintroduced into the subject following contacting. In some embodiments, the cells contacted with the Cas9 protein are derived from a donor (e.g., a healthy donor).
[0195] In some embodiments, the subject to be treated is administered a Cas9 protein as described herein prior to contacting, and the contacting and modification occurs in vivo. Suitable methods of administration are described elsewhere in this disclosure.
[0196] In some embodiments, after modification, the expression and / or activity of the one or more genes is increased by at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 40-fold, or at least about 50-fold or more compared to the expression and / or activity of the one or more genes in a reference cell (e.g., a cell prior to contacting and modification). In some embodiments, after modification, the expression and / or activity of the one or more genes is decreased by at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 40-fold, or at least about 50-fold or more compared to the expression and / or activity of the one or more genes in a reference cell (e.g., a cell prior to contacting and modification).
[0197] In some embodiments, the modified gene differs (e.g., in sequence) from other genes present in the cell ("reference genes"). For example, in some embodiments, the nucleic acid sequence of the modified gene has less than about 99%, less than about 98%, less than about 97%, less than about 96%, less than about 95%, less than about 94%, less than about 93%, less than about 92%, less than about 91%, less than about 90%, less than about 85%, less than about 80%, or less than about 75% sequence identity to the nucleic acid sequence of the reference gene. In some embodiments, both the modified gene and the reference gene comprise a target sequence, a PAM, or both. In some embodiments, the target sequence of the modified gene differs from the target sequence of the reference gene by one or more nucleotides. In some embodiments, the PAM of the modified gene differs from the PAM of the reference gene by one or more nucleotides. In some embodiments, both the target sequence and the PAM of the modified gene differ from those of the reference gene by one or more nucleotides.
[0198] As described herein, the Cas9 proteins of the present disclosure can also recognize single base mismatches in the target sequence and / or PAM, such that the Cas9 proteins will not cleave genes containing such base mismatches. Because of this increased specificity, the Cas9 proteins described herein can increase on-target effects and / or decrease off-target effects.
[0199] Also provided herein is a method for enhancing on-target effect during CRISPR-based gene editing of a cell, comprising contacting the cell with a modified Cas9 protein comprising one or more amino acid modifications that enhance the specificity of the Cas9 protein. In some embodiments, the Cas9 protein comprises any of the modified Cas9 proteins described herein. The Cas9 protein can be contacted with the cell in combination with a guide RNA. In some embodiments, the on-target effect is increased by at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 40-fold, or at least about 50-fold or more compared to the on-target effect observed during CRISPR-based gene editing of a cell using a reference Cas9 protein (e.g., wild-type Cas9 protein).
[0200] Similarly, in some embodiments, provided herein is a method for reducing the occurrence of off-target effects (e.g., cleavage at non-target sequences) during CRISPR-based gene editing of a cell, comprising contacting the cell with a modified Cas9 protein that comprises one or more amino acid modifications that enhance the specificity of the Cas9 protein. In some embodiments, the Cas9 protein comprises any of the modified Cas9 proteins described herein. In some embodiments, the Cas9 protein is contacted with the cell in combination with a guide RNA. In some embodiments, the occurrence of off-target effects is reduced by at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 40-fold, or at least about 50-fold or more compared to the off-target effects observed during CRISPR-based gene editing of a cell using a reference Cas9 protein (e.g., a wild-type Cas9 protein).
[0201] As shown, the Cas9 proteins of the present disclosure observe less than about 70, less than about 65, less than about 60, less than about 55, less than about 50, less than about 45, less than about 40, less than about 35, less than about 30, less than about 25, less than about 20, less than about 15, less than about 10, or less than about 5 off-target site cleavages, as measured, for example, using Digenome-seq analysis. In some embodiments, the Cas9 proteins described herein may produce a single off-target site cleavage. In some embodiments, the Cas9 proteins described herein do not produce off-target site cleavages during CRISPR-based gene editing of a cell.
[0202] As described herein, the present disclosure can be used to treat any suitable disease or disorder known in the art, non-limiting examples of such diseases and disorders are described elsewhere in this disclosure.
[0203] In some embodiments, the therapeutic methods provided herein further include administering one or more additional agents to the subject. For example, if the subject suffers from cancer, the additional agent may include an anti-cancer agent. Non-limiting examples of such anti-cancer agents include chemotherapy, immunotherapy (e.g., checkpoint inhibitors), or both. If the subject suffers from a neurodegenerative disease, the additional therapeutic agent includes an acetylcholinesterase inhibitor. In some embodiments, the additional therapeutic agent includes a dopamine agonist. In some embodiments, the additional therapeutic agent includes a dopamine receptor antagonist. In some embodiments, the additional therapeutic agent includes an antipsychotic. In some embodiments, the additional therapeutic agent includes a monoamine oxidase (MAO) inhibitor. In some embodiments, the additional therapeutic agent includes a catechol O-methyltransferase (COMT) inhibitor. In some embodiments, the additional therapeutic agent includes an N-methyl-D-aspartate (NMDA) receptor antagonist. In some embodiments, the additional therapeutic agent includes an immunomodulatory agent. In some embodiments, the additional therapeutic agent includes an immunosuppressant agent.
[0204] The following examples are illustrative only and should not be construed as limiting the scope of the present disclosure in any way, since numerous variations and equivalents will become apparent to those of skill in the art upon reading this disclosure. EXAMPLES
[0205] Example 1: Materials and Methods In the examples that follow, one or more of the materials and methods described below were used.
[0206] Protein engineering (structural analysis) and cloning The protein structure of FnCas9 (PDB ID 5B2O) was analyzed by Pymol. Also, the residues of FnCas9 that form hydrogen bond distance with DNA were marked with spheres. These residues were changed to alanine using QuikChange II site-directed mutagenesis kit (Agilent). Briefly, wild-type FnCas9 (pET28-a) was used as template, and primers containing alanine point mutations were used to amplify FnCas9 variants. FnCas9 variants with Hisx6 tag at the N-terminus of recombinant FnCas9 protein were cloned by the manufacturer's equipment.
[0207] Protein purification The pET-a vector containing the FnCas9 variant under the T7 promoter was transformed into BL21-DE competent cells (Novagen) by the manufacturer's equipment. The cells carrying the pET-FnCas9 variant were cultured in LB medium (Duchefa, Haarlem, The Netherlands) at 37°C. IPTG (Beams bio) was treated when the OD600nm value reached the range of 0.5-0.7. The cells were harvested after overnight incubation at 18°C. The cells were lysed in lysis buffer (50 mM NaH2PO4, 300 mM NaCl, 10 mM imidazole, 1 mg / ml lysozyme, 1 mM PMSF, 1 mM DTT, pH 8) using an ultrasonic device. The cell lysate was centrifuged at 15000 rpm to remove cell debris. The clear supernatant containing the FnCas9 variant protein was treated with Ni NTA beads (Qiagen). The Ni NTA beads containing FnCas9 protein were washed with washing buffer (50 mM NaH2PO4, 300 mM NaCl, 20 mM imidazole, pH 8). The protein was eluted with elution buffer (50 mM NaH2PO4, 300 mM NaCl, 250 mM imidazole, pH 8) and kept in storage buffer (50 mM HEPES, 200 mM NaCl, 20% glycerol, 1 mM DTT, pH 7.5) until further analysis.
[0208] In vitro transcription of sgRNA We designed and synthesized single guide RNAs (sgRNAs) for SpCas9 and FnCas9 containing single base pair mutations. sgRNAs were synthesized by in vitro transcription. Briefly, RNA templates were incubated with 1000 ng / ml of 40 mM Tris-HCl (pH 7.9), 6 mM MgCl 2 The transcribed sgRNA was purified using a PCR purification kit (GeneAll, Seoul, Korea) and quantified using a NanoDrop spectrophotometer.
[0209] In vitro DNA cleavage assay A 3 kb target DNA containing KRAS, NRAS, and EGFR gene sequences was cleaved with Cas9 protein and sgRNA (Table 3). KRAS, NRAS, and EGFR target sites were synthesized by IDT and cloned into p3 vector. 3 kb target DNA was amplified by PCR from p3 vector containing target sites, two sets of primers, and Q5 DNA polymerase (New England Biolabs). The reaction was cleaned up with PCR cleanup kit (GeneAll). DNA template was incubated with guide RNA, Cas9 variants in Cutsmart buffer (New England Biolabs) (100 mM potassium acetate, 20 mM Tris acetate, 10 mM magnesium acetate, 100 ug / ml BSA, pH 7.9) at 37°C for 1 hour. To analyze nuclease cleavage, digested DNA fragments were run on a TBE 1.5% agarose gel followed by ethidium bromide staining. [Table 3-1] [Table 3-2] [Table 3-3] [Table 3-4]
[0210] Digenome sequencing Digenome was performed as described by Kim et al., Nature methods 12:237-243 (2015). Briefly, 8ug of genomic DNA (gDNA) was extracted from HEK293T using a Blood and Tissue kit (Qiagen) and digested with 40ug of Cas9 and 10ug of gRNA (target sequence: 5'-TTGGACATACTGGATACAGC-3', SEQ ID NO: 280) in 400ul of 1x Cutsmart buffer (New England Biolabs) at 37°C for 16 hours. The digested gDNA was isolated using a Blood and Tissue kit (Qiagen) and then fragmented to a size of 500-600bp by an M220 sonicator (Covaris). NGS libraries for whole genome sequencing were prepared using a TruSeq Nano kit (Illumina) and then sequenced by NovaSeq (illumine). For FnCas9, double-strand break scores were measured using the digenome analysis tool of the Rgenome web server (rgenome.net) with 2 bp overhangs. Loci with DSB scores >1 were sorted and marked as a Manhattan plot of the entire human genome (hg38).
[0211] Analysis of cancer-associated mutations that can be targeted by CRISPR enrichment Methods: We extracted 20-bp sequences above and below all PAM(NGG / CCN) sites from the human genome (GRCh38). We downloaded the COSMIC cancer point mutation data (cancer.sanger.ac.uk / cosmic). We counted the number of mutations located at all PAM(NGG / CCN) sites and the number of mutations located within a 20-bp window above and below all PAM(NGG / CCN) sites. We calculated the ratio of the number of mutations on / within PAM sites to the total number of mutations.
[0212] CRISPR enrichment of mutant DNA wtDNA and EGFR Exon19del, EGFR T790M, EGFR L858R, and KRAS G12D mutants were synthesized by IDT. DNA samples were prepared by mixing wtDNA and mtDNA in the following ratios: 95:5 (95% wtDNA; 5% mtDNA), 99:1 (99% wtDNA; 1% mtDNA), 99.9:0.1 (99.9% wtDNA; 0.1% mtDNA), or 100:0 (100% wtDNA; 0% mtDNA). 5ng of mutant / wtDNA mixture was digested with 500ng FnCas9-AF2 containing 200ng guide in 10ul of 1x Cutsmart buffer (New England Biolabs) for 1 hour at 37°C, and digestion was terminated by adding 10x STOP RXN solution (1% SDS, 100mM EDTA, pH 8). The digested products were amplified with Q5 DNA polymerase (New England Biolabs) using index primers. Index PCR amplicons were purified with AMPure and sequenced on an Illumine Iseq instrument.
[0213] Tissue / blood sample sampling and DNA / cfDNA extraction Patients with stage I non-small cell lung cancer were included in the study, which was approved by (IRB No. 2020AN0005). Cancer tissue was collected during surgical resection and approximately 10 cc of blood was collected in EDTA tubes (BD Vacutainer) before surgical resection. DNA was extracted from the tissue using DNeasy Blood & Tissue kit (Qiagen) according to the manufacturer's protocol. Blood was transferred to Falcon tubes and centrifuged at 1900g. Supernatant (plasma) was collected in e-tubes and centrifuged again at 16000g.
[0214] cfDNA was isolated from 1 ml of plasma using the Maxwell RSC ccfDNA Plasma Kit (Promega) using the manufacturer's equipment. cfDNA was eluted with 60 ul elution buffer from the Maxwell RSC ccfDNA Plasma Kit. cfDNA was run through a cell-free DNA ScreenTape using an Agilent 4150 TapeStation instrument (Agilent). cfDNA concentration and purity were analyzed using the Agilent TapeStation System software (Agilent).
[0215] CRISPR-based enrichment of mutant alleles in cancer patients and preparation of NGS libraries Mutant allele-enriched NGS libraries were prepared from 5–10 ng of gDNA and DNA / cfDNA. Seven genes containing hotspots of interest were amplified with Q5 DNA polymerase (New England Biolabs) using 18 primer pairs (see Table 4).
[0216] Aliquots (1ul) of a 10-fold dilution of the PCR products in DEPC water were treated with 8ug (25pmol / 10ul) FnCas9-AF2 and 2ug (50pmol / 10ul) gRNA mix (see Table 5 showing the gRNA sequences used in different experiments) in 10ul 1x Cutsmart buffer (New England Biolabs) at 37°C for 1 hour to remove wild type DNA alleles, and the reaction was terminated by adding 10x STOP RXN solution (1% SDS, 100mM EDTA, pH 8).
[0217] The wild-type digested product was amplified with Q5 DNA polymerase (New England Biolabs) using index primers. Index PCR amplicons were purified with AMPure and sequenced on an Illumine Iseq instrument.
[0218] 4. Data Analysis To accurately quantify the altered mutation rate in response to Cas9 usage, we analyzed the cfDNA NGS data using an in-house script (written in Python) instead of common NGS analysis methods. The analysis was performed in three steps: QC, targeted read capture, and mutation detection.
[0219] To perform the target read capture and mutation detection steps, target information including target amplicon sites and mutation sites was prepared. The target amplicon sites were extracted from the reference genome sequence, which were sequences expected to be amplified through the gRNA sequence and primer information. The mutation sites were extracted from the COSMIC database, and the expected mutation site sequences were prepared.
[0220] In an initial quality control step, the FASTQC tool ( www.bioinformatics.bbsrc.ac.uk / projects / fastqc ) was used to trim low-quality reads by phred quality score, remove adapter sequences, and perform a length filter (min = 50, max = 150 bp).
[0221] The mutation rate of all variants was expressed as the ratio of the number of reads containing the mutation / total number of reads at a particular variant position. Each ratio was organized in the form of an R data frame and visualized using the heatmap package. Boxplots were presented using the boxplot function in the pubr package in R, and statistical significance of each group was obtained through t-tests for the whole group and Kruskal-Wallis analysis. Correlation plots were presented as a visualization strategy using the plot function and the R cortest function. [Table 4-1] [Table 4-2] [Table 5-1] [Table 5-2] [Table 5-3] [Table 5-4] [Table 5-5] [Table 5-6]
[0222] Example 2: Analysis of the effect of base mismatches on CAS9 endonuclease activity To evaluate the effect of base mismatches between sgRNAs and target DNA sequences on Cas9 activity, sgRNAs (total of 20) with single base mismatches at different positions in the KRAS target sequence (20 base pairs long) were prepared as described in Example 1. Table 5 (above) lists the sequences of sgRNAs targeting KRAS. The specific single base mismatches in the sgRNAs are shown in bold lower case. The activity of the following Cas9 proteins was then evaluated in an in vitro cleavage assay with various KRAS-targeting sgRNAs: (1) wild-type SpCas9, (ii) SpCas9-HF1, (iii) SpCas9-HF4, (iv) eSpCas9(1.0), (v) eSpCas9(1.1), and (vi) FnCas9-wild type.
[0223] As shown in Figures 1A and 1B, except for KRAS-sgRNA#2, wild-type SpCas9 induced significant cleavage of the KRAS target sequence with all other sgRNAs tested. The extent of cleavage was similar to that observed with the control KRAS-sgRNA that did not contain base mismatches. With SpCas9-HF1 and SpCas9-HF4, a significant decrease in target DNA cleavage was observed only with KRAS-sgRNA#2, 7, 13, 14, and 17 (see Figures 2A and 2B). The decrease in cleavage efficiency was more pronounced between SpCas9-HF1 and SpCas9-HF4. In both eSpCas9(1.0) and eSpCas9(1.1), the specificity was similar to that observed with the wild-type SpCas9 protein. With the exception of KRAS-sgRNA#2, significant cleavage of the KRAS target sequence was observed (see Figures 2C and 2D). Finally, similar to wild-type FnCas9 protein, a significant decrease in cleavage was observed for many more sgRNAs, namely KRAS-sgRNA#2, 4, 7, 8, 9, 11, and 17 (see Figures 1C and 1D). A heatmap comparison of the cleavage efficiency of the above Cas9 proteins and various KRAS-sgRNAs is shown in Figure 3.
[0224] The above results indicate that, at least compared to SpCas9 and its highly specific variants, the wild-type FnCas9 protein is much more efficient at discriminating single-base differences within target sequences.
[0225] Example 3: Construction of FNCAS9 protein variants with enhanced specificity To evaluate whether the specificity of the wild-type FnCas9 protein could be further increased, 49 different recombinant FnCas9 proteins were constructed with single amino acid modifications. The amino acid modifications were made at residues within the cavity domain of the wild-type FnCas9 protein (SEQ ID NO: 1) that were predicted to interact with the backbone phosphates of the target DNA sequence. The FnCas9 protein variants contained alanine substitutions at one of the following amino acid residues: K405, R455, K546, K561, K562, K564, K566, K578, K579, R618, R622, K664, R721, R785, K786, K788, K789, R807, K808, R849, R856, K91 The cleavage efficiency of the different FnCas9 protein variants was tested by in vitro cleavage assay using KRAS-sgRNA as described in Example 2.
[0226] As shown in Figure 4A, several of the amino acid modifications (e.g., R455A, R721A, R785A, K789A, R919A, R939A, K941A, K1189A, R1226A, K1228A, and R1241A) significantly reduced the cleavage rate with one or more mismatched sgRNAs while retaining the ability to cleave with perfectly matched sgRNAs. To compare the relative specificity of FnCas9 single residue variants, a specificity score (i.e., the difference between the non-target cleavage rate and the average off-target) was determined for each of the variants. Among the FnCas9 variants tested, variants with alanine substitutions at one of the following amino acid residues had specificity scores greater than 60%: R455, R785, R721, K789, R919, R1241, R939, K941, K1189, R1226, and K1228 (see FIG. 4B). Amino acid residues R455, R785, R721, K789, R919, and R1241 were found to be in the recognition (REC) lobe and interact with the backbone phosphates of the target DNA strand (see FIG. 5). Amino acid residues R939, K941, K1189, R1226, and K1228 were found to be in the nuclease (NUC) lobe and interact with the backbone phosphates of the non-target DNA strand. Among these 11 variants, FnCas9 proteins with alterations at residues K1189 or R1241 had the highest specificity scores.
[0227] Next, we constructed FnCas9 proteins with double and triple amino acid alterations to assess whether specificity could be further improved. Specifically, FnCas9 protein variants were constructed having the following amino acid modifications: (i) K1189A and R1241A ("FnCas9 double mutant #1") (also referred to herein as "FnCas9-advanced fidelity 1" or "FnCas9-AF1"), (ii) R721A and R1241A ("FnCas9 double mutant #2"), (iii) R785A and R1241A ("FnCas9 double mutant #3"), (iv) K1189A and R1241A ("FnCas9 double mutant #3"), (v) R785A, K1189A, R1241A ("FnCas9 triple mutant #1") (also referred to herein as "FnCas9-advanced fidelity 1" or "FnCas9-AF1"). (vi) R721A, K1189A, and R1241A ("FnCas9 triple mutant #2"), and (vii) K1189, K1228, and R1241 ("FnCas9 triple mutant #3"). The in vitro cleavage rates of the FnCas9 double and triple mutants were then evaluated using sgRNAs (60 in total) covering all possible single-base mismatches at all 20 positions within the target NRAS sequence. NRAS-targeting sgRNAs were generated as described in Example 1. Table 5 shows the sgRNA sequences.
[0228] As shown in FIG. 6, compared to the wild-type and single mutant (K1189A or R1241A) FnCas9 proteins, the FnCas9 double mutant #1 (both K1189A and R1241A) showed reduced cleavage of the NRAS target sequence and higher specificity as evidenced by many NRAS-sgRNAs containing a single base mismatch. Similar results were observed with FnCas9 double mutants #2 and #3 (see FIG. 7B). Even higher specificity was observed with the triple mutant (see FIGS. 6 and 7B). To further confirm, the cleavage efficiency of the double and triple FnCas9 protein variants was also tested using the KRAS-sgRNA containing a single base mismatch described in Example 2. As shown in FIG. 7A, similar results were observed, with improved specificity for both the double and triple mutants and the greatest specificity observed with the triple mutant.
[0229] The above results collectively demonstrate the improved specificity of the FnCas9 proteins described herein and suggest that specific amino acid modifications, especially at residues that interact with the backbone phosphates of the DNA sequence, may be useful in improving the specificity of the Cas9 protein.
[0230] Example 4: Analysis of the effect of SGRNA sequence on the specificity of FNCAS9 protein To assess whether the specific sequences of the sgRNAs described herein affect the enhanced specificity of the FnCas9 protein described above, in vitro cleavage assays were performed using FnCas9-AF2 (i.e., FnCas9 triple mutant #1) and sgRNAs containing single-base mismatches to different positions of the target KRAS or EGFR target sequence (20 base pairs long). The sgRNAs were constructed as described in Example 1, and the sequences of the sgRNAs are shown in Table 5.
[0231] As shown in Figures 8B, 8C, and 14A-14D, the specificity of FnCas9-AF2 with KRAS and EGFR sgRNAs was comparable to that previously observed with sgRNAs targeting NRAS (see Example 3). The results presented herein demonstrate that the enhanced specificity of FnCas9-AF2 can be applied to sgRNAs with a variety of sequences.
[0232] Example 5: Off-target analysis of FNCAS9 protein with enhanced specificity To assess whether the increased specificity of the FnCas9 proteins described herein is associated with reduced off-target effects, genome-wide off-target analysis using Digenome-seq was performed for the following Cas9 proteins (as described in Example 1): (i) wild-type SpCas9 protein, (ii) wild-type FnCas9 protein, (iii) eSpCas9(1.1), (iv) FnCas9-AF1, (v) SpCas9-H4, and (vi) FnCas9-AF2.
[0233] As shown in Figures 9A and 9B, the wild-type SpCas9 protein had 654 potential off-target sites with Digenome-seq cleavage scores greater than 1. Also, consistent with the improved specificity observed in the previous examples (see, e.g., Example 1), the number of potential off-target sites was observed to be significantly reduced for the other Cas9 proteins tested. For example, 77 potential off-target sites were observed for the wild-type FnCas9 protein. For eSpCas9(1.1) and SpCas9-H4, there were 37 and 13 potential off-target sites, respectively. Also, for the FnCas9-AF1 and FnCas9-AF2 variants described herein, 1 and 0 potential off-target sites were observed, respectively.
[0234] The above results confirm the improved specificity of the FnCas9 protein variants described herein and suggest that the FnCas9 proteins described herein are less prone to off-target effects observed with many Cas9 proteins available in the art (e.g., wild-type SpCas9 protein).
[0235] Example 6: Enrichment of low-frequency gene mutations using FNCAS9 protein with enhanced selectivity As described herein, current CRISPR-based enrichment methods (e.g., used to identify the presence of circulating tumor DNA in biological samples) require that the mutation be located within the PAM region of the target gene sequence. Lee et al., Oncogene 36:6823-6829 (2017). Thus, such methods cannot identify mutations elsewhere in the target gene sequence. Also, analysis of cancer-associated mutations from the COSMIC database showed that current CRISPR-based enrichment methods may only be applicable to about 30% of all mutations (see FIG. 11). Using the FnCas9 protein variants described herein (e.g., with enhanced specificity and the ability to distinguish single nucleotide differences both within and outside the PAM region), in some embodiments, DNA sequences containing nearly 95% of cancer-associated mutations can be identified.
[0236] To further evaluate the specificity of FnCas9 protein for pathogenic mutations (e.g., mutations associated with cancer), we performed an in vitro enrichment experiment to screen synthetic DNA containing any of the following mutations, all located outside the PAM region: EGFR Exon19del, EGFR T790M, EGFR L858R, and KRAS G12D. Mixtures were prepared containing mutant and wild-type DNA sequences at the following ratios: 5%, 1%, 0.1%, or 0%. The mixtures were then digested with FnCas9-AF2 protein (i.e., FnCas9 triple mutant #1) to determine whether FnCas9 protein variants could be used to enrich for mutant alleles within the different mixtures. Next-generation sequencing was used to measure the frequency of occurrence of mutant alleles within the various mixtures both before and after digestion.
[0237] Before digestion with the FnCas9 protein variants, the frequency of mutant alleles in the different mixtures was roughly as expected (Figures 12A-12C). However, after digestion, the frequency of mutant alleles increased significantly. Notably, no enrichment was observed in the negative control mixture (i.e., containing 0% mutant DNA sequence).
[0238] As further confirmation, the above enrichment steps were applied to cancer patient samples (cancer tissues and blood from 10 non-small cell lung cancer patients (9 stage I and 1 stage II)) and targeted NGS sequencing using general targeted NGS and the above enrichment process, respectively, was performed to confirm the concordance rate. In total, the status of 1,056 genomic variants was analyzed. On average, the concordance rate of mutations between tissue and blood was confirmed to be about 70%. The heat map shown showed that CRISPR enrichment increased the overall correlation of mutational signatures between tissue and ctDNA (see Figure 13). For individual patients, genomic variants were classified into eight conditions (or categories), as shown in Table 6 below. [Table 6]
[0239] Of the eight conditions, category 7 had the most mutations, with 6592, while categories 4, 5, and 6, which correspond to CRISPR-enriched variants, were found with 57, 190, and 198, respectively. CRISPR enrichment increased the detection of pathogenic variants from 445 to 1,480 out of a total of 10,560 possible cases.
[0240] Next, the statistical significance and change of the mutations detected before and after enrichment were analyzed for each group of tissue and cfDNA. Eleven and seventeen variants that met p-value <0.05 and absolute fold change >0.1 were analyzed in tissue DNA and cfDNA, respectively. As shown in Figures 16A-16B, 17A-17D, and 18A-18P, enrichment statistically significantly increased the mutation detection rate.
[0241] In summary, the above results show that the Cas9 protein described herein can efficiently discriminate single base mutations at all 20 target positions and induce DNA cleavage only in targets with perfect base matches. The above results further show that the accuracy of FnCas9-AF is superior to the highly accurate eSpCas9 and SpCas9-HF. The modified FnCas9-AF provided herein can efficiently discriminate base changes at all PAM and non-PAM positions and can be utilized for flexible sgRNA design to detect mutations in circulating tumor DNA (ctDNA) from cancer cells. [Table 7-1] [Table 7-2] [Table 7-3] [Table 7-4] [Table 7-5]
Table 7-6
Table 7-7
Table 7-8
Table 7-9
Table 7-10
Table 7-11
Table 7-12
Table 7-13
Table 7-14
Table 7-15
Table 7-16
Table 7-17
Table 7-18
Table 7-19
Table 8-1
Table 8-2
Table 8-3
Claims
1. A Cas9 protein comprising a cavity domain comprising a plurality of positively charged amino acids, wherein at least one of the plurality of positively charged amino acids is modified (an "amino acid modification") compared to a corresponding wild-type Cas9 protein, and wherein the amino acid modification can increase the specificity of the Cas9 protein.
2. 1, wherein the amino acid modifications are at the following residues of the sequence set forth in SEQ ID NO: 1: K1189, R1241, R785, R721, K1228, K405, R455, K546, K561, K562, K564, K566, K578, K579, R618, R622, K664, K786, K788, K789, R807, K808, R849, R856 , K914, K917, R919, K921, K922, R926, K934, K936, R939, K941, K945, R1047, R1131, R1137, K1142, K1152, K1155, R1178, K1198, K1206, K1213, K1223, R1226, K1227, or a combination thereof.
3. 3. The Cas9 protein of claim 2, wherein the amino acid modification is at residue R785, K1189, R1241, or a combination thereof.
4. 3. The Cas9 protein of claim 2, wherein the amino acid modifications are made at residues: (a) K1189 and R1241 of SEQ ID NO:1; (b) R721 and R1241 of SEQ ID NO:1; (c) R785 and R1241 of SEQ ID NO:1; (d) R785, K1189, and R1241 of SEQ ID NO:1; (e) R721, K1189, and R1241 of SEQ ID NO:1; or (f) K1189, K1228, and R1241 of SEQ ID NO:
1. (a) an amino acid sequence set forth in SEQ ID NO: 7; (b) the amino acid sequence set forth in SEQ ID NO: 2; (c) the amino acid sequence set forth in SEQ ID NO: 3; (d) the amino acid sequence set forth in SEQ ID NO: 4; (e) the amino acid sequence set forth in SEQ ID NO: 5; (f) the amino acid sequence set forth in SEQ ID NO: 6; (g) the amino acid sequence set forth in SEQ ID NO: 8, or (h) the amino acid sequence set forth in SEQ ID NO: 9; 2. The Cas9 protein of claim 1, comprising, consisting of, or consisting essentially of:
6. A composition comprising the Cas9 protein of any one of claims 1 to 5.
7. An isolated polynucleotide encoding the Cas9 protein of any one of claims 1 to 5.
8. A vector comprising the isolated polynucleotide of claim 7.
9. A cell comprising the vector of claim 8.
10. 10. A method for enriching a first nucleotide sequence in a biological sample comprising the first nucleotide sequence and a second nucleotide sequence, the method comprising contacting the biological sample with a Cas9 protein according to any one of claims 1 to 5, wherein the first nucleotide sequence and the second nucleotide sequence are different and the Cas9 protein is capable of cleaving the second nucleotide sequence but is incapable of cleaving the first nucleotide sequence.
11. 10. A method for measuring the amount of a first nucleotide sequence in a biological sample, the biological sample comprising the first nucleotide sequence and a second nucleotide sequence different from the first nucleotide sequence, the method comprising contacting the biological sample with a Cas9 protein of any one of claims 1 to 5, wherein the contacting reduces the amount of the second nucleotide sequence present in the biological sample.
12. 11. The method of claim 10, wherein the first nucleotide sequence comprises a mutation and the second nucleotide sequence does not comprise the mutation.
13. The method described in claim 11, wherein the first nucleotide sequence contains a mutation and the second nucleotide sequence does not contain the mutation.
14. 10. A composition for use in diagnosing a disease in a subject in need thereof, said composition comprising a Cas9 protein according to any one of claims 1 to 5, wherein said Cas9 protein is suitable for detecting whether the amount of a nucleotide sequence is elevated in a biological sample obtained from said subject compared to a corresponding amount present in a reference sample (e.g. a biological sample obtained from a subject not suffering from said disease), said nucleotide sequence comprising a mutation associated with said disease.
15. 10. A method for reducing the occurrence of off-target cleavage of a nucleic acid sequence during CRISPR-based gene editing, said method comprising contacting the nucleic acid sequence with a complex comprising the Cas9 protein and a guide polynucleotide of any one of claims 1-5, wherein said method is not performed in a human subject.
16. A composition for use in a method for reducing the occurrence of off-target cleavage of a nucleic acid sequence during CRISPR-based gene editing, the composition comprising a complex comprising a Cas9 protein and a guide polynucleotide described in any one of claims 1 to 5, the method comprising contacting the nucleic acid sequence with the complex.
17. 1. A method for increasing the specificity of a Cas9 protein, comprising modifying at least one amino acid residue of the Cas9 protein, wherein the at least one amino acid residue is capable of interacting with a backbone phosphate of a DNA sequence, wherein the method is not performed in a human subject.
18. The Cas9 protein comprises the amino acid sequence set forth in SEQ ID NO: 1, and the at least one amino acid residue to be modified is K1189, R1241, R785, R721, K1228, K405, R455, K546, K561, K562, K564, K566, K578, K579, R618, R622, K664, K786, K788, K789, or R80 corresponding to the amino acid sequence set forth in SEQ ID NO:
1. 7, K808, R849, R856, K914, K917, R919, R920, K921, K922, R926, K934, K936, R939, K941, K945, R1047, R1131, R1137, K1142, K1152, K1155, R1178, K1198, K1206, K1213, K1223, R1226, K1227, or a combination thereof.
19. 10. A method for genetic modification of a cell, comprising contacting a cell with a Cas9 protein according to any one of claims 1 to 5, wherein said contacting results in said modification of one or more DNA sequences of said cell, wherein said method is not performed in a human subject.
20. A composition for use in a method for genetically modifying a cell, the composition comprising a Cas9 protein according to any one of claims 1 to 5, the method comprising contacting the cell with the Cas9 protein, the contact resulting in the modification of one or more DNA sequences of the cell.