Type II Cas protein, CRISPR-Cas system and its use
Engineered type II CRISPR-Cas proteins with enhanced PAM recognition and modifications improve genome editing accuracy and versatility, addressing limitations in existing systems for diverse applications.
Patent Information
- Application Number
- JP2026509307
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-06
- Filing Date
- 2024-08-16
- Publication Date
- 2026-08-26
AI Technical Summary
Existing CRISPR-Cas systems face limitations in recognizing diverse protospacer adjacent motifs (PAMs), limiting their versatility and accuracy in genome editing.
Engineered type II CRISPR-related Cas proteins with enhanced PAM recognition capabilities, including specific sequences such as NRRANH, NRHACT, and NGG, and modifications like nuclear localization signals and affinity tags for improved targeting and efficiency.
Enhances the versatility and accuracy of genome editing by allowing broader DNA sequence editing and improved targeting capabilities, facilitating advanced applications in research and biotechnology.
Smart Images

Figure 2026528957000606 
Figure 2026528957000607 
Figure 2026528957000608
Abstract
Description
[Technical Field]
[0001] This disclosure relates to type II Cas proteins, the CRISPR-Cas system, and their uses. In particular, type II Cas proteins and the CRISPR-Cas system are used for gene targeting and gene editing. This application claims priority to PCT applications PCT / CN2023 / 113355, PCT / CN2023 / 116757, PCT / CN2023 / 136724, PCT / CN2024 / 091203, PCT / CN2024 / 091198, and PCT / CN2024 / 091211. The entire contents of the aforementioned applications are incorporated herein by reference. [Background technology]
[0002] Targeted genome editing or modification is rapidly becoming an important tool in both basic and applied research. Clustered, regularly spaced short palindromic repeats and CRISPR-related protein (CRISPR-Cas) systems are the most promising because target specificity can be easily modified by designing the associated guide RNA. Recent advances in genome sequencing technologies and analytical techniques have significantly accelerated the ability to catalog and map genetic factors associated with diverse biological functions and diseases. Precise genome targeting techniques are necessary to enable the systematic reverse engineering of causative genetic mutations by allowing the selective modification of individual genetic elements, and to advance synthetic biology, biotechnology, and medical applications. [Overview of the project] [Problems that the invention aims to solve]
[0003] Various CRISPR-Cas systems have been explored, and different CRISPR-Cas systems exhibit different characteristics. For example, the CRISPR-Cas9 system, which belongs to the class 2 CRISPR-Cas system, is used in genome editing and holds great potential in biomedical research. [Means for solving the problem]
[0004] (Summary of the invention) The present invention provides an engineered, non-natural type II CRISPR-related (Cas) protein or a variant thereof having at least 70% sequence identity to any one amino acid sequence of SEQ ID NOs: 1 to 71. In some embodiments, the Cas protein has at least 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% sequence identity to any one amino acid sequence of SEQ ID NOs: 1 to 71.
[0005] The present invention also provides an engineered, non-natural type II CRISPR-related (Cas) protein having at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% sequence identity with respect to any one of the amino acid sequences of SEQ ID NOs. 1 to 71, except for the amino acid "M" at position 1 of the sequence.
[0006] In some embodiments, the Cas protein further comprises an effector domain (or functional domain). Such an effector domain may have one or more types of enzymatic activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain comprises a nuclease, nickase, deaminase, reverse transcriptase, recombinase, methyltransferase, methylase, acetylase, acetyltransferase, transcription activator, transcription repressor domain, cryptochrome, photoinducible / controllable domain, or chemoinducible / controllable domain.
[0007] In some embodiments, the Cas protein further includes one or more nuclear localization signal sequences, nuclear export signal sequences, cell membrane permeable peptide sequences, and affinity tags. A type II Cas protein includes one or more nuclear localization signals (NLS). The NLS can be located at the end or other site of the peptide. The NLS located at each end or other site of the Cas9 amino acid sequence may be identical or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are identical. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains one NLS, and the C-terminus of the Cas9 amino acid sequence contains one NLS. The amino acid sequences of the NLS are fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence, respectively. The NLS may be an SV40 (simian virus 40) NLS, a c-Myc NLS, or another suitable monosegmented NLS. The NLS may be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, affinity tags (such as GST, FLAG, or hexahistidine sequences) are used for the purification of Cas proteins by affinity chromatography. In some embodiments, the amino acid sequence of the C-terminal NLS is described in SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is described in SEQ ID NO: 883. Other available sequences and different combinations of NLS and FLAG sequences are also selectable.
[0008] In some embodiments, the Cas protein contains an amino acid sequence having 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to any one of the amino acid sequences of SEQ ID NOs: 1 to 71.
[0009] In some preferred embodiments, the Cas protein has sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 44%, 51%, 56%, 59, 60, or 100% with respect to any one amino acid sequence of SEQ ID NOs. 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, or 68. In several other preferred embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 44%, 51%, 56%, 59, 60, or 68 sequence identity with respect to any one amino acid sequence of sequence numbers 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, or 68, except for the amino acid "M" at position 1 of the sequence.
[0010] One of the key elements of the modification process associated with the CRISPR-Cas system is the protospacer adjacent motif (PAM), which is a short DNA sequence adjacent to the target DNA sequence. PAM sequences are essential for the binding and cleavage activity of Cas proteins, ensuring that only the intended DNA sequence is edited. The ability of Cas proteins to recognize these specific PAM sequences is due to their unique structural features and binding interactions with DNA. Each PAM sequence provides a distinct motif that Cas proteins can recognize, ensuring accurate and efficient targeting. For example, the NRRANH sequence may provide a specific nucleotide sequence to which Cas proteins can bind with high affinity. These diverse PAM sequences expand the potential application range of the CRISPR-Cas system. This allows researchers and scientists to broaden the range of DNA sequences they can edit, increasing the versatility and effectiveness of this powerful gene editing tool. Furthermore, understanding these PAM sequences contributes to the development of Cas proteins with more advanced targeting capabilities, further advancing the field of gene editing and its potential benefits in research, medicine, and biotechnology. The Cas proteins disclosed in this invention exhibit a unique ability to recognize diverse PAM sequences. These sequences include NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. Specific recognition of these Cas proteins provides greater flexibility in selecting the DNA sequences to be edited.
[0011] In some embodiments, the Cas proteins disclosed herein are capable of recognizing at least one protospacer-adjacent motif (PAM) having or containing the sequence NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one amino acid sequence of SEQ ID NO: 9, and is capable of recognizing protospacer adjacent motifs (PAMs) having the sequence NRRANH; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 12, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the sequence NRHACT. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one amino acid sequence of SEQ ID NO: 19, and is capable of recognizing protospacer adjacent motifs (PAMs) having the sequence NRAAR; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 21, and is capable of recognizing protospacer adjacent motifs (PAMs) having the sequence NNNCCY; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one amino acid sequence of SEQ ID NO: 22, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NNRYYYY; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one amino acid sequence of SEQ ID NO: 24, and is capable of recognizing protospacer adjacent motifs (PAMs) having sequence NGG; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 25, and is capable of recognizing protospacer adjacent motifs (PAMs) having the sequence NNNNCAA; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 27, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the sequence NRNACN; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one amino acid sequence of SEQ ID NO: 29, and is capable of recognizing protospacer adjacent motifs (PAMs) having sequence NNGR; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one amino acid sequence of SEQ ID NO: 30, and is capable of recognizing protospacer adjacent motifs (PAMs) having sequence NGGNR; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 31, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NNNCCH; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 36 and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRRAAG; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 37 and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRHRAC; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 38 and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRYART; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 43 and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRHACC; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 44 and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRAAR; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 51 and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRNVHH; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 56 and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence YMACAW; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 59 and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NAHAA; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one amino acid sequence of SEQ ID NO: 60, and is capable of recognizing protospacer adjacent motifs (PAMs) having the sequence NRHAYYY; In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of the amino acid sequences of SEQ ID NO: 68, and is capable of recognizing protospacer adjacent motifs (PAMs) having the sequence NGGHA.
[0012] In some embodiments, the Cas protein is a nickase or dead Cas protein. The DNA cleavage domain of the active Cas protein in this invention comprises two subdomains: an HNH nuclease subdomain and a RuvC subdomain. Mutations within these subdomains can silence the nuclease activity of the Cas protein. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 9, and includes a mutation at residue D11 or H859; or having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12, and containing a mutation at residue D10 or H862; Alternatively, having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 31, and containing a mutation at residue D12 or H903. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 9, and includes mutations at residue D11 or H859, except for amino acid "M" at position 1 of the sequence; Or having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12, and containing a mutation at residue D10 or H862, except for the amino acid "M" at position 1 of the sequence; Alternatively, it has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 31, and includes a mutation in residue D12 or H903, except for the amino acid "M" at position 1 of the sequence. In some embodiments, the mutation in residue D11 or H859 of SEQ ID NO: 9 is D11A or H859A; Mutations in residues D10 or H862 of sequence number 12 are D10A or H862A; Alternatively, a mutation in residue D12 or H903 of SEQ ID NO: 31 is D12A or H903A. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one amino acid sequence of SEQ ID NOs: 869–877.
[0013] The present invention also provides engineered, non-natural polynucleotides encoding type II CRISPR-related (Cas) proteins disclosed herein.
[0014] In some embodiments, the polynucleotide encoding the Cas protein is operably ligated to a promoter and present within the vector; optionally, the vector is selected from the group consisting of retroviral vectors, lentiviral vectors, phage vectors, adenovirus vectors, adeno-associated virus vectors, herpes simplex virus vectors, and plasmid vectors.
[0015] In some embodiments, the polynucleotide is a ribonucleotide sequence or a deoxyribonucleotide sequence, or an analog thereof; optionally, the polynucleotide is codon-optimized for expression in the target cell; in some embodiments, the polynucleotide is codon-optimized for expression in eukaryotic cells. In some embodiments, the eukaryotic cell is selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cell is a mammalian cell, preferably a human cell.
[0016] In some embodiments, the polynucleotide is mRNA and further comprises a 5' cap sequence and / or a poly-A tail sequence. In some embodiments of the present invention, the mRNA used may be modified to enhance its functional properties and stability. Specifically, in some embodiments, the modification process involves substituting uridine (represented by the letter "U") with N1-methylpseudouridine or pseudouridine. This substitution is designed to improve the mRNA's resistance to ribonuclease degradation and potentially increase its intracellular half-life and translation efficiency. Incorporating N1-methylpseudouridine or pseudouridine into the mRNA structure may also have a positive impact on the immune response profile, as these modifications have been shown to reduce the immunogenicity of the mRNA molecule compared to unmodified mRNA molecules. This is particularly important in the development of mRNA-based therapeutics and vaccines where minimizing adverse immune responses is paramount.
[0017] In some embodiments, the polynucleotides of the present invention are codon-optimized for expression in eukaryotic cells; optionally, the eukaryotic cells are selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and / or human cells.
[0018] In some embodiments, the polynucleotide has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to any one nucleotide sequence among sequence numbers 161-231, 241-311, 851-852, and 861-863.
[0019] The present invention also provides an engineered non-natural type CRISPR-Cas system comprising: a) a type II Cas protein as described herein, or a polynucleotide encoding the Cas protein; b) at least one engineered guide RNA, or at least one engineered nucleic acid encoding the guide RNA, wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein-binding segment that interacts with the Cas protein, and the Cas protein-binding segment comprises a tracrRNA sequence that hybridizes to form a double-stranded RNA (dsRNA) duplex and a direct repeat (DR) sequence.
[0020] In some embodiments, the guide RNA is a dual guide RNA. In some embodiments, the guide RNA is a single guide RNA. In such embodiments, the guide RNA further includes a linker sequence that ligates the tracrRNA sequence and the DR sequence. In some typical embodiments, the linker includes a short sequence GAAA. In some embodiments, the linker functions as an artificial loop. In some embodiments, the sgRNA includes a) a spacer sequence that can hybridize with the sequence of the target nucleic acid to be manipulated; b) a DR sequence; c) a linker sequence; and d) a tracrRNA sequence. The serial arrangement of the spacer sequence, DR sequence, linker sequence, and tracrRNA sequence is in the 5' to 3' direction or the 3' to 5' direction. In some embodiments, the sgRNA scaffold contains sequences having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% homology to any one of sequence numbers 431-562.
[0021] In some embodiments, the sgRNA comprises a spacer sequence (such as one of sequence numbers 571-835, 885, or 901) and a scaffold sequence, the spacer sequence being located at the 5' end of the scaffold sequence (such as sequence number 903). In some embodiments, the sgRNA comprises a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% homology to any one of sequence numbers 903.
[0022] In some embodiments, the spacer sequence hybridizes with one or more nucleic acids in a prokaryotic or eukaryotic cell. In some embodiments, the eukaryotic cell is selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the eukaryotic cell includes mammalian cells. In some embodiments, the mammalian cell includes human cells. In some embodiments, the eukaryotic cell includes plant cells.
[0023] In some embodiments, the system further includes a donor template nucleic acid.
[0024] In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid. In some embodiments, the donor template nucleic acid is a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear. In some embodiments, the donor template nucleic acid is circular (e.g., a plasmid). In some embodiments, the donor template nucleic acid is an exogenous nucleic acid molecule. In some embodiments, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, the donor template nucleic acid is DNA, RNA, or a DNA-RNA hybrid.
[0025] The present invention also provides a modified vector comprising the polynucleotide described herein.
[0026] In some embodiments, the vector is an expression vector. In some embodiments, the vector is an inductive, conditional, or constitutive expression vector. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA are located on the same vector or on different vectors.
[0027] The present invention also provides a vector system comprising one or more polynucleotides encoding one or more polynucleotides and a guide RNA as described herein, wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein-binding segment that interacts with the Cas protein, and the Cas protein-binding segment comprises a tracrRNA sequence that hybridizes to form a double-stranded RNA (dsRNA) duplex and a direct repeat (DR) sequence. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA reside on the same vector or on different vectors.
[0028] In some embodiments, the vector (e.g., plasmid or viral vector) is delivered to the target tissue by means of, for example, intramuscular injection, intravenous administration, transdermal administration, nasal administration, oral administration, or mucosal administration. Such delivery may be performed as a single dose or multiple doses. Those skilled in the art will understand that the actual dose administered herein may vary considerably depending on various factors, including the choice of vector, target cells, organisms, tissues, the general condition of the subject being treated, the degree of transformation or modification desired, the route of administration, the method of administration, and the type of transformation or modification desired.
[0029] The present invention also provides engineered non-natural cells comprising the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, or the vector system described herein.
[0030] The present invention also provides cells modified using the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, or the vector system described herein.
[0031] In some embodiments, the cells are eukaryotic or prokaryotic. In some embodiments, the eukaryotic cells are selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cells are mammalian cells, human cells, or plant cells.
[0032] In some embodiments, the cells are cells of vertebrates, mammals, rodents, goats, pigs, birds, chickens, turkeys, cattle, horses, sheep, fish, primates, or humans. In some embodiments, the cells are mammalian cells. In one embodiment, the cells are human cells. In some embodiments, the cells are somatic cells, germ cells, or embryonic cells. In some embodiments, the cells are zygote cells, blastocyst cells, embryonic cells, stem cells, mitotic cells, or meiotic cells. In some embodiments, the cells are not part of a human embryo. In some embodiments, the cells are somatic cells. In one embodiment, the cells include T cells, CD8+ T cells, CD8+ naive T cells, central memory T cells, effector memory T cells, CD4+ T cells, stem cell-like memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricting progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythroid progenitor cells, megakaryocyte-erythroid progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, trabecular meshwork cells, cochlear hair cells, outer hair cells, inner hair cells, and lung epithelial cells. These include bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, neurons, neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells (e.g., progenitor B cells, pre-B cells, pro-B cells, memory B cells, plasma B cells), gastrointestinal epithelial cells, bile duct epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, pre-adipocytes, islet cells (e.g., β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0033] In some embodiments, the disclosure includes a modified target locus, where the target locus is modified according to the method, use of a composition, or use of a system of the present invention.
[0034] In some embodiments, the cells are eukaryotic or prokaryotic. In some embodiments, the eukaryotic cells are selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cells are mammalian cells, human cells, or plant cells.
[0035] In some embodiments, the cells are cells of vertebrates, mammals, rodents, goats, pigs, birds, chickens, turkeys, cattle, horses, sheep, fish, primates, or humans. In some embodiments, the cells are mammalian cells. In one embodiment, the cells are human cells. In some embodiments, the cells are somatic cells, germ cells, or embryonic cells. In some embodiments, the cells are zygote cells, blastocyst cells, embryonic cells, stem cells, mitotic cells, or meiotic cells. In some embodiments, the cells are not part of a human embryo. In some embodiments, the cells are somatic cells. In one embodiment, the cells include T cells, CD8+ T cells, CD8+ naive T cells, central memory T cells, effector memory T cells, CD4+ T cells, stem cell-like memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricting progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythroid progenitor cells, megakaryocyte-erythroid progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, trabecular meshwork cells, cochlear hair cells, outer hair cells, inner hair cells, and lung epithelial cells. These include bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, neurons, neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells (e.g., progenitor B cells, pre-B cells, pro-B cells, memory B cells, plasma B cells), gastrointestinal epithelial cells, bile duct epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, pre-adipocytes, islet cells (e.g., β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0036] The present invention also provides a kit comprising the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, the vector system described herein, or the cells described herein.
[0037] The kits described herein contain in one or more containers the components essential for carrying out the methods described herein and may include instructions for use as appropriate. Any of the kits may additionally include auxiliary components necessary for carrying out the preparation method. Each component in the kit is provided, where applicable, in liquid form (e.g., dissolved in solution) or solid form (e.g., lyophilized powder). In certain embodiments, some components can be reconstituted or otherwise processed (e.g., to an active state) by adding a suitable solvent or other substance (e.g., water or buffer), and such solvent or substance may or may not be included in the kit. In some embodiments, the kit may further include other suitable excipients, such as buffers and reagents, to facilitate the application of the kit. Preferably, the kits can be applied to a variety of uses, such as medical applications including therapeutic and diagnostic applications, and research applications. Thus, the type II Cas nucleases and kits of the present invention can be used for the preparation of pharmaceuticals for therapeutic purposes and / or for the preparation of reagents for research purposes.
[0038] The Cas proteins, CRISPR-Cas systems, and polynucleotides described herein can be delivered by various delivery systems, such as vectors (e.g., plasmids), viral delivery vectors such as adeno-associated viruses (AAV), lentiviruses, and other viral vectors, or by nuclear translocation or electroporation of ribonucleoprotein complexes consisting of a type VI effector and its corresponding RNA guide(s). The proteins and one or more RNA guides can be packaged into one or more vectors (e.g., plasmids or viral vectors). In bacterial applications, nucleic acids encoding any component of the CRISPR systems described herein can be delivered to bacteria using phages. Representative phages include, but are not limited to, T4 phage, Mu, λ phage, T5 phage, T7 phage, T3 phage, Φ29, M13, MS2, Qβ, and ΦX174.
[0039] The present invention also provides a pharmaceutical composition comprising any of the Cas proteins, polynucleotides, CRISPR-Cas systems, vectors, vector systems, or cells described herein.
[0040] As used herein, the term “pharmaceutical composition” refers to a formulation intended for pharmaceutical use. In certain embodiments, the pharmaceutical composition further comprises pharmaceutically acceptable excipients. In some embodiments, the pharmaceutical composition may contain additional therapeutic agents. In some embodiments, the pharmaceutical composition is prepared according to standard procedures for administration to a subject, such as a human patient, via intravenous, intramuscular, intradermal, intraarticular, intralesional, intraperitoneal, intracardiac, intrathecal, intraventricular, epidural, topical, subconjunctival, intrastromal, periorbital, intravitreous, parascleral, perscleral, choroidal, retroorbital, subretinal, subtenon’s capsule, nasal inhalation, pressurized inhalation, oral, subcutaneous, or topical routes. For example, a composition for injection may be provided as a sterile isotonic aqueous solution. If necessary, the pharmaceutical composition may also contain solubilizers such as lidocaine and local anesthetics to minimize discomfort at the injection site. Generally, the components are supplied alone or as mixed unit doses (e.g., lyophilized powder or anhydrous concentrated solution) in sealed containers indicating the amount of the active ingredient. When a pharmaceutical composition is intended for intravenous administration, it can be combined with an infusion bottle containing sterile pharmaceutical-grade water or saline solution. When a pharmaceutical composition is intended for injection, it may include sterile water or saline solution for injection to allow for pre-administration mixing of components. Furthermore, humectants, colorants, release agents, coating agents, sweeteners, fragrances, aromas, preservatives, and antioxidants may be incorporated into the formulation as needed.
[0041] In some embodiments, the pharmaceutical composition further comprises a delivery system selected from AAV (adeno-associated virus), adenovirus, retrovirus, HSV (herpes simplex virus), gamma retrovirus, LV (lentivirus), eCIS (extracellular contractile injection system), eVLPs (modified virus-like particles), VLPs (virus-like particles), liposomes, plasmids, LNPs (lipid nanoparticles), exosomes, microvesicles, nucleic acid nanoassemblies, gene guns, and / or implantable devices.
[0042] In certain embodiments, delivery is carried out via adeno-associated viruses (AAVs), such as AAV2, AAV8, or AAV9, which are at least 1 × 10⁻¹⁶ 5 It can be administered as a single dose containing adenovirus or adeno-associated virus particles (also called particle units or pu).
[0043] In some embodiments, delivery is carried out via recombinant adeno-associated virus (rAAV) vectors. For example, in some embodiments, modified AAV vectors can be used for delivery. Modified AAV vectors may be based on one or more capsid types, including AAV1, AAV2, AAV5, AAV6, AAV8, AAV8.2, AAV9, AAV rh10, modified AAV vectors (e.g., modified AAV2, modified AAV3, modified AAV6), and pseudotype AAVs (e.g., AAV2 / 8, AAV2 / 5, AAV2 / 6).
[0044] In some embodiments, delivery is via plasmids. The dose can be a sufficient number of plasmids to induce a response. In some embodiments, the appropriate amount of plasmid DNA in a plasmid composition is in the range of about 0.1 mg to about 2 mg. A plasmid generally includes (i) a promoter; (ii) a sequence encoding a nucleic acid-targeted CRISPR enzyme operably ligated to the promoter; (iii) a selection marker; (iv) an origin of replication; and (v) a transcription termination sequence located downstream of (ii) and operably ligated. Plasmids can also encode RNA components of the CRISPR-Cas system, although one or more of these may instead be encoded on different vectors. The frequency of administration is at the discretion of medical or veterinary professionals (e.g., physicians, veterinarians) or those skilled in the art.
[0045] The present invention also provides for the use of the Cas proteins, polynucleotides, CRISPR-Cas systems, vectors, vector systems, cells, kits, or pharmaceutical compositions described herein for the treatment, prevention, diagnosis, or detection of diseases.
[0046] The present invention also provides a method for modifying or targeting a target DNA locus, comprising delivering to the locus a Cas protein as described herein; a polynucleotide as described herein; a CRISPR-Cas system as described herein; a vector as described herein; a vector system as described herein; a kit as described herein; or a pharmaceutical composition as described herein.
[0047] In some embodiments, the Disclosure also provides methods for targeting and cleaving target DNA, comprising contacting the target DNA with a Cas protein as described herein; a polynucleotide as described herein; a CRISPR-Cas system as described herein; a vector as described herein; a vector system as described herein; a kit as described herein; or a pharmaceutical composition as described herein.
[0048] In some embodiments, the modification or targeting of the target locus includes inducing DNA strand breaks. In some embodiments, the modification or targeting of the target locus includes inducing DNA double-strand breaks or DNA single-strand breaks. In some embodiments, the modification or targeting of the target locus includes altering the gene expression of one or more genes. In some embodiments, the modification or targeting of the target locus includes epigenetic modification of the target DNA locus. In some embodiments, the method is a method for modifying cells, cell lines, or organisms by manipulating one or more target sequences at a genomic locus of interest.
[0049] In some embodiments, cleavage of target DNA or target sequence results in indel formation or insertion of a nucleotide sequence. In some embodiments, cleavage of target DNA or target nucleotide involves cleaving the target DNA or target sequence at two locations, resulting in a deletion or inversion of the sequence between the two locations. In some embodiments, the target DNA is double-stranded DNA, single-stranded DNA, or a DNA-RNA hybrid.
[0050] In some embodiments, modification or targeting of the target gene locus includes induction of DNA strand breaks, alteration of gene expression of one or more genes, or epigenetic modification of the target DNA gene locus; optionally, the DNA strand breaks include DNA double-strand breaks or DNA single-strand breaks.
[0051] In some embodiments, the method is performed in vitro or in vivo.
[0052] The present invention also provides isolated eukaryotic cells comprising a modified target locus of interest, wherein the target locus of interest is modified by the method described in this disclosure, or by using the system described in this disclosure, or by using the Cas protein described in this disclosure; or by using the polynucleotide described in this disclosure; or by using the CRISPR-Cas system described in this disclosure; or by using the vector described in this disclosure, or by using the vector system described in this disclosure, or by using the kit described in this disclosure, or by using the pharmaceutical composition described in this disclosure.
[0053] The present invention also provides a system for detecting the presence of a nucleic acid target sequence in an in vitro sample, comprising: a) a Cas protein as described herein; b) at least one guide polynucleotide comprising a guide sequence capable of binding to a target sequence and designed to form a complex with the Cas protein; and c) a nucleic acid-based masking construct comprising a non-target sequence, wherein the Cas protein exhibits contingent cleavage activity of RNA and / or ssDNA, cleaving the non-target sequence of the nucleic acid-based masking construct activated by the target sequence.
[0054] The present invention also provides a method for detecting target nucleic acids in a sample, comprising the steps of contacting one or more samples with a nucleic acid-based masking construct comprising a) a Cas protein as described herein; b) at least one guide polynucleotide having a degree of complementarity with a target sequence and having a guide sequence designed to form a complex with the Cas protein; and c) a non-target sequence, wherein the Cas protein exhibits contingent cleavage activity of RNA and / or ssDNA, cleaving the non-target sequence of the nucleic acid-based masking construct activated by the target sequence, and detecting a signal from the cleavage of the non-target sequence, thereby detecting one or more target sequences in the sample.
[0055] The present invention also provides a guide RNA (gRNA) comprising: a) a spacer sequence of Sequence ID No. 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of Sequence ID No. 901; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of Sequence ID No. 901. In some embodiments, the gRNA further comprises a Cas protein-binding segment; and the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) duplex. In some embodiments, the gRNA is a dual guide RNA. In some embodiments, the gRNA is a single guide RNA. In some embodiments, the gRNA is modified. In some embodiments, at least three nucleotides of the gRNA are modified. In some embodiments, the gRNA includes a 5'-terminus modification containing at least two phosphorothioate (PS) bonds within the first seven nucleotides from the 5' end of the 5' terminus. In some embodiments, the gRNA includes a 3'-terminus modification containing at least two phosphorothioate (PS) bonds within the first seven nucleotides from the 3' end of the 3' terminus. In some embodiments, the gRNA has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one nucleotide sequence of SEQ ID NO: 903.
[0056] The present invention also provides a polynucleotide encoding a guide RNA (gRNA) comprising: a) a spacer sequence of sequence number 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of sequence number 901; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 901. In some embodiments, the gRNA contains sequences that have 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any one of sequence numbers 903.
[0057] The present invention also provides an engineered non-natural CRISPR-Cas system comprising: a) a Cas protein, or a polynucleotide encoding the Cas protein; b) at least one guide RNA (gRNA) described herein, or at least one engineered nucleic acid encoding the guide RNA, wherein the gRNA further comprises a Cas protein-binding segment that interacts with the Cas protein; and wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) duplex; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of Sequence ID No. 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of Sequence ID No. 901; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with respect to the sequence of Sequence ID No. 901. In some embodiments, the gRNA comprises a) the sequence of sequence number 903; or b) a sequence having 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the sequence of sequence number 903.
[0058] In some embodiments, the Cas protein comprises a sequence having 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to any of SEQ ID NOs: 12; or C The as protein contains a sequence that has 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to any of sequence number 12, except for the amino acid "M" at position 1 of the sequence.
[0059] In some embodiments, the Cas protein further comprises an effector domain (or functional domain). Such an effector domain may have one or more types of enzymatic activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain comprises a nuclease, nickase, deaminase, reverse transcriptase, recombinase, methyltransferase, methylase, acetylase, acetyltransferase, transcription activator, transcription repressor domain, cryptochrome, photoinducible / controllable domain, or chemoinducible / controllable domain.
[0060] In some embodiments, the Cas protein further includes one or more nuclear localization signal sequences, nuclear export signal sequences, cell membrane permeable peptide sequences, and affinity tags. A type II Cas protein includes one or more nuclear localization signals (NLS). The NLS can be located at the end or other site of the peptide. The NLS located at each end or other site of the Cas9 amino acid sequence may be identical or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are identical. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains one NLS, and the C-terminus of the Cas9 amino acid sequence contains one NLS. The amino acid sequences of the NLS are fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence, respectively. The NLS may be an SV40 (simian virus 40) NLS, a c-Myc NLS, or another suitable monosegmented NLS. The NLS may be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, affinity tags (such as GST, FLAG, or hexahistidine sequences) are used for the purification of Cas proteins by affinity chromatography. In some embodiments, the amino acid sequence of the C-terminal NLS is described in SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is described in SEQ ID NO: 883. Other available sequences and different combinations of NLS and FLAG sequences are also selectable.
[0061] In some embodiments, the Cas protein contains an amino acid sequence having 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12.
[0062] In some embodiments, the Cas protein is capable of recognizing a protospacer adjacent motif (PAM) having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 12, and having the sequence of NRHACT.
[0063] In some embodiments, the Cas protein is a nickase or dead Cas protein. The DNA cleavage domain of the active Cas protein in the present invention comprises two subdomains: an HNH nuclease subdomain and a RuvC subdomain. Mutations within these subdomains can silence the nuclease activity of the Cas protein. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12, and includes a mutation at residue D10 or H862; several embodiments In this case, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence of Sequence ID No. 12, except for the amino acid "M" at position 1 of the sequence, and contains a mutation at residue D10 or H862.
[0064] In some embodiments, the mutation in residue D10 or H862 of SEQ ID NO: 12 is D10A or H862A; in some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one amino acid sequence of SEQ ID NOs: 872-874.
[0065] The present invention also provides a modified vector comprising a polynucleotide encoding the gRNA described herein, comprising: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 901. In some embodiments, the gRNA comprises a) the sequence of sequence number 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to sequence number 903.
[0066] In some embodiments, the vector is an inducible, conditional, or constitutive expression vector.
[0067] The present invention also provides a vector system comprising one or more polynucleotides encoding the gRNA described herein, comprising: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 871. In some embodiments, the gRNA comprises a) the sequence of sequence number 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to sequence number 903.
[0068] The present invention also includes the gRNA described herein, a polynucleotide encoding such gRNA, a CRISPR-Cas system comprising such gRNA or polynucleotide, a vector comprising such gRNA coding sequence, and a vector system comprising such gRNA coding sequence; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; c) at least 99%, 98%, of the sequence of SEQ ID NO: 901. The present invention provides a pharmaceutical composition comprising a spacer sequence having 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity; in some embodiments, the gRNA comprises a) the sequence of SEQ ID NO: 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO: 903.
[0069] The present invention also provides a method for treating, preventing, or diagnosing a disease associated with the RHO locus in a subject, wherein the guide RNA (gRNA) comprises: a) the spacer sequence of Sequence ID No. 901; b) the spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of Sequence ID No. 901; c) the spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of Sequence ID No. 901. In some embodiments, the gRNA comprises a) the sequence of sequence number 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to sequence number 903.
[0070] The present invention also provides a method for treating, preventing, or diagnosing a locus-related disease in a subject, wherein the guide RNA (gRNA) comprises i) the spacer sequence of Sequence ID No. 901; ii) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of Sequence ID No. 901; or iii) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity to the sequence of Sequence ID No. 901. In some embodiments, the gRNA comprises the sequence of sequence number 903; or a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to sequence number 903.
[0071] The present invention also provides compositions comprising: (i) Cas protein; where, a. The Cas protein contains a sequence that has at least 90% identity with SEQ ID NO: 12 or 92; and / or b.Cas proteins contain sequences that are at least 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 12 or 92; and / or (ii) an sgRNA containing the sgRNA sequence of Sequence ID No. 903, or a vector encoding said sgRNA.
[0072] The present invention also provides a method for modifying an RHO gene locus, comprising delivering a composition to cells, wherein the composition is a. Guide RNA containing the guide sequence of SEQ ID NO: 901; b. A guide RNA containing at least 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; or c. A guide RNA containing a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 901;
[0073] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases in a subject, comprising administering a composition to a subject of interest, wherein the composition is a. Guide RNA containing the guide sequence of SEQ ID NO: 901; b. A guide RNA containing at least 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; or c. A guide RNA containing a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 901;
[0074] The present invention also provides a method for modifying an RHO gene locus, comprising delivering a composition to cells, wherein the composition is a. sgRNA containing the sgRNA sequence of sequence number 903; b. An sgRNA containing an sgRNA sequence that has at least 90% identity with the sequence of Sequence ID No. 903; or c. An sgRNA containing an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 903;
[0075] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases in a subject, comprising administering a composition to a subject of interest, wherein the composition is a. sgRNA containing the sgRNA sequence of sequence number 903; b. An sgRNA containing an sgRNA sequence that has at least 90% identity with the sequence of Sequence ID No. 903; or c. An sgRNA containing an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 903;
[0076] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases in a subject, comprising administering a composition to a subject of interest, wherein the composition is a. Guide RNA containing the spacer sequence of SEQ ID NO: 901; b. A guide RNA containing at least 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; or c. A guide RNA containing a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 901;
[0077] The present invention also provides a method for modifying an RHO gene locus, comprising the step of delivering a composition to cells, wherein the composition is (i) Cas protein, here a. The Cas protein contains a sequence that has at least 90% identity with SEQ ID NO: 12 or 92; and / or b. The RNA-guide DNA binder contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% homology to SEQ ID NO: 12 or 92; and / or (ii) A guide RNA or vector encoding a guide RNA containing the spacer sequence of SEQ ID NO: 901;
[0078] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases in a subject, comprising administering a composition to a subject of interest, wherein the composition is (i) RNA guide DNA binding agent, however, a. The RNA-guide DNA binder contains a sequence having at least 90% homology to SEQ ID NO: 12 or 92; and / or b. The RNA-guide DNA binder contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% homology to SEQ ID NO: 12 or 92; and / or (ii) an sgRNA containing the sequence of sequence number 903 or a vector encoding said sgRNA.
[0079] These and other aspects, purposes, features, and advantages of the embodiments will become apparent to those skilled in the art by considering the following detailed description of the illustrated embodiments. [Brief explanation of the drawing]
[0080] An understanding of the features and benefits of this disclosure will be obtained by referring to the following detailed description and accompanying drawings that specify exemplary embodiments in which the principles of the disclosure may be utilized: [Figure 1] This shows the domain configuration of the GEBx type II Cas protein. [Figure 2A] 2A-2E exhibits PAM preference for Cas9 in the HEK293 cell line. [Figure 2B] 2A-2E exhibits PAM preference for Cas9 in the HEK293 cell line. [Figure 2C] 2A-2E exhibits PAM preference for Cas9 in the HEK293 cell line. [Figure 2D] 2A-2E exhibits PAM preference for Cas9 in the HEK293 cell line. [Figure 2E] 2A-2E exhibits PAM preference for Cas9 in the HEK293 cell line. [Figure 3] This shows the indel levels of GEBx0305 against 16 targets containing GGAAAA-PAM in the HEK293T cell line (n=3). [Figure 4] This shows the indel levels of GEBx0308 against 19 targets containing GGTACT-PAM in the HEK293T cell line (n=3). [Figure 5]This shows the indel levels of GEBx0308 against 15 targets containing CATACT-PAM in the HEK293T cell line (n=3). [Figure 6] This shows the indel levels of GEBx0328 against 20 targets containing NGGCCT-PAM in the HEK293T cell line (n=3). [Figure 7] This shows the indel levels of GEBx0361 against 11 targets containing GGTACC-PAM in the HEK293T cell line (n=2). [Figure 8] This shows the indel levels of GEBx0361 against eight targets containing TGTACC-PAM in the HEK293T cell line (n=2). [Figure 9] A and B show the change in indel level depending on the length of the guide array of GEBx0305 (n=3). [Figure 10] A and B show the change in indel level depending on the length of the guide array of GEBx0308 (n=3). [Figure 11] A and B show the change in indel level depending on the length of the guide array of GEBx0328 (n=3). [Figure 12] A and B show the indel levels of GEBx0305 targeting endogenous genes using a modified RNA scaffold (n=3). [Figure 13] A and B show the indel levels of GEBx0308 targeting the endogenous gene using a modified RNA scaffold (n=3). [Figure 14] A and B show the indel levels of GEBx0305 targeting endogenous genes under optimal conditions (n=3). [Figure 15] A and B show the indel levels of GEBx0308 targeting endogenous genes under optimal conditions (n=3). [Figure 16] A and B show the indel levels of GEBx0328 targeting endogenous genes using a modified RNA scaffold (n=3). [Figure 17]A and B show the indel levels of GEBx0328 targeting endogenous genes under optimal conditions (n=3). [Figure 18] This shows the indel levels of GEBx0305, which targets the endogenous gene, after transfection of HEK293T cells with a lipoplex containing a fixed amount (20 ng) of sgRNA and mRNA in different ratios. SpCas9 is used as a positive control. [Figure 19] This shows the indel levels of GEBx0308, which targets endogenous genes, after transfection of HEK293T cells with a lipoplex containing a fixed amount (20 ng) of sgRNA and mRNA in different ratios. [Figure 20] This shows the indel levels of GEBx0305, which targets the endogenous gene, after transfection of PHH cells with lipoplexes containing a fixed amount (20 ng) of sgRNA and mRNA in different ratios. SpCas9 is used as a positive control. [Figure 21] This shows the indel levels of GEBx0308, which targets endogenous genes, after transfection of PHH cells with a lipoplex containing a fixed amount (20 ng) of sgRNA and mRNA in different ratios. [Figure 22] This shows an overview of the higher-level Guide-seq insertion sites for target site 1 (CFTR-NGGAAAA-T5) and site 2 (EMX1-NGGAAAA-T5) of GEBx0305. [Figure 23] This shows an overview of the top guide-seq insertion sites for target site 1 (CD34-NGGTACT-T4) and site 2 (POLQ-NGGTACT-T1) of GEBx0308. [Figure 24] This shows an overview of the top guide-seq insertion sites of GEBx0328 for the target sequence (CFTR-NGGCCT-T3) and site 2 (CFTR-NGGCCT-T5). [Figure 25] This shows the frequency of A-to-G base editing in HEK293T cells using GEBx0305-ABE for adenine at four sites. [Figure 26]This shows the frequency of A-to-G conversion editing at five adenine base sites in HEK293T cells using GEBx0308-ABE. [Figure 27] This shows the frequency of A-to-G conversion editing at five adenine base sites in HEK293T cells by GEBx0328-ABE. [Figure 28] This shows 4-allyle specific editing of GEBx0308 at the RHO-P23H pathogenic site. [Modes for carrying out the invention]
[0081] (Detailed description of preferred embodiments) The following embodiments further illustrate the present disclosure, but are not limited thereto.
[0082] It should be noted that, as used herein and in the claims, the singular form or the terms “one,” “one,” “this,” “the foregoing,” and similar terms as used in the context of this disclosure (particularly in the context of the claims) shall be construed to encompass both singular and plural unless otherwise specifically indicated herein or the context clearly contradicts this. In some embodiments, the foregoing terms can be reasonably understood as “one” or “one or more.” Furthermore, unless the context specifically requires otherwise, the singular form of a term shall include the plural form, and the plural form of a term shall include the singular form.
[0083] Unless otherwise specified, singular / plural terms are intended to encompass their active and past tenses and should be understood according to the context.
[0084] In this disclosure, particularly in the claims and / or paragraphs, terms such as “comprises,” “omprised,” and “comprising” may have the meanings given in U.S. patent law. For example, they may have the meanings of “includes,” “included,” and “including,” and terms such as “consisting essentially of” and “consists essentially of” have the meanings given in U.S. patent law. Terms such as “a group consisting of” refer to a particular set or collection of elements, components, or features. It may include one or more of the specified elements, components, or features. For example, a group consisting of A, B, or C may refer to a set that includes one or more of the specified elements A, B, or C. The claims encompass the possibility of having any one single element (A, B, or C) individually, any combination of any two elements (A and B, A and C, or B and C), or any combination of all three elements (A, B, and C). This wording defines the variability of the invention within the specified options while maintaining the scope of the claims, even while allowing different combinations of the enumerated elements.
[0085] In this invention, when "t" or "T" appears in the RNA sequence as a nucleotide, it should be understood as "u" or "U".
[0086] The term "identity" of two or more nucleic acid or polypeptide sequences refers to two or more sequences or subsequences that are identical or have the same proportion of amino acid residues or nucleotides, as measured by a sequence comparison algorithm such as BLAST, BLAST 2.0, or FASTA using the default parameters described below.
[0087] The term “exemplary” is used herein to mean an example, case, or illustration. An embodiment or design described as “exemplary” should not necessarily be construed as preferable or advantageous to any other embodiment or design.
[0088] In this specification, the terms “any” or “optionally” include both cases in which the events, situations, or substituents described below occur and do not occur, and such descriptions include both cases in which such events or situations occur and do not occur.
[0089] Unless otherwise specified, the use of "or" or " / " is inclusive and means "and / or". Furthermore, different interpretations are possible depending on the context. As used herein, the term "and / or" is intended to include both A and B, A or B, A (alone), and B (alone) in phrases such as "A and / or B". Similarly, as used herein, the term "and / or" is intended to include each of the following embodiments in phrases such as "A, B, and / or C": A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0090] In this specification, the terms “approximately” and “~” used to refer to measurable values such as parameters, quantities, and temporal durations mean that they include variation from a specified value. The value itself that the modifier “approximately” or “~” refers to should also be understood to be specifically and preferably disclosed.
[0091] The term “exemplary” is used herein to mean an example, case, or illustration. An embodiment or design described as “exemplary” should not necessarily be construed as preferable or advantageous to any other embodiment or design.
[0092] The terms “subject,” “individual,” and “patient” are used interchangeably herein to refer to vertebrates, preferably mammals, more preferably humans. Mammals include, but are not limited to, mice, primates, humans, livestock, sports animals, and pets. Tissues, cells, and their offspring obtained in vivo or cultured in vitro from biological entities are also included.
[0093] A “variant” of a sequence disclosed herein includes a sequence having one or more additions, deletions, stop positions, or substitutions compared to the sequence disclosed herein.
[0094] "Code" refers to the property of a specific sequence of nucleotides in a gene, such as cDNA or mRNA, to function as a template for the synthesis of other macromolecules, such as a defined amino acid sequence. Therefore, if the transcription and translation of mRNA corresponding to a gene produce a protein in a cell or other biological system, that gene codes for a protein. Protein-coding polynucleotides are degenerate versions of each other and include all nucleotide sequences that code for the same amino acid sequence or a substantially similar amino acid sequence in terms of form and function.
[0095] The terms “unnatural” and “modified” are used interchangeably to indicate artificial involvement. When referring to nucleic acid molecules or polypeptides, these terms mean that the nucleic acid molecule or polypeptide is substantially free from at least one other naturally occurring component in nature. In all aspects and embodiments, with or without these terms, they are preferably optional and therefore preferably included or preferably not included. Furthermore, the terms “unnatural” and “modified” are interchangeable and can be used alone or in combination, with either being replaced by both. In particular, “modified” is preferably used instead of “unnatural,” “unnatural and / or modified,” or “artificially engineered unnatural.”
[0096] As presented herein, the term “cleavage event” as used herein refers to a DNA break in a target nucleic acid caused by a type IICas nuclease of the CRISPR system described herein. In some embodiments, the cleavage event is a double-strand DNA break. In some embodiments, the cleavage event is a single-strand DNA break.
[0097] As presented in the present invention, the term "targeting" refers to the ability of a complex comprising a CRISPR-related protein and an RNA guide to preferentially or specifically bind (e.g., hybridize) to a particular target nucleic acid compared to other nucleic acids that do not have the same or similar sequence as the target nucleic acid.
[0098] As presented in this disclosure, the term "GEBx" followed by a numerical suffix is used as a general code to represent either a nucleic acid or a protein. It is important to note that the use of the same code for nucleic acids and proteins, or their derivatives, does not mean that the substances represented by these codes are identical. In other words, GEBx-0305 may refer to a specific nucleic acid sequence in some cases and to a different protein in other cases. In some embodiments, a direct correspondence may be shown between nucleic acids and proteins represented by the same or derived codes. Thus, the code "GEBx" functions as an index system for organizing and referencing the diverse biomolecules described in this disclosure, and the meaning of the code is understood based on the context provided.
[0099] Unless otherwise defined herein, scientific and technical terms used in connection with this disclosure shall have the meanings generally understood by an ordinary technician of the art. The meaning and scope of terms should be clear; however, in the event of any potential ambiguity, the definitions provided herein shall prevail over dictionary or external definitions.
[0100] Various embodiments are described below. It should be noted that specific embodiments are not intended as exhaustive descriptions or as limitations on broader aspects discussed herein. Aspects described in relation to a particular embodiment are not necessarily limited to that embodiment and can be implemented in any one other embodiment. Throughout this specification, phrases such as “in some embodiments,” “in certain embodiments,” “in some preferred embodiments,” “in some typical embodiments,” “in typical embodiments,” or similar expressions mean that certain features, structures, or characteristics described in relation to that embodiment are included in at least one embodiment of this disclosure. Furthermore, certain features, structures, or characteristics can be combined in any suitable way in one or more embodiments, as will be apparent to those skilled in the art from this disclosure. Furthermore, some embodiments described herein include some features included in other embodiments but not others, but combinations of features from different embodiments are within the scope of the disclosure. For example, any of the claimed embodiments in the appended claims can be used in any combination.
[0101] Numerical ranges indicated by endpoints include all numbers and fractions contained within each range, as well as the endpoints indicated.
[0102] Various embodiments are described below. Note that specific embodiments are not intended as exhaustive descriptions or as limitations on broader aspects discussed herein. An aspect described in relation to a particular embodiment is not necessarily limited to that embodiment and can be implemented in any one other embodiment. Throughout this specification, references to “in a particular embodiment,” “in some embodiments,” or “in a certain embodiment” mean that a particular feature, structure, or characteristic described in relation to that embodiment is included in at least one embodiment of this disclosure. Therefore, phrases such as “in a particular embodiment,” “in one embodiment,” or “in a certain embodiment” appearing in various places throughout this specification do not necessarily refer to the same embodiment, but may refer to it. Furthermore, specific features, structures, or characteristics can be combined in any suitable way in one or more embodiments, as will be apparent to those skilled in the art from this disclosure. Additionally, some embodiments described herein include some features included in other embodiments but not others, but combinations of features from different embodiments are within the scope of the disclosure. For example, any of the claimed embodiments in the appended claims can be used in any combination.
[0103] All publications, published patent documents, and patent applications cited herein are incorporated by reference to the same extent as any individual publication, published patent document, or patent application is incorporated by reference.
[0104] The present invention provides an engineered, non-natural type II CRISPR-related (Cas) protein or a variant thereof having at least 70% sequence identity to any one amino acid sequence of SEQ ID NOs: 1 to 71. In some embodiments, the Cas protein has at least 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% sequence identity to any one amino acid sequence of SEQ ID NOs: 1 to 71.
[0105] As presented herein, "M" as used herein represents the starting amino acid methionine, which is the starting point for protein synthesis in many proteins. Naturally occurring Cas proteins often begin with methionine as the first amino acid in their sequence. However, when scientists modify these proteins, for example, by fusing them with nuclear localization signals (NLS) or other domains, this initial amino acid "M" may be substituted or altered to introduce new functions or properties.
[0106] This disclosure focuses particularly on the flexibility and functionality of modified Cas proteins. Except for the intentionally modified start amino acid, methionine (M) at position 1, the designed Cas proteins exhibit high sequence identity ranging from 70% to 100% compared to the reference sequence. This strategic modification not only aligns with our goal of tuning the protein's properties but also ensures that the core or expected functionality inherent in the Cas protein is maintained. By manipulating the initial amino acids, we enhance specific properties such as improved intracellular localization and the introduction of other advantageous features, without compromising overall sequence homology, while maintaining the fundamental properties that make the Cas protein an indispensable tool in genomic manipulation.
[0107] The present invention also provides an engineered, non-natural type II CRISPR-related (Cas) protein having at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% sequence identity with respect to any one of the amino acid sequences of SEQ ID NOs. 1 to 71, except for the amino acid "M" at position 1 of the sequence.
[0108] As presented in this disclosure, “Cas protein,” “CRISPR-related protein,” or other similar terms refer to a class of CRISPR-related proteins that are essential components of the CRISPR-Cas system. These proteins possess endogenous nuclease activity and can cleave double-stranded DNA or RNA molecules in a sequence-specific manner guided by complementary RNA molecules; for example, Cas9 and Cas12 are widely used in genome editing applications. Furthermore, some Cas proteins can be modified to retain only one of the two active sites necessary for double-strand breaks, resulting in nickase activity and enabling controlled cleavage events by introducing single-strand breaks into target nucleic acid sequences. In addition, certain Cas proteins can be modified to completely lack endogenous nuclease activity; these are called dead Cas or nuclease-inactive Cas proteins. Despite lacking enzymatic function, these dead Cas proteins retain the ability to specifically bind to target nucleic acid sequences and are often used in combination with other effector domains (or functional domains) for applications such as gene regulation, epigenome editing, and components of advanced imaging systems. In some embodiments, Cas proteins may be used to reduce off-target effects. In some embodiments, an active Cas nuclease, nickase, or dead Cas may be part of a fusion protein containing another effector domain. Fusion proteins containing such other effector domains (or functional domains) and such active Cas nuclease, nickase, or dead Cas are also included in the range of Cas proteins. In some embodiments, Cas proteins may be fragmented. In some embodiments, Cas proteins may be inducible Cas proteins. In some embodiments, type II Cas proteins may be part of a self-inactivation system (SIN); in some embodiments, type II Cas nucleases may be part of a synergistic activation system (SAM) as defined elsewhere herein.
[0109] In some embodiments, the domain configuration of type II Cas proteins includes a RuvC domain, a BH (bridge helix) domain, an REC domain, an HNH domain, and / or a CTD (C-terminal domain). The RuvC domain is a critical catalytic site responsible for cleaving the target DNA strand. It contains three divided RuvC subdomains, which fold intricately to form the active site where DNA cleavage occurs. These subdomains work together to recognize and cleave DNA at a specific location indicated by the guide RNA. The BH domain, or bridge helix domain, functions as a structural link between different domains of the Cas protein. It is characterized as an arginine-rich region composed of numerous arginine amino acids. These abundant arginine residues are crucial for interaction with the phosphate backbone of the target DNA strand. The arginine residues form hydrogen bonds with the phosphate group, assisting in the proper positioning and orientation of the DNA to be cleaved. The REC domain (recognition lobe) is involved in the recognition of the target DNA sequence. By distinguishing between target and non-target sequences, it contributes to the specificity of the Cas protein, ensuring that only the intended DNA fragment is cleaved. The HNH domain is another catalytic site that works in conjunction with the RuvC domain to cleave the complementary strand of target DNA. Named after its characteristic histidine-asparagine-histidine sequence, it is essential for the nuclease activity of the Cas protein. The CTD, or C-terminal domain, is often involved in interactions with other proteins and cellular structures, contributing to the localization and regulation of the Cas protein within cells. It also plays a role in maintaining the stability and overall structure of the Cas protein, ensuring its functionality and specificity for its target. This complex domain configuration allows type II Cas proteins to perform precise and crucial functions within the CRISPR system, making them invaluable tools in genome editing and manipulation.
[0110] In some embodiments, the Cas protein further comprises an effector domain (or functional domain). Such an effector domain may have one or more types of enzymatic activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain comprises a nuclease, nickase, deaminase, reverse transcriptase, recombinase, methyltransferase, methylase, acetylase, acetyltransferase, transcription activator, transcription repressor domain, cryptochrome, photoinducible / controllable domain, or chemoinducible / controllable domain.
[0111] In some embodiments, the Cas protein further includes one or more nuclear localization signal sequences, nuclear export signal sequences, cell membrane permeable peptide sequences, and affinity tags. A type II Cas protein includes one or more nuclear localization signals (NLS). The NLS can be located at the end or other site of the peptide. The NLS located at each end or other site of the Cas9 amino acid sequence may be identical or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are identical. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains one NLS, and the C-terminus of the Cas9 amino acid sequence contains one NLS. The amino acid sequences of the NLS are fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence, respectively. The NLS may be an SV40 (simian virus 40) NLS, a c-Myc NLS, or another suitable monosegmented NLS. The NLS may be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, affinity tags (such as GST, FLAG, or hexahistidine sequences) are used for the purification of Cas proteins by affinity chromatography. In some embodiments, the amino acid sequence of the C-terminal NLS is described in SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is described in SEQ ID NO: 883. Other available sequences and different combinations of NLS and FLAG sequences are also selectable.
[0112] In some embodiments, the Cas protein contains an amino acid sequence having 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to any one of the amino acid sequences of SEQ ID NOs: 1 to 71.
[0113] In some preferred embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 44%, 51, 56, 59, 60, and 68 sequence identity with respect to the amino acid sequences of SEQ ID NOs. 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, and 68. In several other preferred embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 44%, 51%, 56%, 59, 60, or 68 sequence identity with respect to any one amino acid sequence of sequence numbers 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, or 68, except for the amino acid "M" at position 1 of the sequence.
[0114] One of the key elements of the modification process associated with the CRISPR-Cas system is the protospacer adjacent motif (PAM), which is a short DNA sequence adjacent to the target DNA sequence. PAM sequences are essential for the binding and cleavage activity of Cas proteins, ensuring that only the intended DNA sequence is edited. The ability of Cas proteins to recognize these specific PAM sequences is due to their unique structural features and binding interactions with DNA. Each PAM sequence provides a distinct motif that Cas proteins can recognize, ensuring accurate and efficient targeting. For example, the NRRANH sequence may provide a specific nucleotide sequence to which Cas proteins can bind with high affinity. These diverse PAM sequences expand the potential application range of the CRISPR-Cas system. This allows researchers and scientists to broaden the range of DNA sequences they can edit, increasing the versatility and effectiveness of this powerful gene editing tool. Furthermore, understanding these PAM sequences contributes to the development of Cas proteins with more advanced targeting capabilities, further advancing the field of gene editing and its potential benefits in research, medicine, and biotechnology. The Cas proteins disclosed in this invention exhibit a unique ability to recognize diverse PAM sequences. These sequences include NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. Specific recognition of these Cas proteins provides greater flexibility in selecting the DNA sequences to be edited.
[0115] In the PAM sequences presented in this invention, N represents one of four standard DNA bases: adenine (A), thymine (T), cytosine (C), or guanine (G). This code facilitates the inclusion of any nucleotide at specific positions, eliminating the need for individual specification. R represents a purine base in particular, i.e., adenine (A) or guanine (G). Purine bases are important in the structure and function of DNA, and this code simplifies the incorporation of these larger nucleotides into PAM sequences. H represents any nucleotide except guanine (G). Thus, it can represent adenine (A), thymine (T), or cytosine (C). This code is useful for excluding guanine (G) from specific positions where its size might affect the structural or functional aspects of the sequence. Y represents a pyrimidine nucleotide, representing either thymine (T) or cytosine (C). Pyrimidines are frequently found in specific regions of DNA and RNA, and this code allows for their inclusion without specifying the exact nucleotide. W: Represents a weak base, indicating either adenine (A) or thymine (T). This distinction is biochemically important, as adenine and thymine exhibit similar properties under certain conditions, such as hydrogen bonding. V: Reflects any nucleotide except thymine (T), including adenine (A), cytosine (C), or guanine (G). This code is useful in situations where thymine is undesirable or unnecessary in pyrimidines due to its specific chemical properties. M: Represents either adenine (A) or cytosine (C).
[0116] In some embodiments, the Cas proteins disclosed herein are capable of recognizing at least one protospacer adjacent motif (PAM) having or containing the sequence NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 9, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the sequence of NRRANH. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the sequence of NRHACT. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 19, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the sequence of NRAAR.In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 21, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the NNNCCY sequence. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 22, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the sequence NNRYYYY. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 24, and is capable of recognizing protospacer adjacent motifs (PAMs) having the sequence of NGG. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 25, and is capable of recognizing protospacer adjacent motifs (PAMs) having the sequence NNNNCAA.In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 27, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the sequence of NRNACN. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 29, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the NNGR sequence. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 30, and is capable of recognizing protospacer adjacent motifs (PAMs) having the sequence of NGGNR. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 31, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the NNNCCH sequence.In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 36, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the sequence of NRRAAG. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 37, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the sequence of NRHRAC. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 38, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the sequence of NRYART. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 43, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the NRHACC sequence.In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 44, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the sequence of NRAAR. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 51, and is capable of recognizing protospacer adjacent motifs (PAMs) having the NRNVHH sequence. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 56, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the YMACAW sequence. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 59, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the NAHAA sequence.In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 60, and is capable of recognizing protospacer-adjacent motifs (PAMs) having the NRHAYY sequence. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 68, and is capable of recognizing protospacer adjacent motifs (PAMs) having the sequence of NGGHA.
[0117] As presented in this invention, the terms “recognized,” “recognize,” or “recognize” in this context refer to the ability of a Cas protein to form a functional complex with sgRNA at a DNA target site where sgRNA hybridizes (i.e., the sgRNA's spacer sequence hybridizes) and which is laterally surrounded by a PAM sequence, and the ability of the Cas protein to perform its innate function (i.e., DNA cleavage or DNA binding). In this context, such DNA cleavage excludes the type II Cas protein from being a catalytically inactive type II Cas nuclease. For example, in the case of an inactivated type II Cas nuclease (e.g., dead type II Cas nuclease), if the necessary PAM sequence is present, a complex may be formed between the type II Cas nuclease, sgRNA, and the corresponding target, but this does not result in DNA cleavage.
[0118] In some embodiments, the Cas protein is a nickase or a dead Cas protein. The DNA cleavage domain of the active Cas protein in this invention comprises two subdomains: an HNH nuclease subdomain and a RuvC subdomain. Mutations within these subdomains can silence the nuclease activity of the Cas protein. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12, and includes a mutation at residue D11 or H859; or at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83% , having 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and containing a mutation at residue D10 or H862; or having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 31 and containing a mutation at residue D12 or H903.In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 9, and includes mutations in residue D11 or H859, except for amino acid "M" at position 1 of the sequence; or at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 8 Having 6%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, and containing mutations in residue D10 or H862 except for amino acid "M" at position 1 of the sequence; or having at least 70%, 71%, 72%, 73%, 74% of the amino acid sequence of SEQ ID NO: 31, The sequences have 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, and contain a mutation in residue D12 or H903, except for the amino acid "M" at position 1 of the sequence. In some embodiments, a mutation in residue D11 or H859 in SEQ ID NO: 9 is D11A or H859A; a mutation in residue D10 or H862 in SEQ ID NO: 12 is D10A or H862A; or a mutation in residue D12 or H903 in SEQ ID NO: 31 is D12A or H903A. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to any one amino acid sequence of sequence numbers 869-877.
[0119] The present invention also provides engineered, non-natural polynucleotides encoding type II CRISPR-related (Cas) proteins disclosed herein.
[0120] As presented herein, the term “polynucleotide” refers to a polymer of nucleotides of any length, consisting of deoxyribonucleotides, ribonucleotides, or their analogues. Polynucleotides may have any known or unknown three-dimensional structure and may perform any function. Accordingly, the term includes, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and pyrimidine bases, or other natural, chemically or biochemically modified, unnatural, or derivatized nucleotide bases. In some embodiments, polynucleotides encoding multiple portions of the expressive type II Cas protein described herein may be functionally linked to each other and to relevant regulatory sequences (such as promoters, enhancers, or termination regions). For example, a functional linkage may exist between a regulatory sequence and an exogenous nucleic acid sequence, resulting in the expression of the latter. In other embodiments, the first nucleic acid sequence may be operably linked to the second nucleic acid sequence if the first nucleic acid sequence is arranged to have a functional relationship with the second nucleic acid sequence. For example, if a promoter affects the transcription or expression of a coding sequence, the promoter is operably ligated to the coding sequence. Generally, operably ligated DNA sequences are contiguous and, where necessary or useful, link coding regions to the same reading frame. In some embodiments, the promoter is a constitutive promoter, a tissue-specific promoter, or an inducible promoter. In some embodiments, the term may further include all introns and other DNA sequences spliced from mRNA transcripts, along with variants derived from alternative splicing sites. These nucleic acid sequences may be DNA strand sequences transcribed to RNA, or RNA sequences translated to proteins. Nucleic acid sequences include both full-length nucleic acid sequences and incomplete-length sequences derived from full-length proteins. Sequences may also include degenerate codons of native sequences, or sequences that may be introduced to provide codon preference in specific cell types.
[0121] In some embodiments, the polynucleotide encoding the Cas protein is operably ligated to a promoter and present within the vector; optionally, the vector is selected from the group consisting of retroviral vectors, lentiviral vectors, phage vectors, adenovirus vectors, adeno-associated virus vectors, herpes simplex virus vectors, and plasmid vectors.
[0122] In some embodiments, the polynucleotide is a ribonucleotide sequence or a deoxyribonucleotide sequence, or an analog thereof; optionally, the polynucleotide is codon-optimized for expression in the target cell; in some embodiments, the polynucleotide is codon-optimized for expression in eukaryotic cells. In some embodiments, the eukaryotic cell is selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cell is a mammalian cell, preferably a human cell.
[0123] In some embodiments, the polynucleotide is mRNA and further comprises a 5' cap sequence and / or a poly-A tail sequence. In some embodiments of the present invention, the mRNA used may be modified to enhance its functional properties and stability. Specifically, in some embodiments, the modification process involves substituting uridine (represented by the letter "U") with N1-methylpseudouridine or pseudouridine. This substitution is designed to improve the mRNA's resistance to ribonuclease degradation and potentially increase its intracellular half-life and translation efficiency. Incorporating N1-methylpseudouridine or pseudouridine into the mRNA structure may also have a positive impact on the immune response profile, as these modifications have been shown to reduce the immunogenicity of the mRNA molecule compared to unmodified mRNA molecules. This is particularly important in the development of mRNA-based therapeutics and vaccines where minimizing adverse immune responses is paramount.
[0124] In some embodiments, the polynucleotides of the present invention are codon-optimized for expression in eukaryotic cells; optionally, the eukaryotic cells are selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and / or human cells.
[0125] In some embodiments, the polynucleotide has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to any one nucleotide sequence among sequence numbers 161-231, 241-311, 851-852, and 861-863.
[0126] The present invention also provides an engineered non-natural type CRISPR-Cas system comprising: a) a type II Cas protein as described herein, or a polynucleotide encoding the Cas protein; b) at least one engineered guide RNA, or at least one engineered nucleic acid encoding the guide RNA, wherein the guide RNA comprises a spacer sequence complementary to the target nucleic acid and a Cas protein-binding segment that interacts with the Cas protein, and the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) duplex.
[0127] As presented in this invention, the term “complementary” describes the ability of two nucleic acid strands to pair with each other via their bases. This complementarity can occur via perfect base pairing, where each base along one strand forms a specific hydrogen bond pair with its complementary base on the opposite strand according to the standard Watson-Crick base pairing rules: adenine (A) pairs with thymine (T) or uracil (U), and cytosine (C) pairs with guanine (G). This perfect base pairing enables precise recognition and binding between the two strands. Furthermore, complementarity may also involve incomplete base pairing, including situations where mismatches, insertions, or deletions result in non-standard base pairing or reduced affinity between strands. Despite these incompletenesses, the strands retain sufficient complementarity to maintain interaction, although this may potentially reduce specificity and stability in the double helix.
[0128] As described in this invention, direct repeat (DR) sequences originate from short repetitive DNA sequence elements within a CRISPR array. These sequences are typically interspersed among spacer sequences derived from exogenous genetic material, such as phages or plasmid DNA. After transcription, the DR sequences become part of the pre-crRNA transcript and are processed into mature CRISPR RNA (crRNA). The DR sequences in the crRNA function as crucial components in forming a double-stranded RNA (dsRNA) duplex by pairing with complementary sequences within trans-activated CRISPR RNA (tracrRNA). This duplex facilitates Cas protein binding and is essential for the function of the guide RNA (gRNA) complex in the CRISPR / Cas system. In the context of this invention, the DR sequence includes a portion of gRNA that hybridizes with the tracrRNA sequence to form a dsRNA duplex, thereby enabling the formation of a Cas protein-binding segment. The phrase "hybridizes to form a double-stranded RNA (dsRNA) duplex," or similar expressions in this invention, refers to the process by which the direct repeat (DR) sequence of CRISPR RNA (crRNA) pairs with a complementary sequence within transactivated CRISPR RNA (tracrRNA) to form a stable double-stranded RNA structure. This dsRNA duplex is an important component of the guide RNA (gRNA) complex that facilitates Cas protein binding. Specifically, the dsRNA duplex is formed when the DR sequence in crRNA and the complementary sequence in tracrRNA form base pairs, and hydrogen bonds are generated between the complementary bases. This duplex is essential for the function of the Cas protein-binding segment, which is composed of the tracrRNA sequence and DR sequence that hybridize to form the dsRNA duplex.
[0129] As presented in this invention, the term “target nucleic acid” refers to a specific nucleic acid substrate containing a nucleic acid sequence complementary to all or part of the spacer in the RNA guide. In some embodiments, the target nucleic acid includes a gene or a sequence within a gene. In certain embodiments, the target nucleic acid includes a non-coding region (e.g., a promoter). In certain embodiments, the target nucleic acid is single-stranded. In certain embodiments, the target nucleic acid is double-stranded. The terms “target nucleic acid” or “target sequence” should be understood in accordance with the context of this disclosure.
[0130] In some embodiments, the guide RNA is a dual guide RNA. In some embodiments, the guide RNA is a single guide RNA. In such embodiments, the guide RNA further includes a linker sequence that ligates the tracrRNA sequence and the DR sequence. In some typical embodiments, the linker includes a short sequence GAAA. In some embodiments, the linker functions as an artificial loop. In some embodiments, the sgRNA includes a) a spacer sequence that can hybridize with the sequence of the target nucleic acid to be manipulated; b) a DR sequence; c) a linker sequence; and d) a tracrRNA sequence. The serial arrangement of the spacer sequence, DR sequence, linker sequence, and tracrRNA sequence is in the 5' to 3' direction or the 3' to 5' direction. In some embodiments, the sgRNA scaffold contains sequences having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% homology to any one of sequence numbers 431-562.
[0131] In some embodiments, the sgRNA comprises a spacer sequence (e.g., one of sequence numbers 571-835, 885, or 901) and a scaffold sequence, the spacer sequence located at the 5' end of the scaffold sequence (e.g., the sequence of sequence number 903). In some embodiments, the sgRNA comprises a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to sequence number 903.
[0132] In some embodiments, the spacer sequence hybridizes with one or more nucleic acids in a prokaryotic or eukaryotic cell. In some embodiments, the eukaryotic cell is selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the eukaryotic cell includes mammalian cells. In some embodiments, the mammalian cell includes human cells. In some embodiments, the eukaryotic cell includes plant cells.
[0133] In some embodiments, the system further includes a donor template nucleic acid.
[0134] As presented herein, the term “donor template nucleic acid” as used herein refers to a nucleic acid molecule that can be used to alter the structure of a target nucleic acid by one or more cellular proteins after the Cas protein described herein has altered the target nucleic acid. In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid. In some embodiments, the donor template nucleic acid is a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear. In some embodiments, the donor template nucleic acid is circular (e.g., a plasmid). In some embodiments, the donor template nucleic acid is an exogenous nucleic acid molecule. In some embodiments, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, the donor template nucleic acid is DNA, RNA, or a DNA-RNA hybrid.
[0135] The present invention also provides a modified vector comprising the polynucleotide described herein.
[0136] As presented herein, a “vector” is a means of enabling or facilitating the transfer of an entity from one environment to another. It is a replicon such as a plasmid, phage, or cosmid into which another DNA segment can be inserted, resulting in the replication of the inserted segment. Generally, a vector is replicable when associated with appropriate regulatory elements. Generally, the term “vector” refers to a nucleic acid molecule capable of carrying other bound nucleic acids. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules containing one or more free ends, nucleic acid molecules without free ends (e.g., circular); nucleic acid molecules containing DNA, RNA, or both; and other types of polynucleotides known in the art. One form of vector is a “plasmid,” which refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted using standard molecular cloning techniques, for example. Another form of vector is the viral vector, in which viral DNA or RNA sequences are present within the vector for packaging into viruses (e.g., retroviruses, replication-deficient retroviruses, adenoviruses, replication-deficient adenoviruses, adeno-associated viruses (AAVs)). Viral vectors also contain polynucleotides carried by the virus for transfection into host cells. Certain vectors can autonomously replicate within the host cell into which they are introduced (e.g., bacterial vectors with bacterial origins of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the host cell's genome upon introduction and thereby replicate together with the host genome. Furthermore, certain vectors can induce the expression of functionally linked genes. Such vectors are referred to herein as “expression vectors.” Common expression vectors useful in recombinant DNA technology often take the form of plasmids.A recombinant expression vector may contain the nucleic acid of the present invention in a form suitable for nucleic acid expression in a host cell, meaning that the recombinant expression vector includes one or more regulatory elements that can be selected based on the host cell used for expression and are functionally linked to the nucleic acid sequence to be expressed. Within a recombinant expression vector, "functionally linked" means that the nucleotide sequence of interest is linked to the regulatory element and intended to enable the expression of that nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). In some embodiments, the vector is an expression vector. In some embodiments, the vector is an inductive, conditional, or constitutive expression vector. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA reside on the same vector or on different vectors.
[0137] The present invention also provides a vector system comprising one or more polynucleotides encoding one or more polynucleotides and a guide RNA as described herein, wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein-binding segment that interacts with the Cas protein, and the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) duplex. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA reside on the same vector or on different vectors.
[0138] The present invention also provides engineered non-natural cells comprising the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, or the vector system described herein.
[0139] The present invention also provides cells modified using the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, or the vector system described herein.
[0140] In some embodiments, the cells are eukaryotic or prokaryotic. In some embodiments, the eukaryotic cells are selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cells are mammalian cells, human cells, or plant cells.
[0141] In some embodiments, the cells are cells of vertebrates, mammals, rodents, goats, pigs, birds, chickens, turkeys, cattle, horses, sheep, fish, primates, or humans. In some embodiments, the cells are mammalian cells. In one embodiment, the cells are human cells. In some embodiments, the cells are somatic cells, germ cells, or embryonic cells. In some embodiments, the cells are zygote cells, blastocyst cells, embryonic cells, stem cells, mitotic cells, or meiotic cells. In some embodiments, the cells are not part of a human embryo. In some embodiments, the cells are somatic cells. In one embodiment, the cells include T cells, CD8+ T cells, CD8+ naive T cells, central memory T cells, effector memory T cells, CD4+ T cells, stem cell-like memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricting progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythroid progenitor cells, megakaryocyte-erythroid progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, trabecular meshwork cells, cochlear hair cells, outer hair cells, inner hair cells, and lung epithelial cells. These include bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, neurons, neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells (e.g., progenitor B cells, pre-B cells, pro-B cells, memory B cells, plasma B cells), gastrointestinal epithelial cells, bile duct epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, pre-adipocytes, islet cells (e.g., β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0142] In some embodiments, the disclosure includes a modified target locus, where the target locus is modified according to the method, use of a composition, or use of a system of the present invention.
[0143] In some embodiments, the cells are eukaryotic or prokaryotic. In some embodiments, the eukaryotic cells are selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cells are mammalian cells, human cells, or plant cells.
[0144] In some embodiments, the cells are cells of vertebrates, mammals, rodents, goats, pigs, birds, chickens, turkeys, cattle, horses, sheep, fish, primates, or humans. In some embodiments, the cells are mammalian cells. In one embodiment, the cells are human cells. In some embodiments, the cells are somatic cells, germ cells, or embryonic cells. In some embodiments, the cells are zygote cells, blastocyst cells, embryonic cells, stem cells, mitotic cells, or meiotic cells. In some embodiments, the cells are not part of a human embryo. In some embodiments, the cells are somatic cells. In one embodiment, the cells include T cells, CD8+ T cells, CD8+ naive T cells, central memory T cells, effector memory T cells, CD4+ T cells, stem cell-like memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricting progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythroid progenitor cells, megakaryocyte-erythroid progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, trabecular meshwork cells, cochlear hair cells, outer hair cells, inner hair cells, and lung epithelial cells. These include bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, neurons, neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells (e.g., progenitor B cells, pre-B cells, pro-B cells, memory B cells, plasma B cells), gastrointestinal epithelial cells, bile duct epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, pre-adipocytes, islet cells (e.g., β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0145] In some embodiments, the vector (e.g., plasmid or viral vector) is delivered to the target tissue by means of, for example, intramuscular injection, intravenous administration, transdermal administration, nasal administration, oral administration, or mucosal administration. Such delivery may be performed as a single dose or multiple doses. Those skilled in the art will understand that the actual dose administered herein may vary considerably depending on various factors, including the choice of vector, target cells, organisms, tissues, the general condition of the subject being treated, the degree of transformation or modification desired, the route of administration, the method of administration, and the type of transformation or modification desired.
[0146] In certain embodiments, delivery is carried out via adeno-associated viruses (AAVs), such as AAV2, AAV8, or AAV9, which are at least 1 × 10⁻¹⁶ 5 It can be administered as a single dose containing adenovirus or adeno-associated virus particles (also called particle units or pu). In some embodiments, the dose is at least about 1 × 10⁻⁶ 6 Particles, at least about 1 × 10⁻⁶ 7 Particles, at least about 1 × 10⁻⁶ 8 Particles, or at least about 1 × 10⁻⁶ 9 It is a particle-type adeno-associated virus. Due to the limited genome payload of recombinant AAV, the smaller size of the type II Cas nuclease described herein increases the versatility of packaging the effector and RNA guide together with the appropriate regulatory sequences (e.g., promoter) necessary for efficient and cell-type-specific expression.
[0147] In some embodiments, delivery is carried out via recombinant adeno-associated virus (rAAV) vectors. For example, in some embodiments, modified AAV vectors can be used for delivery. Modified AAV vectors may be based on one or more capsid types, including AAV1, AAV2, AAV5, AAV6, AAV8, AAV8.2, AAV9, AAV rh10, modified AAV vectors (e.g., modified AAV2, modified AAV3, modified AAV6), and pseudotype AAVs (e.g., AAV2 / 8, AAV2 / 5, AAV2 / 6).
[0148] In some embodiments, delivery is via plasmids. The dose can be a sufficient number of plasmids to induce a response. In some embodiments, the appropriate amount of plasmid DNA in a plasmid composition is in the range of about 0.1 mg to about 2 mg. A plasmid generally includes (i) a promoter; (ii) a sequence encoding a nucleic acid-targeted CRISPR enzyme operably ligated to the promoter; (iii) a selection marker; (iv) an origin of replication; and (v) a transcription termination sequence located downstream of (ii) and operably ligated. Plasmids can also encode RNA components of the CRISPR-Cas system, although one or more of these may instead be encoded on different vectors. The frequency of administration is at the discretion of medical or veterinary professionals (e.g., physicians, veterinarians) or those skilled in the art.
[0149] The present invention also provides a kit comprising the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, the vector system described herein, or the cells described herein.
[0150] The kits described herein contain in one or more containers the components essential for carrying out the methods described herein and may include instructions for use as appropriate. Any of the kits may additionally include auxiliary components necessary for carrying out the preparation method. Each component in the kit is provided, where applicable, in liquid form (e.g., dissolved in solution) or solid form (e.g., lyophilized powder). In certain embodiments, some components can be reconstituted or otherwise processed (e.g., to an active state) by adding a suitable solvent or other substance (e.g., water or buffer), and such solvent or substance may or may not be included in the kit. In some embodiments, the kit may further include other suitable excipients, such as buffers and reagents, to facilitate the application of the kit. Preferably, the kits can be applied to a variety of uses, such as medical applications including therapeutic and diagnostic applications, and research applications. Thus, the type II Cas nucleases and kits of the present invention can be used for the preparation of pharmaceuticals for therapeutic purposes and / or for the preparation of reagents for research purposes.
[0151] The Cas proteins, CRISPR-Cas systems, and polynucleotides described herein can be delivered by various delivery systems, such as vectors (e.g., plasmids), viral delivery vectors such as adeno-associated viruses (AAV), lentiviruses, and other viral vectors, or by nuclear translocation or electroporation of ribonucleoprotein complexes consisting of a type VI effector and its corresponding RNA guide(s). The proteins and one or more RNA guides can be packaged into one or more vectors (e.g., plasmids or viral vectors). In bacterial applications, nucleic acids encoding any component of the CRISPR systems described herein can be delivered to bacteria using phages. Representative phages include, but are not limited to, T4 phage, Mu, λ phage, T5 phage, T7 phage, T3 phage, Φ29, M13, MS2, Qβ, and ΦX174.
[0152] The present invention also provides a pharmaceutical composition comprising any of the Cas proteins, polynucleotides, CRISPR-Cas systems, vectors, vector systems, or cells described herein.
[0153] As used herein, the term “pharmaceutical composition” refers to a formulation intended for pharmaceutical use. In certain embodiments, the pharmaceutical composition further comprises pharmaceutically acceptable excipients. In some embodiments, the pharmaceutical composition may contain additional therapeutic agents. In some embodiments, the pharmaceutical composition is prepared according to standard procedures for administration to a subject, such as a human patient, via intravenous, intramuscular, intradermal, intraarticular, intralesional, intraperitoneal, intracardiac, intrathecal, intraventricular, epidural, topical, subconjunctival, intrastromal, periorbital, intravitreous, parascleral, perscleral, choroidal, retroorbital, subretinal, subtenon’s capsule, nasal inhalation, pressurized inhalation, oral, subcutaneous, or topical routes. For example, a composition for injection may be provided as a sterile isotonic aqueous solution. If necessary, the pharmaceutical composition may also contain solubilizers such as lidocaine and local anesthetics to minimize discomfort at the injection site. Generally, the components are supplied alone or as mixed unit doses (e.g., lyophilized powder or anhydrous concentrated solution) in sealed containers indicating the amount of the active ingredient. When a pharmaceutical composition is intended for intravenous administration, it can be combined with an infusion bottle containing sterile pharmaceutical-grade water or saline solution. When a pharmaceutical composition is intended for injection, it may include sterile water or saline solution for injection to allow for pre-administration mixing of components. Furthermore, humectants, colorants, release agents, coating agents, sweeteners, fragrances, aromas, preservatives, and antioxidants may be incorporated into the formulation as needed.
[0154] In some embodiments, the pharmaceutical composition further comprises a delivery system selected from AAV (adeno-associated virus), adenovirus, retrovirus, HSV (herpes simplex virus), gamma retrovirus, LV (lentivirus), eCIS (extracellular contractile injection system), eVLPs (modified virus-like particles), VLPs (virus-like particles), liposomes, plasmids, LNPs (lipid nanoparticles), exosomes, microvesicles, nucleic acid nanoassemblies, gene guns, and / or implantable devices.
[0155] The present invention also provides for the use of the Cas proteins, polynucleotides, CRISPR-Cas systems, vectors, vector systems, cells, kits, or pharmaceutical compositions described herein for the treatment, prevention, diagnosis, or detection of diseases.
[0156] The present invention also provides a method for modifying or targeting a target DNA locus, comprising delivering to the locus a Cas protein as described herein; a polynucleotide as described herein; a CRISPR-Cas system as described herein; a vector as described herein; a vector system as described herein; a kit as described herein; or a pharmaceutical composition as described herein.
[0157] In some embodiments, the Disclosure also provides methods for targeting and cleaving target DNA, comprising contacting the target DNA with a Cas protein as described herein; a polynucleotide as described herein; a CRISPR-Cas system as described herein; a vector as described herein; a vector system as described herein; a kit as described herein; or a pharmaceutical composition as described herein.
[0158] In some embodiments, the modification or targeting of the target locus includes inducing DNA strand breaks. In some embodiments, the modification or targeting of the target locus includes inducing DNA double-strand breaks or DNA single-strand breaks. In some embodiments, the modification or targeting of the target locus includes altering the gene expression of one or more genes. In some embodiments, the modification or targeting of the target locus includes epigenetic modification of the target DNA locus. In some embodiments, the method is a method for modifying cells, cell lines, or organisms by manipulating one or more target sequences at a genomic locus of interest.
[0159] In some embodiments, cleavage of target DNA or target sequence results in indel formation or insertion of a nucleotide sequence. In some embodiments, cleavage of target DNA or target nucleotide involves cleaving the target DNA or target sequence at two locations, resulting in a deletion or inversion of the sequence between the two locations. In some embodiments, the target DNA is double-stranded DNA, single-stranded DNA, or a DNA-RNA hybrid.
[0160] In some embodiments, modification or targeting of the target gene locus includes induction of DNA strand breaks, alteration of gene expression of one or more genes, or epigenetic modification of the target DNA gene locus; optionally, the DNA strand breaks include DNA double-strand breaks or DNA single-strand breaks.
[0161] In some embodiments, the method is performed in vitro or in vivo.
[0162] The present invention also provides isolated eukaryotic cells comprising a modified target locus of interest, wherein the target locus of interest is modified by the method described in this disclosure, or by using the system described in this disclosure, or by using the Cas protein described in this disclosure; or by using the polynucleotide described in this disclosure; or by using the CRISPR-Cas system described in this disclosure; or by using the vector described in this disclosure, or by using the vector system described in this disclosure, or by using the kit described in this disclosure, or by using the pharmaceutical composition described in this disclosure.
[0163] The present invention also provides a system for detecting the presence of a nucleic acid target sequence in an in vitro sample, comprising: a) a Cas protein as described herein; b) at least one guide polynucleotide comprising a guide sequence capable of binding to a target sequence and designed to form a complex with the Cas protein; and c) a nucleic acid-based masking construct comprising a non-target sequence, wherein the Cas protein exhibits contingent cleavage activity of RNA and / or ssDNA, cleaving the non-target sequence of the nucleic acid-based masking construct activated by the target sequence.
[0164] The present invention also provides a method for detecting target nucleic acids in a sample, comprising the steps of contacting one or more samples with a nucleic acid-based masking construct comprising a) a Cas protein as described herein; b) at least one guide polynucleotide having a degree of complementarity with a target sequence and having a guide sequence designed to form a complex with the Cas protein; and c) a non-target sequence, wherein the Cas protein exhibits contingent cleavage activity of RNA and / or ssDNA, cleaving the non-target sequence of the nucleic acid-based masking construct activated by the target sequence, and detecting a signal from the cleavage of the non-target sequence, thereby detecting one or more target sequences in the sample.
[0165] As presented in this disclosure, “Sample” may include whole cells and / or live cells and / or cell residue. Sample may include (or be derived from) “Body fluids.” This disclosure includes embodiments in which body fluids are selected from amniotic fluid, aqueous humor, vitreous fluid, bile, plasma, breast milk, cerebrospinal fluid, earwax, chyle, cerebrospinal fluid, endolymph, perilymph, exudate, feces, female ejaculate, gastric acid, gastric juice, lymph, mucus (including nasal discharge and sputum), pericardial fluid, peritoneal fluid, pleural fluid, pus, eye discharge, saliva, sebum, semen, sputum, synovial fluid, sweat, tears, urine, vaginal secretions, vomit, and mixtures of one or more of these. Samples include cell cultures, body fluids, and cell cultures derived from body fluids. Body fluids can be obtained from mammalian individuals, for example, by puncture or other collection / sampling procedures.
[0166] The present invention also provides a guide RNA (gRNA) comprising: a) a spacer sequence of Sequence ID No. 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of Sequence ID No. 901; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of Sequence ID No. 901. In some embodiments, the gRNA further comprises a Cas protein-binding segment; and the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) duplex. In some embodiments, the gRNA is a dual guide RNA. In some embodiments, the gRNA is a single guide RNA. In some embodiments, the gRNA is modified. In some embodiments, at least three nucleotides of the gRNA are modified. In some embodiments, the gRNA includes a 5'-terminus modification containing at least two phosphorothioate (PS) bonds within the first seven nucleotides from the 5' end. In some embodiments, the gRNA includes a 3'-terminus modification containing at least two phosphorothioate (PS) bonds within the first seven nucleotides from the 3' end. In some embodiments, the gRNA has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the nucleotide sequence of SEQ ID NO: 903.
[0167] The present invention also provides a polynucleotide encoding a gRNA, comprising: a) a spacer sequence of sequence number 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of sequence number 901; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 901. In some embodiments, the gRNA contains a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with respect to sequence number 903.
[0168] An engineered non-natural CRISPR-Cas system comprising: a) a Cas protein, or a polynucleotide encoding the Cas protein; b) at least one guide RNA (gRNA) as described herein, or at least one design nucleic acid encoding the guide RNA; wherein the gRNA further comprises a Cas protein-binding segment that interacts with the Cas protein; and the Cas protein-binding segment comprises a tracrRNA sequence and hybridizes to double-stranded RNA (dsR The present invention provides an engineered, non-natural CRISPR-Cas system comprising: NA) a direct repeat (DR) sequence forming a duplex; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of sequence number 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of sequence number 901; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 901. In some embodiments, the gRNA comprises a) the sequence of sequence number 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to sequence number 903.
[0169] In some embodiments, the Cas protein comprises a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the sequence of SEQ ID NO: 12; or, The Cas protein contains a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of Sequence ID No. 12, except for the amino acid "M" at position 1 of the sequence.
[0170] In some embodiments, the Cas protein further comprises an effector domain (or functional domain). Such an effector domain may have one or more types of enzymatic activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain comprises a nuclease, nickase, deaminase, reverse transcriptase, recombinase, methyltransferase, methylase, acetylase, acetyltransferase, transcription activator, transcription repressor domain, cryptochrome, photoinducible / controllable domain, or chemoinducible / controllable domain.
[0171] In some embodiments, the Cas protein further includes one or more nuclear localization signal sequences, nuclear export signal sequences, cell membrane permeable peptide sequences, and affinity tags. A type II Cas protein includes one or more nuclear localization signals (NLS). The NLS can be located at the end or other site of the peptide. The NLS located at each end or other site of the Cas9 amino acid sequence may be identical or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are identical. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains one NLS, and the C-terminus of the Cas9 amino acid sequence contains one NLS. The amino acid sequences of the NLS are fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence, respectively. The NLS may be an SV40 (simian virus 40) NLS, a c-Myc NLS, or another suitable monosegmented NLS. The NLS may be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, affinity tags (such as GST, FLAG, or hexahistidine sequences) are used for the purification of Cas proteins by affinity chromatography. In some embodiments, the amino acid sequence of the C-terminal NLS is described in SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is described in SEQ ID NO: 883. Other available sequences and different combinations of NLS and FLAG sequences are also selectable.
[0172] In some embodiments, the Cas protein contains an amino acid sequence having 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12.
[0173] In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12, and is capable of recognizing protospacer adjacent motifs (PAMs) having the sequence of NRHACT.
[0174] In some embodiments, the Cas protein is a nickase or dead Cas protein. The DNA cleavage domain of the active Cas protein in the present invention comprises two subdomains: an HNH nuclease subdomain and a RuvC subdomain. Mutations within these subdomains can silence the nuclease activity of the Cas protein. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12, and includes a mutation at residue D10 or H862. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12, and the Cas protein contains a mutation at residue D10 or H862, except for the amino acid "M" at position 1 of the sequence.
[0175] In some embodiments, the mutation at residue D10 or H862 of SEQ ID NO: 12 is D10A or H862A. In some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to any one amino acid sequence of SEQ ID NOs: 872-874.
[0176] The present invention also provides a modified vector comprising a polynucleotide encoding the gRNA described herein, comprising: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 901. In some embodiments, the gRNA comprises a) the sequence of sequence number 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to sequence number 903.
[0177] In some embodiments, the vector is an inducible, conditional, or constitutive expression vector.
[0178] The present invention also provides a vector system comprising one or more polynucleotides encoding the gRNA described herein, comprising: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 871. In some embodiments, the gRNA comprises a) the sequence of sequence number 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to sequence number 903.
[0179] The present invention also provides pharmaceutical compositions comprising the gRNA described herein; a polynucleotide encoding the gRNA; a CRISPR-Cas system comprising the gRNA or polynucleotide; a vector comprising the gRNA coding sequence; and a vector system comprising the gRNA coding sequence; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; and c) at least 99%, 98%, or 9% of the sequence of SEQ ID NO: 901. The gRNA includes a spacer sequence having 7%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity; in some embodiments, the gRNA includes a) the sequence of SEQ ID NO: 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO: 903.
[0180] The present invention also provides a method for treating, preventing, or diagnosing a disease associated with the RHO locus in a subject, wherein the guide RNA (gRNA) comprises: a) the spacer sequence of Sequence ID No. 901; b) the spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of Sequence ID No. 901; c) the spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of Sequence ID No. 901. In some embodiments, the gRNA comprises a) the sequence of sequence number 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to sequence number 903.
[0181] As described in this disclosure, the “RHO locus”-related diseases refer to a group of disorders characterized by mutations in the RHO gene. The RHO gene encodes the rhodopsin protein, which is essential for photoreceptor function in the retina. Mutations in this gene cause a variety of retinal degenerative diseases that primarily affect rod photoreceptor cells, potentially leading to vision loss or blindness. These diseases usually manifest as hereditary retinal dystrophy, including retinitis pigmentosa, which is characterized in particular by progressive vision loss due to degeneration of rod photoreceptor cells. Patients with mutations in the RHO gene may experience night blindness or gradual loss of peripheral vision, and these conditions are often inherited in an autosomal dominant manner.
[0182] The present invention also provides a method for treating, preventing, or diagnosing a locus-related disease in a subject, wherein the guide RNA (gRNA) comprises i) the spacer sequence of Sequence ID No. 901; ii) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of Sequence ID No. 901; or iii) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity to the sequence of Sequence ID No. 901. In some embodiments, the gRNA comprises the sequence of sequence number 903; or a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to sequence number 903.
[0183] The present invention also provides a composition comprising the following components: (i) an RNA guide type DNA binder, a. The RNA-guide DNA binder contains a sequence having at least 90% homology to SEQ ID NO: 12 or 92; and / or RNA-guide DNA conjugates include sequences having at least 95%, 96%, 97%, 98%, 99%, or 100% homology to SEQ ID NO: 12 or 92; and / or (ii) an sgRNA containing the sgRNA sequence of Sequence ID No. 903, or a vector encoding said sgRNA.
[0184] The present invention also provides a method for modifying an RHO gene locus, comprising the step of delivering a composition to cells, wherein the composition is a. Guide RNA containing the guide sequence of SEQ ID NO: 901; b. A guide RNA containing at least 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; or c. A guide RNA containing a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 901;
[0185] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases in a subject, comprising administering a composition to a subject of interest, wherein the composition is a. Guide RNA containing the guide sequence of SEQ ID NO: 901; b. A guide RNA containing at least 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; or c. A guide RNA containing a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 901;
[0186] The present invention also provides a method for modifying an RHO gene locus, comprising the step of delivering a composition to cells, wherein the composition is a. sgRNA containing the sgRNA sequence of sequence number 903; b. An sgRNA containing an sgRNA sequence that has at least 90% identity with the sequence of Sequence ID No. 903; or c. An sgRNA containing an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 903;
[0187] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases in a subject, comprising administering a composition to a subject of interest, wherein the composition is a. sgRNA containing the sgRNA sequence of sequence number 903; b. An sgRNA containing an sgRNA sequence that has at least 90% identity with the sequence of Sequence ID No. 903; or c. An sgRNA containing an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 903;
[0188] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases in a subject, comprising administering a composition to a subject of interest, wherein the composition is a. Guide RNA containing the spacer sequence of SEQ ID NO: 901; b. A guide RNA containing at least 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; or c. A guide RNA containing a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of sequence number 901;
[0189] The present invention also provides a method for modifying an RHO gene locus, comprising the step of delivering a composition to cells, wherein the composition is (i) Cas protein, here a. The Cas protein contains a sequence that has at least 90% identity with SEQ ID NO: 12 or 92; and / or RNA-guide DNA conjugates include sequences having at least 95%, 96%, 97%, 98%, 99%, or 100% homology to SEQ ID NO: 12 or 92; and / or (ii) A guide RNA or vector encoding a guide RNA containing the spacer sequence of SEQ ID NO: 901;
[0190] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases in a subject, comprising administering a composition to a subject of interest, wherein the composition is (i) RNA guide DNA binding agent, however, a. The RNA-guide DNA binder contains a sequence having at least 90% identity with SEQ ID NO: 12 or 92; and / or RNA-guide DNA conjugates include sequences having at least 95%, 96%, 97%, 98%, 99%, or 100% homology to SEQ ID NO: 12 or 92; and / or (ii) an sgRNA containing the sequence of sequence number 903 or a vector encoding said sgRNA.
[0191] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8] [Table 1-9] [Table 1-10] [Table 1-11] Table 1-12 Table 1-13 Table 1-14 Table 1-15 Table 1-16 Table 1-17 Table 1-18 Table 1-19 Table 1-20 Table 1-21 Table 1-22 Table 1-23 Table 1-24 Table 1-25 Table 1-26 Table 1-27 Table 1-28 Table 1-29 Table 1-30 Table 1-31 Table 1-32 Table 1-33 Table 1-34 Table 1-35 Table 1-36 Table 1-37 Table 1-38 Table 1-39 Table 1-40 Table 1-41 Table 1-42 Table 1-43 Table 1-44 Table 1-45 Table 1-46 Table 1-47 Table 1-48 Table 1-49 Table 1-50 Table 1-51 Table 1-52 Table 1-53 Table 1-54 Table 1-55 Table 1-56 Table 1-57 Table 1-58 Table 1-59 Table 1-60 Table 1-61 Table 1-62 Table 1-63 Table 1-64 Table 1-65 Table 1-66
Table 1-67
Table 1-90
Table 1-97
Table 1-100
Table 1-110
Table 1-120
[0192] Table 2-1 Table 2-2 Table 2-3 Table 2-4 Table 2-5 Table 2-6 Table 2-7 Table 2-8 Table 2-9 Table 2-10 Table 2-11 Table 2-12 Table 2-13 Table 2-14 Table 2-15 Table 2-16 Table 2-17 Table 2-18 Table 2-19 Table 2-20 Table 2-21 Table 2-22 Table 2-23 Table 2-24 Table 2-25 Table 2-26 Table 2-27 Table 2-28 Table 2-29 Table 2-30 Table 2-31 Table 2-32 Table 2-33 Table 2-34 Table 2-35 Table 2-36 Table 2-37 Table 2-38 Table 2-39 Table 2-40 Table 2-41 Table 2-42 Table 2-43 Table 2-44 Table 2-45 Table 2-46 Table 2-47 Table 2-48 Table 2-49 Table 2-50 Table 2-51 Table 2-52 Table 2-53 Table 2-54 Table 2-55 Table 2-56 Table 2-57 Table 2-58 Table 2-59 Table 2-60 Table 2-61 Table 2-62 Table 2-63 Table 2-64 Table 2-65 Table 2-66 Table 2-67 Table 2-68 Table 2-69 Table 2-70 Table 2-71 Table 2-72 Table 2-73 Table 2-74 Table 2-75 Table 2-76 Table 2-77 Table 2-78 Table 2-79 Table 2-80 Table 2-81 Table 2-82 Table 2-83 Table 2-84 Table 2-85 Table 2-86 Table 2-87 Table 2-88 Table 2-89 Table 2-90 Table 2-91 Table 2-92 Table 2-93 Table 2-94 Table 2-95 Table 2-96 Table 2-97 Table 2-98 Table 2-99 Table 2-100 Table 2-101 Table 2-102
Table 2-103
Table 2-120
Table 2-203
Table 2-300
Table 2-301
Table 2-302
Table 2-309
Table 2-341
Table 2-343
Table 2-345
Table 2-347
Table 2-351
Table 2-356
Table 2-360
Table 2-362
Table 2-365
Table 2-370
Table 2-378
Table 2-379
Table 2-380
Table 2-381
Table 2-387
Table 2-388
Table 2-389
[0193] Table 3-1 Table 3-2 [Table 3-3] [Table 3-4] [Table 3-5]
[0194] In some embodiments, the disclosure provides engineered non-natural crRNAs comprising a nucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs: 341-417 (Table 3), or a variant thereof. In some embodiments, the crRNA comprises a nucleotide sequence having at least 95% or 98% sequence identity to any one of SEQ ID NOs: 341-417. In some embodiments, the crRNA comprises a nucleotide sequence described in any one of SEQ ID NOs: 341-417. [Examples]
[0195] (Embodiments for carrying out the invention) The following non-limiting embodiments are provided to further illustrate embodiments of the disclosure disclosed herein. Those skilled in the art will understand that the techniques disclosed in the following embodiments represent approaches that have been found to work well in the implementation of the disclosure and can therefore be considered to constitute exemplary forms for implementation. Those skilled in the art will understand that, in light of the disclosure, many modifications can be made to the specific embodiments disclosed to obtain similar or similar results without departing from the spirit and scope of the disclosure.
[0196] Example 1: Method for metagenomic analysis of proteins We searched for metagenomic sequence data from publicly available databases using a hidden Markov model generated based on known Cas protein sequences, including class 2 type II Cas effector proteins. The CRISPR-Cas proteins identified by the search were aligned with known proteins to identify potential active sites. After screening hundreds of potential sequences, this metagenomic workflow led to the identification of type II Cas proteins, detailed in Table 1.
[0197] A phylogenetic tree was generated using MUSCLE 3 (Veen et al., 2020) to explore the relationships between orthologs at the primary amino acid level. This study utilized hundreds of class II type II-A / B / C sequences collected from the National Center for Biotechnology Information (NCBI) and various publications and patents. Notably, the phylogenetic tree suggests that the Cas protein described in this invention is different from known ones. The type II Cas protein detailed in this disclosure has extremely low homology to other known Cas proteins.
[0198] Structural modeling of type II Cas proteins was achieved using AlphaFold2. Annotation of domain configuration revealed that the Cas proteins disclosed herein contain a RuvC domain, a BH (bridge helix) domain, a REC domain, an HNH domain, and a CTD (C-terminal domain). In particular, the RuvC domain, along with the BH domain, contains three distinct RuvC subdomains. Figure 1 visually shows the structures of several representative proteins.
[0199] Example 2: Protocol for predictive RNA folding The predicted RNA folding of the Cas protein's guide RNA sequence and the sequence it is presumed to follow was calculated using the RNAfold web server developed by Lorenz et al. in 2011.
[0200] Example 3: PAM determination in mammalian cell lines In a series of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco TM ). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco TM ). 450 μL of cells at a density of 120,000 cells / well were mixed with 50 μL of a mixture containing Lipofectamine TM 3000 (ThermoFisher Scientific, Cat. L3000008), Opti-Mem adjusted to a volume of 50 μL, 1 μL of dsODN (2.5 pM), 100 ng (about 1 μL) of psgRNA encoding the sgRNA scaffold (Table 6) and the Humanspacer3 spacer (SEQ ID NO: 885, Table 4), and 400 ng (about 1 μL) of pCas protein plasmid (Table 2) containing the Cas9 CDS encoding NLS and FLAG, according to the manufacturer's protocol. Next, the cell mixture was seeded into a 24-well plate and cultured under conditions of 37 °C and 5% CO2. 10 pM of dsODN was annealed using dsODN-Top and dsODN-Bot oligonucleotides before transfection.
[0201] <F 72 hours after transfection, the supernatant was removed and the cell layer was washed with PBS. Next, genomic DNA was extracted from each well of the 24-well plate using a DNA extraction solution (Denogen (Beijing) Bio Sci & Tech Co., Ltd, Cat. DNS033-48). All DNA samples (500 ng, 260 / 280 value: 1.8 - 2.0) were subjected to Guide-Seq NGS analysis.
[0202] The basic method for preparing Guide-Seq libraries is described by Nikolay et al. (Nat. Protoc. 2021). The extracted DNA sample was first fragmented using the KAPA Frag Kit (Cat#KK8602, Roche). After purification of the fragmented DNA, it was phosphorylated with T4 polynucleotide kinase (Cat#M0201S, NEB). SS5 adapters (10 μM SS5TOP oligo and 10 μM SS5B) were used. TM (Generated by oligo annealing) via Quick Ligation TM We performed a two-step off-target PCR using Kit (Cat#M2200S, NEB) to ligate fragmented DNA, followed by the addition of chemical modifications for sequencing.
[0203] Off-target PCR1 is Platinum TM Off-target PCR 2 was performed using Taq DNA Polymerase (product number #15966005, Invitrogen), GSP1 (a mixture of GSP1-Top and GSP1-BoT), and Y_XX oligonucleotides. TM The procedure was performed using Taq DNA Polymerase and GSP2 (a mixture of GSP2-TopA / B / C and GSP1-BoTA / B / C), Y_XX (the same as in PCR1), and i753_XX oligonucleotides. DNA products from each step must be purified using SPRI Select (product number #B23318, Beckman Coulter). The final library was quantified by qPCR and sequenced on an Illumina NextSeq 1000. Reads with low quality scores were removed and then mapped to the reference genome. The Q30 rate was greater than 0.9. Read lengths ranged from 130 bp to 140 bp. The result file containing the reads was mapped to the reference genome (BAM file), and reads overlapping with the target region were selected.
[0204] [Table 4]
[0205] Note: "p" indicates phosphorylation modification; "*" indicates phosphorothioate (PS) bond; "N" indicates any natural or non-natural nucleotide.
[0206] Figure 2 and Table 5 show the PAM selectivity of representative Cas proteins in the HEK293 cell line when using the corresponding sgRNA scaffold.
[0207] [Table 5-1] [Table 5-2]
[0208] Example 4: In vitro editing efficiency screening of the CRISPR-Cas system in mammalian cell lines In a series of experiments, HEK293T cells were treated with 10% fetal bovine serum (Gibco). TM They were cultured in DMEM medium supplemented with ). In reverse transfection, HEK293T cells were transfected with 10% fetal bovine serum (Gibco). TM Cells were cultured in DMEM medium supplemented with ). 24 hours prior to transfection, 250 μL of cells at a cell density of 50,000 cells / well were seeded into 48-well plates. Lipofectamine was added according to the manufacturer's protocol. TMCells were transfected with lipoplex containing 3000 (0.4 μL / well), P3000 (2 μL / well), pgRNA / pCas protein plasmids (125 ng / well and 375 ng / well, respectively), and Opti-Mem up to a total of 25 μL / well. Seeded cells were fixed and adhered for 72 hours in a tissue culture incubator at 37°C and 5% CO2. The nucleotide sequences of the pgRNAs used in this example included sequences encoding the corresponding Cas sgRNA scaffolds (Table 6; SEQ ID NOs. 465 (GEBx0305), SEQ ID NOs. 451 (GEBx0308), SEQ ID NOs. 454 (GEBx0308-Rq-V3), SEQ ID NOs. 478 (spCas9), SEQ ID NOs. 491 and SEQ ID NOs. 548) and corresponding spacer sequences (Table 7).
[0209] 72 hours after transfection, the supernatant was removed and the cell layer was washed with PBS. Next, genomic DNA was extracted from each well of a 24-well plate using DNA extraction solution (Denogen (Beijing) Bio Sci & Tech Co. Ltd, Cat.DNS033-48). Amplicon NGS analysis was performed on all DNA samples (500 ng, 260 / 280 values: 1.8-2.0).
[0210] To quantitatively evaluate the editing efficiency in target regions of the genome, NGS was used to detect the presence or absence of insertions and deletions introduced by gene editing. NGS primers targeting the area around the target region within endogenous genes were designed. Additional PCR was performed according to the manufacturer's protocol to add the chemical modifications necessary for sequencing. The amplified products were sequenced using Illumina iSeq 100. Reads with low quality scores were removed and then mapped to the reference genome. The Q30 rate was greater than 0.9. Read lengths were between 130 bp and 140 bp. The file containing the obtained reads was mapped to the reference genome (BAM file), and reads overlapping with the target region were selected. The number of wild-type reads and the number of reads containing insertions, substitutions, and deletions were calculated. The number of reads mapped to the reference genome was over 1000.
[0211] Table 6-1 Table 6-2 Table 6-3 Table 6-4 Table 6-5 Table 6-6 Table 6-7 Table 6-8 Table 6-9 Table 6-10 Table 6-11 Table 6-12 Table 6-13 Table 6-14 [Table 6-15] [Table 6-16] [Table 6-17]
[0212] The spacer arrays were located at the 5' end of the scaffold. Furthermore, the 3' end of each spacer array was directly connected to the 5' end of the subsequent scaffold array, forming a characteristic repeat-spacer pattern.
[0213] [Table 7-1] [Table 7-2] [Table 7-3] [Table 7-4] [Table 7-5] [Table 7-6] [Table 7-7] [Table 7-8] [Table 7-9] Table 7-10 Table 7-11 Table 7-12 Table 7-13 Table 7-14
[0214] Figure 3 shows the indel levels of GEBx0305 against 16 targets with GGAAAA-PAM in the HEK293T cell line. The sgRNA sequence used for GEBx0305 in this experiment was the GEBx0305-HPT-V1 scaffold (WT, SEQ ID NO: 465) with 20nt spacer sequences (SEQ ID NOs: 571, 573, 575, 577, 579, 581, 583, 585, 587, 589, 591, 593, 595, 597, 599, and 601). SpCas9 targeting the corresponding sites (SEQ ID NOs: 572, 574, 576, 578, 580, 582, 584, 586, 588, 590, 592, 594, 596, 598, 600, 602) was used as a positive control. Figure 4 shows the indel levels of GEBx0308 against 19 targets containing GGTACT-PAM in the HEK293T cell line. The sgRNA sequence used for GEBx0308 in this experiment consisted of the GEBx0308-PT-V1 scaffold (WT, SEQ ID NO: 451) and 20nt spacer sequences (SEQ ID NOs: 611, 613, 615, 617, 619, 621, 623, 625, 627, 629, 631, 633, 635, 637, 639, 641, 643, 645, 647, 649). SpCas9 targeting the corresponding sites (SEQ ID NOs: 612, 614, 616, 618, 620, 622, 624, 626, 628, 630, 632, 634, 636, 638, 640, 642, 644, 646, 648, 650) are used as positive controls.
[0215] Figure 5 shows the indel levels of GEBx0308 against 15 targets containing CATACT-PAM in the HEK293T cell line. The psgRNA sequence used for GEBx0308 in this experiment contains the GEBx0308-Rq-V3 scaffold (M0, SEQ ID NO: 454) and a 20nt spacer sequence (SEQ ID NOs: 659-674).
[0216] Figure 6 shows the indel levels of GEBx0328 against 20 targets containing NNGCCT-PAM in the HEK293T cell line. The sgRNA sequence used for GEBx0328 in this experiment was the GEBx0328-PT-V1 scaffold (WT, SEQ ID NO: 491) with a 20nt spacer sequence (SEQ ID NOs: 765-784). GEBx0328 showed moderate indel activity across all 20 targets.
[0217] Figures 7 and 8 show the indel levels of GEBx0361 against 19 targets containing GGTACC or TGTACC PAM in the HEK293T cell line. The sgRNA sequence used for GEBx0361 in this experiment was the GEBx0361-HPT-V2 scaffold (SEQ ID NO: 548) with a 20nt spacer sequence (SEQ ID NOs: 817-835). GEBx0361 showed moderate indel activity across all 19 targets.
[0218] Example 5: Optimization of an exemplary CRISPR / Cas system To further improve Cas protein editing efficiency, variable guide length and sgRNA scaffolds were tested. In a series of experiments, HEK293T cells were treated with 10% fetal bovine serum (Gibco). TM They were cultured in DMEM medium supplemented with ). In reverse transfection, HEK293T cells were transfected with 10% fetal bovine serum (Gibco). TM Cells were cultured in DMEM medium supplemented with ). 24 hours prior to transfection, 250 μL of cells at a cell density of 50,000 cells / well were seeded into 48-well plates. Lipofectamine was added according to the manufacturer's protocol. TMCells were transfected with lipoplex containing 3000 (0.4 μL / well), P3000 (2 μL / well), pgRNA / pCas protein plasmids (125 ng / well and 375 ng / well, respectively), and Opti-Mem, up to a total of 25 μL / well. Seeded cells were fixed and adhered for 72 hours in a tissue culture incubator at 37°C and 5% CO2. The pgRNA sequence used in this embodiment includes sequences encoding the corresponding sgRNA scaffold (SEQ ID NO: 465 (GEBx0305), SEQ ID NO: 451 (GEBx0308), or SEQ ID NO: 491 (GEBx0328)) and one of the corresponding spacers (SEQ ID NOs: 575, 599, or 695-707 for GEBx0305; SEQ ID NOs: 625, 676, or 708-721 for GEBx0308; SEQ ID NOs: 769, 782, 785-798 for GEBx0328).
[0219] As shown in Figure 9, the indel level varies depending on the length of the GEBx0305 guide sequence. The sgRNA sequences used in this experiment had a GEBx0305-HPT-V1 scaffold (WT, SEQ ID NO: 465) and incorporated either a CFTR-NGGAAAA-T3 or POLQ-NGGAAAA-T5 spacer sequence (18nt to 25nt, SEQ ID NOs: 575, 599, or 695-707). The CFTR-NGGAAAA-T3-23nt and POLQ-NGGAAAA-T5-25nt spacers showed the best indel levels in each group.
[0220] Figure 10(A and B) shows that the indel level changes depending on the length of the GEBx0308 guide sequence. The sgRNA sequences used in this experiment were the GEBx0308-PT-V1 scaffold (WT, SEQ ID NO: 451) and CFTR-NGGTACT-T4 or CD34-TATACT-T2 spacer sequences (18nt to 25nt, SEQ ID NOs: 625, 676, or 708-721). The CFTR-NGGTACT-T4-21nt spacer showed the highest indel levels compared to spacers of other lengths.
[0221] Figure 11(A and B) shows that the indel level changes depending on the length of the GEBx0328 guide sequence. The sgRNA sequences used in this experiment were the GEBx0328-PT-V1 scaffold (WT, SEQ ID NO: 491) and TTR-NGGCCT-T1 or CFTR-NGGCCT-T5 spacer sequences (18nt to 25nt, SEQ ID NOs: 769, 782, 785-798). The TTR-NGGCCT-T1-24nt and CFTR-NGGCCT-T5-24nt spacers showed the highest indel levels in each group.
[0222] In another series of experiments, HEK293T cells were introduced into 10% fetal bovine serum (Gibco). TM The cells were cultured in DMEM medium supplemented with (Gibco). In reverse transfection, HEK293T cells were transfected with 10% fetal bovine serum (Gibco). TM Cells were cultured in DMEM medium supplemented with ). 24 hours prior to transfection, 100 μL of cell suspension at a cell density of 25,000 cells / well was seeded into 96-well plates. Following the manufacturer's protocol, Lipofectamine was added. TM Cells were transfected with lipoplex prepared to a total of 25 μL / well of 3000 (0.4 μL / well), P3000 (2 μL / well), pCas protein-gRNA plasmid (300 ng / well), and Opti-Mem. The seeded cells were fixed and adhered for 72 hours in a tissue culture incubator at 37°C and 5% CO2. The Cas protein-gRNA sequence used in this example consists of Cas CDS, Cas sgRNA scaffold (Table 6), and corresponding spacer sequence (Table 7).
[0223] In another series of experiments, HEK293T cells were introduced into 10% fetal bovine serum (Gibco). TM The cells were cultured in DMEM medium supplemented with (Gibco). In reverse transfection, HEK293T cells were transfected with 10% fetal bovine serum (Gibco). TMCells were cultured in DMEM medium supplemented with ). 24 hours prior to transfection, 100 μL of cell suspension at a cell density of 25,000 cells / well was seeded into 96-well plates. Following the manufacturer's protocol, Lipofectamine was added. TM Cells were transfected with lipoplex prepared to a total of 25 μL / well of 3000 (0.4 μL / well), P3000 (2 μL / well), pCas protein-gRNA plasmid (300 ng / well), and Opti-Mem. The seeded cells were fixed and adhered for 72 hours in a tissue culture incubator at 37°C and 5% CO2. The nucleotide sequence of the Cas protein-gRNA used in this example consisted of Cas CDS, Cas sgRNA scaffold, and corresponding spacers.
[0224] Figure 12(A and B) shows the indel levels of GEBx0305 targeting endogenous genes using modified RNA scaffolds. The sgRNA sequences used in this experiment had scaffolds of GEBx0305-HPT-V1 (WT, SEQ ID NO: 465), GEBx0305-Rq-V2 (M0, SEQ ID NO: 466), or GEBx0305-M1~M6 (SEQ ID NOs: 467~472) and included 20nt spacers targeting eight endogenous gene sites (SEQ ID NOs: 575, 579, 585, 587, 589, 591, 599, 601). The GEBx0305-M5 and M6 scaffolds showed the mean values of the top two indels.
[0225] Figure 13(A and B) shows the indel levels of GEBx0308 targeting endogenous genes using modified RNA scaffolds. The sgRNA sequences used in this experiment had scaffolds of GEBx0308-PT-V1 (WT, SEQ ID NO: 451), GEBx0308-Rq-V3 (M0, SEQ ID NO: 454), or GEBx0308-M1~M6 (SEQ ID NOs: 455~460) and contained 20nt spacers targeting seven endogenous gene sites (SEQ ID NOs: 621, 625, 627, 631, 637, 639, 641). The GEBx0308-M3 and M4 scaffolds showed the mean values of the top two indels.
[0226] Figure 14(A and B) shows the indel levels of GEBx0305 targeting endogenous genes under optimal conditions. The optimized sgRNA sequence used in this experiment had a GEBx0305-M5 (SEQ ID NO: 471) scaffold and included 21nt spacers targeting 27 endogenous gene sites (SEQ ID NOs: 652, 654, 655, 658, 696, 703, 722-725, 729-745). Compared to a WT sgRNA (GEBx0305-HPT-V1 scaffold (SEQ ID NO: 465) with 20nt spacers (SEQ ID NOs: 571, 573, 575, 577, 579, 581, 583, 585, 587, 589, 591, 593, 595, 597, 599, 651-658, 746-749)), the optimized sgRNA significantly improved the indel level of GEBx0305 (P<0.0001).
[0227] Figure 15(A and B) shows the indel levels of GEBx0308 targeting endogenous genes under optimal conditions. The optimized sgRNA sequence used in this experiment had a GEBx0308-M4 (SEQ ID NO: 458) scaffold and included 21nt spacers targeting 25 endogenous gene sites (SEQ ID NOs: 710, 659, 665, 667, 668, 672, 674, 726, 727, 728, 750-764). Compared to WT sgRNA (GEBx0308-PT-V1 scaffold (SEQ ID NO: 451) with a 20nt spacer (SEQ ID NOs: 611, 613, 615, 617, 619, 621, 623, 625, 627, 629, 631, 633, 635, 637, 639, 641, 645, 647, 659, 662, 663, 665, 667, 668, 672, 674) added), the optimized sgRNA slightly improved the indel level of GEBx0308 (P=0.03).
[0228] Figure 16(A and B) shows the indel levels of GEBx0328 in endogenous gene targeting using modified RNA scaffolds. The sgRNA sequences used in this experiment were GEBx0328-PT-V1 (WT, SEQ ID NO: 491), GEBx0328-Rq-V1 (M0, SEQ ID NO: 492), or GEBx0328-M1 to M8 (SEQ ID NOs: 493-500) scaffolds, each containing 20nt spacers targeting eight endogenous gene sites (SEQ ID NOs: 765, 767, 769, 770, 771, 776, 782, 783). The GEBx0328-M6 scaffold showed the highest average indel value.
[0229] Figure 17(A and B) shows the indel levels of GEBx0328 targeting endogenous genes under optimal conditions. The optimized sgRNA sequence used in this experiment was the GEBx0328-M6 scaffold (SEQ ID NO: 498) with 24nt spacers targeting 19 endogenous gene sites (SEQ ID NOs: 790, 797, 799-807, 809-816). Compared to WT sgRNA (GEBx0328-PT-V1 scaffold with 20nt spacers), the optimized sgRNA significantly improved the indel levels of GEBx0328 (P=0.0002).
[0230] Example 6: In vitro gene editing using RNA In this experiment, HEK293T cells were incubated with 10% fetal bovine serum (Gibco). TM Cells were cultured in DMEM medium supplemented with ) 24 hours before transfection, 100 μL of cell suspension at a cell density of 25,000 cells / well was seeded into a 96-well plate. Lipofectamine was added according to the manufacturer's protocol. TM RNAiMAX (Invitrogen TM A lipoplex containing sgRNA and RNA (20 ng sgRNA, sgRNA:mRNA = 1:1 to 1:8 w / w ratio) was prepared and adjusted to 25 μL / well using Opti-Mem, and transfection was performed on cells. The seeded cells were fixed and adhered for 72 hours in a tissue culture incubator at 37°C and 5% CO2. The sgRNA and mRNA sequences are shown in Table 8. The mRNA used in this example is N1-methyl-pseudouridine modified uracil.
[0231] After thawing primary human hepatocytes (PHH), they were resuspended in hepatocyte thawing medium with added supplements (Lonza, Cat.MCHT50) and centrifuged at 100g for 10 minutes. The supernatant was removed, and the precipitated cells were resuspended in hepatocyte culture medium (Lonza, product number MP100) supplemented with 10% fetal bovine serum. The number of cells was counted, and they were seeded at a density of 40,000 cells / well in an Ultra Low Adsorption Cell Culture 96-well plate (River Biotech, product number LV-ULA002-96W). The seeded cells were fixed and adhered for 24 hours in a tissue culture incubator at 37°C and 5% CO2. After incubation, cell monolayer formation was confirmed, and the medium was resuspended in hepatocyte culture medium (Lonza, product number CC-3198) supplemented with 10% fetal bovine serum. Lipofectamine was added according to the manufacturer's protocol. TM RNAiMAX (Invitrogen TM Cells were transfected with lipoplex containing ) and RNA (30 ng gRNA, gRNA:mRNA = 1:1 to 1:16 w / w) and Opti-Mem up to 25 μL / well. The seeded cells were fixed and adhered for 72 hours in a tissue culture incubator at 37°C and 5% CO2.
[0232] The methods for genomic DNA extraction and NGS are the same as those described in Example 4.
[0233] Figure 18 shows the indel levels of endogenous gene-targeting GEBx0305 after transfection of HEK293T cells with a lipoplex containing a fixed amount (20 ng) of sgRNA targeting the POLQ-NGGAAAA-T5 site (SEQ ID NO: 855) and GEBx0305 mRNA (SEQ ID NO: 851) in different ratios. Equal amounts of SpCas9 mRNA and sgRNA (SEQ ID NOs: 853 and 854) were used as positive controls. GEBx0305 exhibits indel efficiency comparable to that of SpCas9.
[0234] Figure 19 shows the indel levels of endogenous gene-targeting GEBx0308 after transfection of HEK293T cells with a lipoplex containing a fixed amount (20 ng) of sgRNA targeting the CFTR-NGGTACT-T4 or POLQ-NGGTACT-T1 site (SEQ ID NO: 856 / 857) and GEBx0308 mRNA (SEQ ID NO: 852) in different ratios.
[0235] Figure 20 shows the indel levels of endogenous gene-targeting GEBx0305 after transfection of PHH cells with a lipoplex containing a fixed amount (20 ng) of sgRNA targeting the POLQ-NGGAAAA-T5 site (SEQ ID NO: 855) and GEBx0305 mRNA (SEQ ID NO: 851) in different ratios. Equal amounts of SpCas9 mRNA and sgRNA (SEQ ID NOs: 853, 854, Table 9) were used as positive controls.
[0236] Figure 21 shows the indel levels of endogenous gene-targeting GEBx0308 after transfection of PHH cells with lipoplexes containing a fixed amount (20 ng) of sgRNA (SEQ ID NO: 857) targeting the POLQ-NGGTACT-T1 site and different ratios of GEBx0308 mRNA (SEQ ID NO: 852, Table 8).
[0237] [Table 8-1] [Table 8-2] [Table 8-3] [Table 8-4] [Table 8-5] [Table 8-6] [Table 8-7] [Table 8-8] [Table 8-9] [Table 8-10]
[0238] Example 7: Off-target analysis in cell lines using GUIDE-Seq GUIDE-Seq utilizes a technique that inserts dsODN into double-strand break sites generated by CRISPR / Cas. HEK293T cells are fed 5% fetal bovine serum (Gibco). TM Cells were cultured in advanced DMEM medium supplemented with ). 24 hours prior to transfection, cells were seeded in 24-well plates at a density of 100,000 cells / well. Following the manufacturer's protocol, 400 ng of pCas protein plasmid, 150 ng of pgRNA plasmid, and 2.5 pmol of dsODN were transfected with Lipofectamine 3000 (Invitrogen). TM Transfection was performed using ), followed by incubation at 37°C and 5% CO2 conditions, and the cells were harvested on the third day after transfection.
[0239] For GUIDE-Seq library construction, 500 ng of genomic DNA was used. Briefly, the DNA was fragmented using the KAPA Frag Kit (KAPA Biosystems), followed by adapter ligation, and two heminested PCR amplifications were performed on the dsODN-integrated fragments. The final sequenced library was quantified using the KAPA Library Quantification Kit and sequenced on an Illumina NextSeq 1000 system. Data demultiplexing of index 1 was performed using bcl2fq (version 2.19), followed by demultiplexing of index 2 using a custom script, adapter trimming with the BBduk tool, and analysis using GUIDE-seq software. Briefly, FASTQ files tagged with Unique Molecular Index (UMI) were integrated to generate a UMI consensus sequence, which was aligned to the human reference genome (hg19) using BWA MEM. High-quality alignments (MAPQ≧50) were used to identify genomic loci containing dsODNs as potential off-target sites. Candidate loci with up to six mismatches to the corresponding on-target protospacer were identified as true off-target sites.
[0240] Figure 22 shows an overview of the main Guide-seq insertion sites of GEBx0305. No detectable off-targets were detected at site 1 (CFTR-NGGAAAA-T5) and site 2 (EMX1-NGGAAAA-T5).
[0241] Figure 23 shows an overview of the main Guide-seq insertion sites of GEBx0308. No off-target events were detected at site 1 (CD34-NGGTACT-T4) and site 2 (POLQ-NGGTACT-T1).
[0242] Figure 24 shows an overview of the main Guide-seq insertion sites of GEBx0328. No detectable off-target sites were detected at site 2 (CFTR-NGGCCT-T5), and only one off-target site was detected at site 2 (CFTR-NGGCCT-T3).
[0243] Example 8: Base editing efficiency assay of nCas proteins To generate the Cas base editor, E. coli tRNA adenosine deaminase (TadA-8e) was fused to the N-terminus of GEBx0305, GEBx0308, and GEB328 nicases (GEBx0305-D11A, GEBx0308-D10A, GEB328-D12A) using an ABE linker (SEQ ID NO: 868). HEK293T cells were then fused to 10% fetal bovine serum (Gibco). TM The cells were cultured in DMEM medium supplemented with (Gibco). In reverse transfection, HEK293T cells were transfected with 10% fetal bovine serum (Gibco). TM Cells were cultured in DMEM medium supplemented with ). 24 hours prior to transfection, 100 μL of cell suspension at a cell density of 25,000 cells / well was seeded into 96-well plates. Lipofectamine was added according to the manufacturer's protocol. TM Cells were transfected using lipoplex prepared with 3000 (0.4 μL / well), P3000 (2 μL / well), pABE-gRNA plasmid (300 ng / well), and Opti-Mem, totaling 25 μL / well. Seeded cells were fixed and adhered for 72 hours in a tissue culture incubator at 37°C and 5% CO2. The pABE / gRNA plasmid used in this example consists of Cas-ABE CDS (SEQ ID NOs. 861, 862, 863, Table 9), Cas sgRNA scaffold, and corresponding spacer. 72 hours after transfection, the supernatant was removed and the cell layer was washed with PBS. Genomic DNA was then extracted from each well of a 24-well plate using DNA extraction solution (Denogen (Beijing) Bio Sci & Tech Co. Ltd, Cat. DNS033-48). All DNA samples were subjected to amplicon NGS analysis.
[0244] To quantitatively evaluate the editing efficiency in target regions of the genome, NGS was used to detect the presence or absence of insertions and deletions introduced by gene editing. NGS primers targeting the area around the target region within endogenous genes were designed. Additional PCR was performed according to the manufacturer's protocol to add the chemical modifications necessary for sequencing. The amplified products were sequenced using Illumina iSeq 100. Reads with low quality scores were removed and then mapped to the reference genome. The Q30 rate was greater than 0.9. Read lengths were between 130 bp and 140 bp. The resulting read-containing file was mapped to the reference genome (BAM file), reads overlapping with the target region were selected, and the number of wild-type reads and reads including A-to-G substitutions were calculated. The number of reads mapped to the reference genome was over 1000.
[0245] [Table 9-1] [Table 9-2] [Table 9-3] [Table 9-4] [Table 9-5] [Table 9-6] [Table 9-7] [Table 9-8] Table 9-9 Table 9-10 Table 9-11 Table 9-12 Table 9-13 Table 9-14 Table 9-15 Table 9-16 Table 9-17 Table 9-18 Table 9-19 Table 9-20 Table 9-21 Table 9-22
[0246] Figure 25 shows the frequency of A-to-G conversion editing of adenine bases at five endogenous gene sites by GEBx0305-ABE in HEK293T cells. The sgRNA sequence used in this experiment had a GEBx0305-M5 (SEQ ID NO: 471) scaffold and 21nt spacers (SEQ ID NOs: 703, 722, 723, 725). GEBx0305-ABE showed efficient A-to-G conversion at these sites.
[0247] Figure 26 shows the frequency of adenine A-to-G conversion base editing at five endogenous gene loci in HEK293T cells by GEBx0308-ABE. The sgRNA sequence used in this experiment had a GEBx0308-M4 (sequence number 458) scaffold and 21nt spacers (sequence numbers 710, 726, 727, 728, 667). GEBx0308-ABE showed efficient A-to-G conversion at these sites.
[0248] Figure 27 shows the frequency of adenine A-to-G conversion base editing at five endogenous gene loci in HEK293T cells by GEBx0328-ABE. The sgRNA sequence used in this experiment had a GEBx0328-M6 (SEQ ID NO: 498) scaffold and 24nt spacers (SEQ ID NOs: 790, 797, 801, 809, 815). GEBx0328-ABE showed efficient A-to-G conversion at these loci.
[0249] Example 8: Evaluation of specific knockout of GEBx0308 against the RHO-P23H site in HEK293T cells Editor plasmid and target plasmid design: The editor plasmid was designed and created to contain both the SaCas9 (SEQ ID NO: 906) or GEBx0308 coding gene (SEQ ID NO: 252) and the corresponding guide sequences (RHO-P23H sgRNA (SaCas9) (SEQ ID NO: 904), which is a single-stranded guide RNA for SaCas9, and RHO-P23H sgRNA (SEQ ID NO: 902), which is a single-stranded guide RNA for GEBx0308). The target plasmid was designed and created to contain RHO-WT (approximately 170 bp), RHO-P23H (approximately 170 bp, c.68C>A), and a 500 bp unrelated sequence to separate RHO-WT and RHO-P23H.
[0250] In the specific knockout evaluation of GEBx0308 against the RHO-P23H site in HEK293T cells, HEK293T cells were found to be affected by 2 mM L-glutamine (GlutaMAX). TM 10% fetal bovine serum (GEMINI) was added to Dulbecco's modified Eagle medium (DMEM, CORNING) containing -1, Gibco, and the cells were incubated at 37°C in a 5% CO2 buffered incubator.
[0251] To evaluate the specific knockout efficiency of SaCas9 (SEQ ID NO: 906, Table 10) and GEBx0308 (SEQ ID NO: 252) against the RHO-P23H site, editor and target plasmids were co-introduced into HEK293T cells. 1.0 × 10⁶ cells were placed in each well of a 24-well plate. 5 Cells were seeded. Editor plasmids and target plasmids were mixed in 2:1 and 3:1 (w / w) ratios and co-transfected into HEK293T cells using Lipofectamine 3000 reagent according to the manufacturer's instructions (Life Technologies). After 72 hours of culture, cells were harvested and lysed to release the target plasmid. The editing efficiency of the RHO-WT and RHO-P23H sites was evaluated by NGS analysis using the editable target plasmid as a template and target plasmid-specific primer pairs. The results are shown in Figure 28.
[0252] [Table 10-1] [Table 10-2] [Table 10-3] [Table 10-4]
[0253] Preferred embodiments of the present invention have been shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided for illustrative purposes only. The present invention is not intended to be limited by the specific examples provided herein. The present invention has been described with reference to the foregoing specification, but the descriptions and illustrations of embodiments herein are not intended to be constrained. Those skilled in the art will understand that numerous variations, modifications, and substitutions may occur without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific descriptions, configurations, or relative proportions described herein, depending on various conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in the practice of the present invention. Accordingly, the present invention is intended to also encompass such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and methods and structures within the scope of these claims and their equivalents are intended to be encompassed thereby.
Claims
1. An engineered, non-natural type II CRISPR-related (Cas) protein or a variant thereof, having at least 70% sequence identity with respect to any one amino acid sequence of sequence numbers 1 to 71.
2. The Cas protein according to claim 1, wherein the Cas protein has at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to any one amino acid sequence of SEQ ID NOs. 1 to 71.
3. An engineered non-natural type II CRISPR-related (Cas) protein having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to any one amino acid sequence of sequence numbers 1 to 71, except for the first amino acid "M" of the sequence.
4. The Cas protein according to any one of claims 1 to 3, wherein the Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to any one amino acid sequence of SEQ ID NOs. 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, and 68.
5. The Cas protein according to any one of claims 1 to 3, wherein the Cas protein has sequence identity of at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% with respect to any one amino acid sequence of SEQ ID NOs. 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, and 68, excluding the amino acid "M" at position 1 of the sequence.
6. The Cas protein according to any one of claims 1 to 5, wherein the Cas protein is capable of recognizing a protospacer adjacent motif (PAM) having a sequence selected from the group consisting of NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, and NGGHA.
7. (1) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 9, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRRANH; (2) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12, and is capable of recognizing a protospacer adjacent motif (PAM) having sequence NRHACT; (3) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 19, and is capable of recognizing a protospacer adjacent motif (PAM) having a sequence NRAAR; (4) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 21, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NNNCCY; (5) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 22, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NNRYYYYY; (6) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 24, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NGG; (7) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 25, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NNNNNCAA; (8) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 27, and is capable of recognizing a protospacer adjacent motif (PAM) having a sequence NRNACN; (9) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 29, and is capable of recognizing a protospacer adjacent motif (PAM) having a sequence NNGR; (10) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 30, and is capable of recognizing a protospacer adjacent motif (PAM) having a sequence NGGNR; (11) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 31, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NNNCCH; (12) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 36, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRRAAG; (13) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 37, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRHRAC; (14) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 38, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRYART; (15) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 43, and is capable of recognizing a protospacer adjacent motif (PAM) having sequence NRHACC; (16) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 44, and is capable of recognizing a protospacer adjacent motif (PAM) having sequence NRAAR; (17) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 51, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRNVHH; (18) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 56, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence YMACAW; (19) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of Sequence ID No. 59, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NAHAA; (20) The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 60, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NRHAYY; or (21) The Cas protein according to any one of claims 1 to 6, wherein the Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of Sequence ID No. 68, and is capable of recognizing a protospacer adjacent motif (PAM) having the sequence NGGHA.
8. The Cas protein according to any one of claims 1 to 7, wherein the Cas protein is nickase or dead Cas protein.
9. The Cas protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 9, and has a mutation at residue D11 or H859; or has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92% sequence identity with respect to the amino acid sequence of SEQ ID NO: 12, The Cas protein according to claim 8, having at least 95%, at least 98%, at least 99%, or 100% sequence identity and having a mutation at residue D10 or H862; or having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to the amino acid sequence of SEQ ID NO: 31 and having a mutation at residue D12 or H903.
10. With respect to the amino acid sequence of SEQ ID NO: 9, except for the amino acid "M" at position 1 of the sequence, it has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity, and has a mutation at residue D11 or H859; or, with respect to the amino acid sequence of SEQ ID NO: 12, except for the amino acid "M" at position 1 of the sequence, it has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and The Cas protein according to claim 8, having 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity and having a mutation at residue D10 or H862; or, with respect to the amino acid sequence of SEQ ID NO: 31, having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity except for the amino acid "M" at position 1 of the sequence and having a mutation at residue D12 or H903.
11. The Cas protein according to any one of claims 8 to 10, wherein the mutation in residue D11 or H859 of SEQ ID NO: 9 is D11A or H859A; the mutation in residue D10 or H862 of SEQ ID NO: 12 is D10A or H862A; or the mutation in residue D12 or H903 of SEQ ID NO: 31 is D12A or H903A.
12. The Cas protein according to any one of claims 8 to 11, having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with respect to any one amino acid sequence of sequence numbers 869 to 877.
13. The Cas protein according to any one of claims 8 to 12, further comprising a sequence selected from the group consisting of a nuclear localization signal sequence, a nuclear export signal sequence, a cell membrane permeable peptide sequence, an affinity tag sequence, a deaminase sequence, a reverse transcriptase sequence, a recombinase sequence, a methyltransferase sequence, a methylase sequence, an acetylase sequence, an acetyltransferase sequence, a transcription activator sequence, a transcription repression domain sequence, a cryptochrome sequence, a photoinducible / controllable domain sequence, and a chemiinducible / controllable domain sequence.
14. A Cas protein according to any one of claims 1 to 13, comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to any one amino acid sequence of sequence numbers 81 to 151 and 864 to 866.
15. An engineered, non-natural polynucleotide encoding a type II CRISPR-related (Cas) protein according to any one of claims 1 to 14.
16. The polynucleotide according to claim 15, wherein the polynucleotide is a ribonucleotide sequence or a deoxyribonucleotide sequence or an analog thereof; optionally, the polynucleotide is codon-optimized for expression in target cells; preferably, the polynucleotide is mRNA and further comprises a 5' cap sequence and / or a poly A tail sequence.
17. The polynucleotide according to claim 16, wherein the polynucleotide is codon-optimized for expression in eukaryotic cells; optionally, the eukaryotic cells are selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and / or human cells.
18. The polynucleotide according to any one of claims 15 to 17, wherein the polynucleotide has at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 98%, at least 99%, or 100% sequence identity with respect to any one nucleotide sequence of sequence numbers 161 to 231, 241 to 311, 851 to 852, and 861 to 863.
19. a) a type II Cas protein according to any one of claims 1 to 14, or a polynucleotide encoding the same; b) an engineered non-natural CRISPR-Cas system comprising at least one engineered guide RNA, or at least one engineered nucleic acid encoding the same, wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein-binding segment that interacts with the Cas protein, and the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) duplex.
20. The system according to claim 19, wherein the guide RNA further includes a linker sequence that ligates a tracrRNA sequence and a DR sequence to form an sgRNA scaffold.
21. The system according to claim 20, wherein the sgRNA scaffold includes a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 98%, at least 99%, or 100% identity with any one of sequence numbers 431 to 562.
22. The system according to any one of claims 19 to 21, wherein the polynucleotide encoding the Cas protein is operably ligated to a promoter; optionally, the promoter is a constitutive promoter, a tissue-specific promoter, or an inducible promoter.
23. The system according to claim 22, wherein the polynucleotide encoding the Cas protein is operably ligated to a promoter and present in the vector; optionally, the vector is selected from the group consisting of retroviral vectors, lentiviral vectors, phage vectors, adenovirus vectors, adeno-associated virus vectors, herpes simplex virus vectors, and plasmid vectors.
24. An engineered vector comprising a polynucleotide according to any one of claims 15 to 18, wherein the vector is optionally an inducible, conditional, or constitutive expression vector.
25. A vector system comprising one or more polynucleotides according to any one of claims 15 to 18 and one or more polynucleotides encoding a guide RNA, wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein-binding segment that interacts with the Cas protein, and the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) duplex.
26. An engineered non-natural cell comprising a Cas protein according to any one of claims 1 to 14, a polynucleotide according to any one of claims 15 to 18, a CRISPR-Cas system according to any one of claims 19 to 23, a vector according to claim 24, or a vector system according to claim 25.
27. Cells modified using the Cas protein according to any one of claims 1 to 14, the polynucleotide according to any one of claims 15 to 18, the CRISPR-Cas system according to any one of claims 19 to 23, the vector according to claim 24, or the vector system according to claim 25.
28. A kit comprising the Cas protein according to any one of claims 1 to 14, the polynucleotide according to any one of claims 15 to 18, the CRISPR-Cas system according to any one of claims 19 to 23, the vector according to claim 24, the vector system according to claim 25, or the cells according to claim 26 or 27.
29. A pharmaceutical composition comprising a Cas protein according to any one of claims 1 to 14, a polynucleotide according to any one of claims 15 to 18, a CRISPR-Cas system according to any one of claims 19 to 23, a vector according to claim 24, a vector system according to claim 25, or cells according to claim 26 or 27.
30. The pharmaceutical composition according to claim 29, further comprising a delivery system selected from AAV (adeno-associated virus), adenovirus, retrovirus, HSV (herpes simplex virus), gammaretrovirus (Gammaretrovirus), LV (lentivirus), eCIS (extracellular contractile injection system), eVLPs (artificial virus-like particles), VLPs (virus-like particles), liposomes, plasmids, LNPs (lipid nanoparticles), exosomes, microvesicles, nucleic acid nanoassemblies, gene guns, and / or implantable devices.
31. Use of a Cas protein according to any one of claims 1 to 14, a polynucleotide according to any one of claims 15 to 18, a CRISPR-Cas system according to any one of claims 19 to 23, a vector according to claim 24, a vector system according to claim 25, a cell according to claim 26 or 27, a kit according to claim 28, or a pharmaceutical composition according to claim 29 or 30 in a method for treating, preventing, diagnosing, or detecting a disease.
32. A method for modifying or targeting a target DNA locus, comprising delivering to the locus a Cas protein according to any one of claims 1 to 14, a polynucleotide according to any one of claims 15 to 18, a CRISPR-Cas system according to any one of claims 19 to 23, a vector according to claim 24, a vector system according to claim 25, a cell according to claim 26 or 27, a kit according to claim 28, or a pharmaceutical composition according to claim 29 or 30.
33. The method according to claim 32, wherein the modification or targeting of the target DNA locus includes induction of DNA strand breaks, alteration of gene expression of one or more genes, or epigenetic modification of the target DNA locus; optionally, the DNA strand breaks include DNA double-strand breaks or DNA single-strand breaks.
34. The method according to any one of claims 32 or 33, wherein the method is performed in vitro or in vivo.
35. A method for targeting and cleaving double-stranded target DNA, comprising the step of contacting the double-stranded target DNA with a Cas protein according to any one of claims 1 to 14, a polynucleotide according to any one of claims 15 to 18, a CRISPR-Cas system according to any one of claims 19 to 23, or a pharmaceutical composition according to claim 29 or 30.
36. Isolated eukaryotic cells comprising a modified target locus of interest, wherein the target locus of interest is modified by the method described in any one of claims 32 to 35, by using the pharmaceutical composition described in claim 29 or 30, or by using the CRISPR-Cas system described in any one of claims 19 to 23.
37. A system for detecting the presence of a nucleic acid target sequence in an in vitro sample, comprising: a) a Cas protein according to any one of claims 1 to 14; b) at least one guide polynucleotide comprising a guide sequence capable of binding to the target sequence and designed to form a complex with the Cas protein; and c) a nucleic acid-based masking construct comprising a non-target sequence, wherein the Cas protein exhibits associated cleavage activity of RNA and / or ssDNA and cleaves the non-target sequence of the nucleic acid-based masking construct activated by the target sequence.
38. A method for detecting target nucleic acids in a sample, comprising the steps of contacting one or more samples with a) a Cas protein according to any one of claims 1 to 14; b) at least one guide polynucleotide comprising a guide sequence designed to have a predetermined complementarity with a target sequence and to form a complex with the Cas protein; and c) a nucleic acid-based masking construct comprising a non-target sequence, wherein the Cas protein exhibits contingent cleavage activity of RNA and / or ssDNA, cleaving the non-target sequence of the nucleic acid-based masking construct activated by the target sequence, and detecting a signal from the cleavage of the non-target sequence to detect one or more target sequences in the sample.
39. a) a spacer sequence of Sequence ID No. 901; a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of Sequence ID No. 901; or a spacer sequence having at least 99%, at least 98%, at least 97%, at least 96%, at least 95%, at least 94%, at least 93%, at least 92%, at least 91%, or at least 90% identity with the sequence of Sequence ID No. 901; and b) a Cas protein-binding segment; the Cas protein-binding segment comprising a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the guide RNA (gRNA) comprises: a) a spacer sequence of Sequence ID No. 901; a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of Sequence ID No. 901; or a spacer sequence having at least 99%, at least 98%, at least 97%, at least 96%, at least 95%, at least 94%, at least 93%, at least 92%, at least 91%, or at least 90% identity with the sequence of Sequence ID No. 901; and b) a Cas protein-binding segment; wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) duplex.
40. The gRNA according to claim 39, wherein the gRNA is a double-stranded guide RNA or a single-stranded guide RNA.
41. The gRNA according to claim 40, wherein the gRNA has at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 98%, at least 99%, or 100% sequence identity with respect to any one nucleotide sequence of Sequence ID No.
903.
42. The gRNA according to claim 41, wherein at least three nucleotides of the gRNA are modified.
43. A polynucleotide encoding the gRNA according to any one of claims 39 to 42.
44. a) Cas protein or polynucleotide encoding said Cas protein; b) an engineered non-natural CRISPR-Cas system comprising at least one guide RNA (gRNA) as described in any one of claims 39 to 42, or at least one artificial nucleic acid encoding the guide RNA, wherein the gRNA further comprises a Cas protein-binding segment that interacts with the Cas protein, and the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) duplex.
45. The CRISPR-Cas system according to claim 44, wherein the gRNA comprises a) the sequence of Sequence ID No. 903, or b) a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 98%, at least 99%, or 100% identity with the sequence of Sequence ID No.
903.
46. The CRISPR-Cas system according to claim 44 or 45, wherein the Cas protein consists of a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 98%, at least 99%, or 100% identity with any one of the sequences of SEQ ID NO: 12, except for the amino acid "M" at position 1 of the sequence.
47. A modified vector comprising the polynucleotide of claim 43, wherein the vector is optionally an inducible, conditional, or constitutive expression vector.
48. A vector system comprising one or more polynucleotides according to claim 43.
49. A pharmaceutical composition comprising a gRNA according to any one of claims 39 to 42, a polynucleotide according to claim 43, a CRISPR-Cas system according to any one of claims 44 to 46, a vector according to claim 47, or a vector system according to claim 48.
50. A method for treating, preventing, or diagnosing a disease associated with the RHO locus in a subject as needed, comprising the steps of: a) contacting target cells derived from the subject with the gRNA described in any one of claims 39 to 42; b) contacting target cells derived from the subject with the polynucleotide described in claim 43; c) contacting target cells derived from the subject with the CRISPR-Cas system described in any one of claims 44 to 46; or d) contacting target cells derived from the subject with the pharmaceutical composition described in claim 49.