Type II Cas protein, CRISPR-Cas system and application thereof
By modifying the type II CRISPR-Cas system, multiple PAM sequences can be identified and bound to guide RNA, overcoming the shortcomings of existing systems in targeting DNA sequences, achieving more efficient gene editing results, and improving the flexibility and precision of gene editing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-03-24
AI Technical Summary
Existing CRISPR-Cas systems struggle to effectively identify multiple PAM sequences during gene targeting and editing, resulting in insufficient flexibility and precision in targeting DNA sequences and limiting their potential application in gene editing.
An engineered type II CRISPR-associated Cas protein was developed that can recognize a variety of PAM sequences and bind to specific guide RNA to form a double-stranded RNA dimer, thereby achieving efficient targeted editing of target DNA.
It has expanded the flexibility and precision of the CRISPR-Cas system in DNA sequence editing, improved the versatility and effectiveness of gene editing, and advanced research and medical applications in the field of gene editing.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to type II Cas protein, the CRISPR-Cas system, and their uses. In particular, the type II Cas protein and the CRISPR-Cas system are used for gene targeting or gene editing. This application claims priority to PCT applications PCT / CN2023 / 113355, PCT / CN2023 / 116757, PCT / CN2023 / 136724, PCT / CN2024 / 091203, PCT / CN2024 / 091198, and PCT / CN2024 / 091211. The entire contents of the above applications are hereby incorporated by reference. Background Technology
[0002] Targeted genome editing or modification is rapidly becoming an important tool in both basic and applied research. Clustered regularly spaced short palindromic repeats and CRISPR-associated protein (CRISPR-Cas) systems show great promise due to their ease of specific targeting via engineered guide RNAs. Recent advances in genome sequencing technologies and analytical methods have significantly accelerated the ability to catalog and locate genetic factors associated with a wide range of biological functions and diseases. Precise genome targeting technologies are needed to systematically reverse-engineer causal genetic variations by allowing selective interference with individual genetic elements, and to advance synthetic biology, biotechnology, and medical applications.
[0003] Existing technologies have explored various CRISPR-Cas systems, and different CRISPR-Cas systems exhibit different characteristics. For example, the CRISPR-Cas9 system, which belongs to the second class of CRISPR-Cas systems, has been used for genome editing and has shown great promise in biomedical research. Summary of the Invention
[0004] This invention provides an engineered, non-naturally occurring type II CRISPR-associated (Cas) protein or a variant thereof, wherein the protein has at least 70% sequence identity with any amino acid sequence in SEQ ID NO: 1-71. In some embodiments, the sequence identity of the Cas protein with any amino acid sequence in SEQ ID NO: 1-71 is at least 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%.
[0005] The present invention also provides an engineered, non-naturally occurring type II CRISPR-related (Cas) protein, wherein, except for the amino acid "M" at position 1 in the sequence, the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with any amino acid sequence in SEQ ID NO: 1-71.
[0006] In some embodiments, the Cas protein further includes an effector domain (or functional domain). Such an effector domain may have one or more types of enzymatic activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain includes a nuclease, nickase, deaminase, reverse transcriptase, recombinase, methyltransferase, methyltransferase, acetyltransferase, transcription activator, transcription repressor domain, cryptochrome, photoinducible / controllable domain, or chemically inducible / controllable domain.
[0007] In some embodiments, the Cas protein further comprises one or more nuclear localization signal sequences, nuclear export signal sequences, cell-penetrating peptide sequences, and affinity tags. Type II Cas proteins comprise one or more nuclear localization signals (NLS). The NLS may be located at the ends or other parts of the peptide chain. NLS located at both ends or other parts of the Cas9 amino acid sequence may be the same or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are the same. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains one NLS, and the C-terminus of the Cas9 amino acid sequence contains one NLS. The NLS is fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence, respectively. The NLS may be an SV40 (monkey virus 40) NLS, a c-Myc NLS, or other suitable monomeric NLS. The NLS may be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, the Cas protein is purified by affinity chromatography using an affinity tag (such as GST, FLAG, or a six-histidine sequence). In some embodiments, the amino acid sequence of the C-terminal NLS is given in SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is given in SEQ ID NO: 883. The NLS and FLAG sequences may also be selected from other available sequences and different combinations.
[0008] In some embodiments, the Cas protein comprises an amino acid sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any of the amino acid sequences in SEQ ID NO: 1-71.
[0009] In some preferred embodiments, the Cas protein is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any of the amino acid sequences in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, 68. In some other preferred embodiments, except for amino acid "M" at position 1 in the sequence, the Cas protein is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any of the amino acid sequences in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, 68.
[0010] A key element in the CRISPR-Cas system's modification process is the adjacent short palindromic repeat (PAM) motif, a short DNA sequence immediately adjacent to the target DNA sequence. The PAM sequence is crucial for the binding and cleavage activity of the Cas protein, ensuring that only the predetermined DNA sequence is targeted for editing. The Cas protein's ability to recognize these specific PAM sequences is attributed to their unique structural features and DNA-binding interactions. Each PAM sequence provides a distinct motif that the Cas protein can recognize, ensuring accurate and efficient targeting. For example, an NRRANH sequence may provide a specific nucleotide arrangement that the Cas protein can bind to with high affinity. This diversity of PAM sequences expands the potential applications of the CRISPR-Cas system. It enables researchers and scientists to target a wider range of DNA sequences for editing, increasing the versatility and effectiveness of this powerful gene-editing tool. Furthermore, understanding these PAM sequences can help develop more advanced Cas proteins with higher precision targeting capabilities, further advancing the field of gene editing and its potential benefits for research, medicine, and biotechnology. The Cas protein disclosed in this invention possesses the unique ability to recognize a variety of PAM sequences. These sequences include NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. The specific recognition of these sequences by these Cas proteins allows for greater flexibility in selecting target DNA sequences for editing.
[0011] In some embodiments, the Cas protein disclosed in this invention is capable of recognizing at least one protospacer adjacent motif (PAM) having or containing the following sequences: NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 9, and is capable of recognizing protospacer adjacent motifs (PAMs) with the NRRANH sequence; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 100% identity with any amino acid sequence in SEQ ID NO: 9, and is capable of recognizing protospacer adjacent motifs (PAMs) with the NRRANH sequence; The identity of any amino acid sequence in 12 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRHACT; in some embodiments, the Cas protein is similar to SEQ ID NO: The identity of any amino acid sequence in 19 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with sequence NRAAR; in some embodiments, the Cas protein is related to SEQ ID NO: The identity of any amino acid sequence in 21 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and is capable of recognizing a parsing adjacent motif (PAM) with the sequence NNNCCY;In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 22, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NNRYYYY; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 100% identity with any amino acid sequence in SEQ ID NO: 22, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NNRYYYY; The identity of any amino acid sequence in 24 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NGG; in some embodiments, the Cas protein is related to SEQ ID NO: The identity of any amino acid sequence in 25 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NNNCAA; in some embodiments, the Cas protein is related to SEQ ID NO: The identity of any amino acid sequence in 27 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRNACN; in some embodiments, the Cas protein is associated with SEQ ID NO: The identity of any amino acid sequence in 29 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and is capable of recognizing the original spacer adjacent motif (PAM) with the sequence NNGR;In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 30, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NGGNR; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 100% identity with any amino acid sequence in SEQ ID NO: 30, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NGGNR; The identity of any amino acid sequence in SEQ ID NO: 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NNNCCH; in some embodiments, the identity of the Cas protein with any amino acid sequence in SEQ ID NO: 36 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 7 ... %, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRRAAG; in some embodiments, the Cas protein has an identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% with any amino acid sequence in SEQ ID NO: 37, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRHRAC ... The identity of any amino acid sequence in 38 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRYART;In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 43, and is capable of recognizing protospacer adjacent motifs (PAMs) with NRHACC sequences; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 100% identity with any amino acid sequence in SEQ ID NO: 43, and is capable of recognizing protospacer adjacent motifs (PAMs) with NRHACC sequences; The identity of any amino acid sequence in 44 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with sequence NRAAR; in some embodiments, the Cas protein is related to SEQ ID NO: The identity of any amino acid sequence in 51 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRNVHH; in some embodiments, the Cas protein is similar to SEQ ID NO: The identity of any amino acid sequence in 56 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence YMACAW; in some embodiments, the Cas protein is related to SEQ ID NO: The identity of any amino acid sequence in 59 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NAHAA;In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 60, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRHAYY; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 100% identity with any amino acid sequence in SEQ ID NO: 60, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRHAYY; The identity of any amino acid sequence in 68 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing proto-spacer adjacent motifs (PAMs) with the NGGHA sequence.
[0012] In some embodiments, the Cas protein is a cleavage enzyme-active or inactive Cas protein. The DNA cleavage domain of the active Cas protein of this invention comprises two subdomains: an HNH nuclease subdomain and a RuvC subdomain. Mutations within these subdomains can inhibit the nuclease activity of the Cas protein. In some embodiments, the Cas protein exhibits at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with the SEQ ID NO: 9, and includes mutations at residues D11 or H859; or with SEQ ID NO: The amino acid sequence identity of SEQ ID NO: 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes a mutation at residue D10 or H862; or the amino acid sequence identity with SEQ ID NO: 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes a mutation at residue D12 or H903.In some embodiments, except for amino acid "M" at position 1 in the sequence, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 9, and includes mutations at residues D11 or H859; or except for amino acid "M" at position 1 in the sequence, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 100% amino acid sequence identity with SEQ ID NO: 9, and includes mutations at residues D11 or H859; or except for amino acid "M" at position 1 in the sequence, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 9, and includes mutations at residues D11 or H859; or except for amino acid "M" at position 1 in the sequence, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77 The amino acid sequence identity of 12 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes mutations at residue D10 or H862; or, except for amino acid "M" at position 1 in the sequence, is identical to SEQ ID NO: The amino acid sequence identity of 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes a mutation at residue D12 or H903. In some embodiments, the mutation at residue D11 or H859 in SEQ ID NO: 9 is D11A or H859A; the mutation at residue D10 or H862 in SEQ ID NO: 12 is D10A or H862A; and the mutation at residue D12 or H903 in SEQ ID NO: 31 is D12A or H903A. In some embodiments, the Cas protein is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any of the amino acid sequences in SEQ ID NO: 869-877.
[0013] The present invention also provides an engineered, non-naturally occurring polynucleotide encoding a type II CRISPR-associated (Cas) protein.
[0014] In some embodiments, the polynucleotide encoding the Cas protein is operatively linked to a promoter and presented in a vector; alternatively, the vector is selected from the group consisting of: retroviral vectors, lentiviral vectors, phage vectors, adenovirus vectors, adeno-associated virus vectors, herpes simplex virus vectors, and plasmid vectors.
[0015] In some embodiments, the polynucleotide is a ribonucleotide sequence or a deoxyribonucleotide sequence, or an analogue thereof; optionally, the polynucleotide is codon-optimized for expression in the cells of interest; in some embodiments, the polynucleotide is codon-optimized for expression in eukaryotic cells. In some embodiments, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cells are mammalian cells, preferably human cells.
[0016] In some embodiments, the polynucleotide is mRNA and further comprises a 5' cap sequence and / or a poly-A tail sequence. In some embodiments of the invention, the mRNA used may be modified to enhance its functional properties and stability. Specifically, in some embodiments, the modification process involves replacing uridine (represented by the letter "U") with N1-methylpseuuridine or pseudouridine. This substitution aims to improve the mRNA's resistance to ribonuclease degradation, potentially increasing its intracellular half-life and translation efficiency. Incorporating N1-methylpseuuridine or pseudouridine into the mRNA structure can also positively influence the immunogenicity profile, as these modifications have been shown to reduce the immunogenicity of the mRNA molecule compared to unmodified mRNA molecules. This is particularly critical for the development of mRNA-based therapeutics and vaccines, where minimizing adverse immune responses is essential.
[0017] In some embodiments, the polynucleotides of the present invention are codon-optimized for expression in eukaryotic cells; optionally, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, single-celled eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and / or human cells.
[0018] In some embodiments, the polynucleotide shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the nucleotide sequences in SEQ ID NO: 161-231, 241-311, 851-852, 861-863.
[0019] The present invention also provides an engineered, non-naturally occurring CRISPR-Cas system comprising: a) a type II Cas protein or a polynucleotide encoding a Cas protein as described herein; b) at least one engineered guide RNA or at least one engineered nucleic acid encoding a guide RNA, wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein binding segment that interacts with the Cas protein, wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer.
[0020] In some embodiments, the guide RNA is a dual guide RNA. In some embodiments, the guide RNA is a single guide RNA. In these embodiments, the guide RNA further includes a linker sequence connecting the tracrRNA sequence and the DR sequence. In some typical embodiments, the linker sequence comprises a short GAAA sequence. In some embodiments, the linker sequence is an artificial loop. In some embodiments, the sgRNA comprises the following sequence: a) a spacer sequence capable of hybridizing with the target nucleic acid sequence to be manipulated; b) the DR sequence; c) the linker sequence; and d) the tracrRNA sequence. The spacer sequence, DR sequence, linker sequence, and tracrRNA sequence are tandemly arranged in a 5' to 3' direction or a 3' to 5' direction; in some embodiments, the sgRNA backbone includes a sequence that has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the sequences in SEQ ID NO:431-562.
[0021] In some embodiments, the sgRNA comprises a spacer sequence (such as any of the sequences in SEQ ID NO: 571-835, 885, 901) and a backbone sequence, wherein the spacer sequence is located at the 5' end of the backbone sequence (such as the sequence in SEQ ID NO: 903). In some embodiments, the sgRNA comprises a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the sequences in SEQ ID NO: 903.
[0022] In some embodiments, the spacer sequence hybridizes with one or more nucleic acids in a prokaryotic or eukaryotic cell. In some embodiments, the eukaryotic cell is selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the eukaryotic cell includes mammalian cells. In some embodiments, the mammalian cell includes human cells. In some embodiments, the eukaryotic cell includes plant cells.
[0023] In some embodiments, the system further includes a donor template nucleic acid.
[0024] In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid. In some embodiments, the donor template nucleic acid is a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear. In some embodiments, the donor template nucleic acid is spherical (e.g., a plasmid). In some embodiments, the donor template nucleic acid is an exogenous nucleic acid molecule. In some embodiments, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, the donor template nucleic acid is DNA or RNA or a DNA-RNA hybrid.
[0025] The present invention also provides an engineered vector comprising the polynucleotides described in this disclosure.
[0026] In some embodiments, the vector is an expression vector. In some embodiments, the vector is an inducible, conditional, or constitutive expression vector. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA are located on the same vector or on different vectors.
[0027] The present invention also provides a vector system comprising one or more polynucleotides described herein and one or more polynucleotides encoding guide RNA; wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein-binding segment that interacts with the Cas protein, wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA are located on the same vector or on different vectors.
[0028] In some embodiments, the vector, such as a plasmid or viral vector, is delivered to the target tissue via, for example, intramuscular injection, intravenous administration, transdermal administration, intranasal administration, oral administration, or mucosal administration. Such administration can be a single dose or multiple doses. Those skilled in the art will understand that the actual dose may vary considerably due to a variety of factors, such as the choice of vector, target cells, target organism, target tissue, general condition of the subject, the required degree of transformation / modification, route of administration, manner of administration, and type of transformation / modification required.
[0029] The present invention also provides an engineered, non-naturally occurring cell comprising: the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, or the vector system described herein.
[0030] The present invention also provides a cell modified by utilizing the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, or the vector system described herein.
[0031] In some embodiments, the cell is a eukaryotic or prokaryotic cell. In some embodiments, the eukaryotic cell is selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cell is a mammalian cell, a human cell, or a plant cell.
[0032] In some embodiments, the cell is a vertebrate, mammal, rodent, goat, pig, bird, chicken, turkey, cow, horse, sheep, fish, primate, or human cell. In some embodiments, the cell is a mammalian cell. In one embodiment, the cell is a human cell. In some embodiments, the cell is a somatic cell, germ cell, or fetal cell. In some embodiments, the cell is a zygote, blastocyst, embryonic cell, stem cell, mitotically capable cell, or meiotically capable cell. In some embodiments, the cell is not part of a human embryo. In some embodiments, the cell is a somatic cell. In one embodiment, the cell is a T cell, CD8+ T cell, CD8+ naive T cell, central memory T cell, effector memory T cell, or CD4+ T cell. T cells, stem cell memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricted progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythrocyte progenitor cells, megakaryocytic erythrocyte progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, aqueous reticulum cells, cochlear hair cells, outer hair cells, inner hair cells, lung epithelial cells, bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, neurons The cells may include neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells (e.g., progenitor B cells, pre-B cells, original B cells, memory B cells, plasma cells), gastrointestinal epithelial cells, biliary epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, preadipocytes, pancreatic islet cells (e.g., β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0033] In some embodiments, this disclosure provides an isolated eukaryotic cell comprising a modified target site, wherein the target site has been modified according to the method described in the invention or using the composition or system described in the invention.
[0034] In some embodiments, the cells are eukaryotic or prokaryotic cells. In some embodiments, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cells are mammalian cells, human cells, or plant cells.
[0035] In some embodiments, the cell is a vertebrate, mammal, rodent, goat, pig, bird, chicken, turkey, cow, horse, sheep, fish, primate, or human cell. In some embodiments, the cell is a mammalian cell. In one embodiment, the cell is a human cell. In some embodiments, the cell is a somatic cell, germ cell, or fetal cell. In some embodiments, the cell is a zygote, blastocyst, embryonic cell, stem cell, mitotically capable cell, or meiotically capable cell. In some embodiments, the cell is not part of a human embryo. In some embodiments, the cell is a somatic cell. In one embodiment, the cell is a T cell, CD8+ T cell, CD8+ naive T cell, central memory T cell, effector memory T cell, or CD4+ T cell. T cells, stem cell memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricted progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythrocyte progenitor cells, megakaryocytic erythrocyte progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, aqueous reticulum cells, cochlear hair cells, outer hair cells, inner hair cells, lung epithelial cells, bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, nerve cells. The cells may include: neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells (e.g., progenitor B cells, pre-B cells, memory B cells), plasma cells, gastrointestinal epithelial cells, biliary epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, preadipocytes, pancreatic islet cells (e.g., β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0036] The present invention also provides a kit comprising: the Cas protein described in this disclosure, the polynucleotide described in this disclosure, the CRISPR-Cas system described in this disclosure, the vector described in this disclosure, the vector system described in this disclosure, or the cell described in this disclosure.
[0037] The kits described in this disclosure may include one or more containers containing the components necessary to perform the methods of this disclosure, and may include instructions for use. Any underlined kit may also include auxiliary components necessary to perform the editing methods. Each component in the kit may be provided in liquid form (e.g., dissolved in solution) or solid form (e.g., lyophilized powder), where applicable. In certain embodiments, some components may need to be reformulated or otherwise treated (e.g., activated state) after the addition of suitable solvents or other substances (such as water or buffers), which may or may not be provided with the kit. In some embodiments, the kit may further include other suitable excipients, such as buffers or reagents, to facilitate the application of the kit. The kit can be used in a variety of applications, such as medical applications, including therapeutic and diagnostic, research, etc. Therefore, the type II Cas nuclease and kit of the present invention can be used to prepare pharmaceuticals or reagents for therapeutic and / or research purposes.
[0038] Cas proteins, the CRISPR-Cas system, and the polynucleotides described herein can be delivered via various delivery systems, such as vectors (e.g., plasmids), viral delivery vectors (e.g., adeno-associated virus (AAV), lentivirus, adenovirus, and other viral vectors), or methods (e.g., ribo-electroporation or electroporation of ribonucleoprotein complexes composed of type V-VI effectors and their corresponding guide RNAs). Proteins and one or more guide RNAs can be packaged into one or more vectors, such as plasmids or viral vectors. For bacterial applications, nucleic acids encoding any component of the CRISPR system described herein can be delivered into bacteria using bacteriophages. Exemplary bacteriophages include, but are not limited to, T4 phage, Mu, λ phage, T5 phage, T7 phage, T3 phage, Φ29, M13, MS2, Qβ, and ΦX174.
[0039] The present invention also provides a pharmaceutical composition comprising the following components: the Cas protein described in this disclosure, the polynucleotide described in this disclosure, the CRISPR-Cas system described in this disclosure, the vector described in this disclosure, the vector system described in this disclosure, or the cell described in this disclosure.
[0040] As described in this disclosure, a "pharmaceutical composition" refers to a formulation intended for pharmaceutical use. In some embodiments, the pharmaceutical composition further includes an acceptable pharmaceutical excipient. In some embodiments, the pharmaceutical composition may contain other therapeutic agents. In some embodiments, the pharmaceutical composition is prepared according to standard procedures and can be administered to a subject, such as a human patient, via intravenous, intramuscular, intradermal, intra-articular, intralesional, intraperitoneal, intracardiac, intracerebrospinal fluid, intraventricular, epidural, topical, subconjunctival, periocular, intraocular, vitreous, posterior sclera, penetrating sclera, suprascleral, subretinal, retroretinal, fundus, intranasal inhalation, pressurized inhalation, oral, subcutaneous, or topical routes. For example, compositions for injection may be provided as a sterile isotonic aqueous solution. If necessary, the pharmaceutical composition may also contain a solvent and a local anesthetic, such as lidocaine, to minimize discomfort at the injection site. Typically, components may be provided individually or as a mixture of unit doses, for example as a lyophilized powder or anhydrous concentrated solution, in a sealed container indicating the amount of active ingredient. If the pharmaceutical composition is intended for infusion, it can be combined with an infusion bottle containing sterile pharmaceutical-grade water or physiological saline. If the pharmaceutical composition is intended for injection, sterile water for injection or physiological saline may be included to mix the components before administration. Furthermore, wetting agents, colorants, release agents, coating agents, sweeteners, flavoring agents, preservatives, and antioxidants may also be incorporated into the formulation as needed.
[0041] In some embodiments, the pharmaceutical composition further includes a delivery system selected from: AAV (adeno-associated virus), adenovirus, retrovirus, HSV (herpes simplex virus), gamma retrovirus, lentivirus, eCIS (extracellular contractile injection system), eVLPs (engineered virus-like particles), VLPs (virus-like particles), liposomes, plasmids, LNPs (lipid nanoparticles), exosomes, microvesicles, nucleic acid nanoassemblies, gene guns, and / or implantable devices.
[0042] In some embodiments, delivery is carried out via adeno-associated virus (AAV), such as AAV2, AAV8, or AAV9, which may contain at least 1 × 10^5 adenovirus or adeno-associated virus particles (also known as particle units, pu) in a single dose.
[0043] In some embodiments, delivery is performed via a recombinant adeno-associated virus (rAAV) vector. For example, in some embodiments, a modified AAV vector may be used for delivery. The modified AAV vector may be based on one or more capsid types, including AAV1, AAV2, AAV5, AAV6, AAV8, AAV8.2, AAV9, AAV rh10, modified AAV vectors (e.g., modified AAV2, modified AAV3, modified AAV6), and pseudotyped AAVs (e.g., AAV2 / 8, AAV2 / 5, and AAV2 / 6).
[0044] In some embodiments, delivery is performed via plasmids. The dosage may be sufficient to elicit a response. In some embodiments, the appropriate amount of plasmid DNA in the plasmid composition may range from about 0.1 mg to about 2 mg. Plasmids typically comprise (i) a promoter; (ii) a sequence encoding a CRISPR enzyme targeting a nucleic acid, operatively linked to the promoter; (iii) a selectivity marker; (iv) an origin of replication; and (v) a transcription terminator, downstream of (ii) and operatively linked to (ii). Plasmids may also encode RNA components of the CRISPR-Cas system, but one or more of these may be encoded on different vectors. The frequency of administration is determined by a medical or veterinary practitioner (e.g., a physician, veterinarian) or skilled technician.
[0045] The present invention also provides methods for treating, preventing, diagnosing or detecting diseases using the Cas proteins, polynucleotides, CRISPR-Cas systems, vectors, vector systems, cells, kits or pharmaceutical compositions described in this disclosure.
[0046] The present invention also provides a method for modifying or targeting a target DNA site, the method comprising delivering a Cas protein, polynucleotide, CRISPR-Cas system, vector, vector system, kit or pharmaceutical composition described herein to the site.
[0047] In some embodiments, the disclosure also provides a method for targeting and cleaving target DNA, the method comprising: contacting the target DNA with a Cas protein, polynucleotide, CRISPR-Cas system, vector, vector system, kit, or pharmaceutical composition described herein.
[0048] In some embodiments, modifying or targeting a target site includes inducing DNA strand breaks. In some embodiments, modifying or targeting a target site includes inducing DNA double-strand breaks or DNA single-strand breaks. In some embodiments, modifying or targeting a target site includes altering the gene expression of one or more genes. In some embodiments, modifying or targeting a target site includes epigenetic modification of a target DNA site. In some embodiments, the method is a method of modifying a cell, cell line, or organism by manipulating one or more target sequences at sites of interest in the genome.
[0049] In some embodiments, cleaving the target DNA or target sequence results in the formation of an insertion or deletion (indel) or the insertion of a nucleotide sequence. In some embodiments, cleaving the target DNA or target nucleotide includes cutting the target DNA or target sequence at two sites, resulting in deletion or inversion of the sequence between the two sites. In some embodiments, the target DNA is double-stranded DNA or single-stranded DNA, or a DNA-RNA hybrid.
[0050] In some embodiments, modifying or targeting a target site includes inducing DNA strand breaks, altering the gene expression of one or more genes, or epigenetic modification of the target DNA site; optionally, DNA strand breaks include DNA double-strand breaks or DNA single-strand breaks.
[0051] In some embodiments, the method is performed in vitro or in vivo.
[0052] The present invention also provides an isolated eukaryotic cell comprising a modified target site, wherein the target site has been modified according to the methods described in this disclosure, or using the systems described in this disclosure, or using the Cas protein described in this disclosure, or using the polynucleotide described in this disclosure, or using the CRISPR-Cas system described in this disclosure, or using the vector described in this disclosure, or using the vector system described in this disclosure, or using the kit described in this disclosure, or using the pharmaceutical composition described in this disclosure.
[0053] The present invention also provides a system for detecting the presence of a nucleic acid target sequence in an in vitro sample, comprising: a) the Cas protein described in this disclosure; b) at least one guide polynucleotide comprising a guide sequence capable of binding to the target sequence and designed to form a complex with the Cas protein; and c) a nucleic acid-based masking structure comprising a non-target sequence, wherein the Cas protein exhibits collateral cleavage activity against RNA and / or ssDNA and cleaves the non-target sequence in the nucleic acid-based masking structure activated by the target sequence.
[0054] The present invention also provides a method for detecting target nucleic acids in a sample, comprising: contacting one or more samples with a) the Cas protein described in this disclosure; b) at least one guide polynucleotide comprising a guide sequence designed to be complementary to a target sequence and designed to form a complex with the Cas protein; and c) a nucleic acid-based masking structure comprising a non-target sequence, wherein the Cas protein exhibits incidental cleavage activity against RNA and / or ssDNA and cleaves the non-target sequence in the nucleic acid-based masking structure activated by the target sequence; and detecting a signal of non-target sequence cleavage to detect one or more target sequences in the sample.
[0055] The present invention also provides a guide RNA (gRNA) comprising: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA further comprises a Cas protein-binding segment; wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence, which hybridize to form a double-stranded RNA (dsRNA) dimer. In some embodiments, the gRNA is a dual guide RNA. In some embodiments, the gRNA is a single guide RNA. In some embodiments, the gRNA is modified. In some embodiments, at least three nucleotides of the gRNA are modified. In some embodiments, the gRNA comprises a 5' end modification that contains at least two phosphothioester (PS) bonds in the first seven nucleotides of the 5' end. In some embodiments, the gRNA includes a 3' end modification that contains at least two phosphothioester (PS) bonds in the first seven nucleotides of the 3' end. In some embodiments, the gRNA exhibits at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any nucleotide sequence of SEQ ID NO: 903.
[0056] The present invention also provides a polynucleotide encoding the gRNA described in this document; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises a sequence that has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the sequences in SEQ ID NO: 903.
[0057] The present invention also provides an engineered, non-naturally occurring CRISPR-Cas system comprising: a) a Cas protein or a polynucleotide encoding a Cas protein; b) at least one gRNA or at least one engineered nucleic acid encoding a gRNA as described in this disclosure, wherein the gRNA further comprises a Cas protein binding segment that interacts with the Cas protein; wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence, the sequences hybridizing to form a double-stranded RNA (dsRNA) dimer; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 903.
[0058] In some embodiments, except for the amino acid "M" at position 1 of the sequence, the Cas protein comprises a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to any of the sequences in SEQ ID NO: 12; or the Cas protein comprises a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 98%, 99%, or 100% identity to any of the sequences in SEQ ID NO: 12; 12. Any sequence identity is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0059] In some embodiments, the Cas protein further includes an effector domain (or functional domain). Such an effector domain may have one or more types of enzymatic activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain includes nuclease, nickase, deaminase, reverse transcriptase, recombinase, methyltransferase, methyltransferase, acetyltransferase, transcription activator, transcription repressor domain, cryptochrome, photoinducible / controllable domain, or chemically inducible / controllable domain.
[0060] In some embodiments, the Cas protein further comprises one or more nuclear localization signal sequences, nuclear export signal sequences, cell-penetrating peptide sequences, and affinity tags. Type II Cas proteins comprise one or more nuclear localization signals (NLS). The NLS may be located at the ends or other parts of the peptide chain. NLS located at both ends or other parts of the Cas9 amino acid sequence may be the same or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are the same. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains one NLS, and the C-terminus of the Cas9 amino acid sequence contains one NLS. The amino acid sequences of the NLS are fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence, respectively. The NLS may be an SV40 (monkey virus 40) NLS, a c-Myc NLS, or other suitable monomeric NLS. The NLS may be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, the Cas protein is purified by affinity chromatography using an affinity tag (such as GST, FLAG, or a six-histidine sequence). In some embodiments, the amino acid sequence of the C-terminal NLS is selected from SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is selected from SEQ ID NO: 883. Other available sequences and different combinations of the NLS and FLAG sequences may also be selected.
[0061] In some embodiments, the Cas protein comprises an amino acid sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 12.
[0062] In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence NRHACT, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRHACT.
[0063] In some embodiments, the Cas protein is a cleaving enzyme or an inactivated Cas protein. The DNA cleavage domain of the active Cas protein of the present invention comprises two subdomains, namely the HNH nuclease subdomain and the RuvC subdomain. Mutations within these subdomains can inhibit the nuclease activity of the Cas protein. In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 12, and includes mutations at residues D10 or H862; in some embodiments, except for the amino acid "M" at position 1 in the sequence, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 78%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 12. The amino acid sequence identity of 12 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes mutations at residues D10 or H862.
[0064] In some embodiments, the mutation at residue D10 or H862 in SEQ ID NO: 12 is D10A or H862A; in some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 872-874.
[0065] The present invention also provides an engineered vector comprising a polynucleotide encoding gRNA as described in this disclosure; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 903.
[0066] In some embodiments, the vector is optionally an inducible, conditional, or constitutive expression vector.
[0067] The present invention also provides a vector system comprising one or more polynucleotides encoding the gRNA described herein; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 871 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 903.
[0068] The present invention also provides a pharmaceutical composition comprising the gRNA described herein; a polynucleotide encoding such gRNA; a CRISPR-Cas system comprising such gRNA or polynucleotide; a vector comprising a coding sequence of such gRNA; and a vector system comprising a coding sequence of such gRNA; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 903.
[0069] The present invention also provides a method for treating, preventing, or diagnosing diseases associated with the RHO gene locus, comprising: a) contacting target cells with the gRNA described in this disclosure; b) contacting target cells with a polynucleotide encoding such gRNA; c) contacting target cells with a CRISPR-Cas system comprising such gRNA or polynucleotide; or d) contacting target cells with a pharmaceutical composition described in this disclosure; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 903.
[0070] The present invention also provides a method for treating, preventing, or diagnosing diseases related to a gene locus, comprising administering a) the gRNA described in this disclosure; b) administering a target cell with a polynucleotide encoding such gRNA; c) administering a target cell with a CRISPR-Cas system containing such gRNA or polynucleotide; or d) administering a target cell with a pharmaceutical composition described in this disclosure; wherein the guide RNA (gRNA) comprises: i) a spacer sequence of SEQ ID NO: 901; ii) a spacer sequence comprising at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or iii) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 903.
[0071] The present invention also provides a composition comprising: (i) a Cas protein, wherein: a. the Cas protein contains a sequence that is at least 90% identical to SEQ ID NO: 12 or 92; and / or b. the Cas protein contains a sequence that is at least 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 12 or 92; and / or (ii) an sgRNA or a vector encoding sgRNA, wherein the sgRNA contains the sgRNA sequence of SEQ ID NO: 903.
[0072] The present invention also provides a method for modifying RHO gene sites, comprising delivering a composition to a cell, wherein the composition comprises: a. a guide RNA comprising a guide sequence of SEQ ID NO: 901; b. a guide RNA comprising at least 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c. a guide RNA comprising a guide sequence that is at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identical to the SEQ ID NO: 901 sequence.
[0073] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a. a guide RNA comprising a guide sequence of SEQ ID NO: 901; b. a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c. a guide RNA comprising a guide sequence that is at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identical to the SEQ ID NO: 901 sequence.
[0074] The present invention also provides a method for modifying RHO gene sites, comprising delivering a composition to cells, wherein the composition comprises: a. an sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b. an sgRNA comprising an sgRNA sequence having at least 90% identity with the sequence of SEQ ID NO: 903; or c. an sgRNA comprising an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 903.
[0075] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a. an sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b. an sgRNA comprising an sgRNA sequence having at least 90% identity with the sequence of SEQ ID NO: 903; or c. an sgRNA comprising an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 903.
[0076] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a. a guide RNA comprising a spacer sequence of SEQ ID NO: 901; b. a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c. a guide RNA comprising a guide sequence that is at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identical to the SEQ ID NO: 901 sequence.
[0077] The present invention also provides a method for modifying RHO gene sites, comprising delivering a composition to a cell, wherein the composition comprises: (i) a Cas protein, wherein: a. the Cas protein contains a sequence that is at least 90% identical to SEQ ID NO: 12 or 92; and / or b. an RNA-guided DNA binder that contains a sequence that is at least 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 12 or 92; and / or (ii) a guide RNA or a vector encoding the guide RNA, wherein the guide RNA contains a spacer sequence of SEQ ID NO: 901.
[0078] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: (i) an RNA-guided DNA binder, wherein: a. the RNA-guided DNA binder contains a sequence that is at least 90% identical to SEQ ID NO:12 or 92; and / or b. the RNA-guided DNA binder contains a sequence that is at least 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:12 or 92; and / or (ii) an sgRNA or a vector encoding sgRNA, wherein the sgRNA contains a sequence of SEQ ID NO:903.
[0079] From the following detailed description, those skilled in the art will readily recognize the advantages of the embodiments, other embodiments, objects, features, and exemplary models described herein. Attached Figure Description
[0080] The features and advantages of this disclosure can be understood by referring to the following detailed description, which describes in detail exemplary embodiments that can utilize the principles of this disclosure, and is accompanied by corresponding drawings: Figure 1 The domain arrangement of the GEBx type II Cas protein was shown; Figures 2A to 2E The PAM preference of Cas9 in the HEK293 cell line was shown; Figure 3 ( Figure 3 The results show the indel levels of GEBx0305 against 16 target sites and GGAAAA PAM in the HEK293T cell line (n=3). Figure 4 ( Figure 4 The results show the indel levels of GEBx0308 against 19 target sites and GGTACT PAM in the HEK293T cell line (n=3). Figure 5 ( Figure 5 The results show the indel levels of GEBx0308 against CATACTPAM at 15 target sites in the HEK293T cell line (n=3). Figure 6 ( Figure 6 The results show the indel levels of GEBx0328 against NGGCCTPAM at 20 target sites in the HEK293T cell line (n=3). Figure 7 The indel levels of GEBx0361 against 11 target sites and GGTACC-PAM were shown in the HEK293T cell line (n=2). Figure 8 The indel levels of GEBx0361 against 8 target sites and TGTACC-PAM in HEK293T cell line were shown (n=2). Figure 9 A and B ( Figure 9 A and B show that the indel level of GEBx0305 varies with the length of the bootstrap sequence (n=3). Figure 10 A and B ( Figure 10 (A and B) show that the indel level of GEBx0308 varies with the length of the bootstrap sequence (n=3). Figure 11 A and B ( Figure 11 (A and B) show that the indel level of GEBx0328 varies with the length of the bootstrap sequence (n=3). Figure 12 A and B ( Figure 12 (A and B) show the indel levels of GEBx0305 against endogenous genes, using a modified RNA backbone (n=3). Figure 13 A and B ( Figure 13 (A and B) show the indel levels of GEBx0308 against endogenous genes, using a modified RNA backbone (n=3). Figure 14 A and B ( Figure 14 A and B show the indel levels of GEBx0305 against endogenous genes under optimal conditions (n=3). Figure 15 A and B ( Figure 15 (A and B) show the indel levels of GEBx0308 against endogenous genes under optimal conditions (n=3). Figure 16 A and B ( Figure 16(A and B) show the indel levels of GEBx0328 targeting endogenous genes using a modified RNA backbone (n=3). Figure 17 A and B ( Figure 17 (A and B) show the indel levels of GEBx0328 against endogenous genes under optimal conditions (n=3). Figure 18 ( Figure 18 The results show the indel levels of GEBx0305 targeting endogenous genes in HEK293T cells after transfection with a liposome complex containing a fixed amount (20 ng) of sgRNA and different proportions of mRNA. SpCas9 was used as a positive control. Figure 19 ( Figure 19 The results showed the indel levels of GEBx0308 targeting endogenous genes after HEK293T cells were transfected with a liposome complex containing a fixed amount (20 ng) of sgRNA and different proportions of mRNA. Figure 20 ( Figure 20 The results showed the indel levels of GEBx0305 targeting endogenous genes in PHH cells after transfection with a liposome complex containing a fixed amount (20 ng) of sgRNA and different proportions of mRNA. SpCas9 was used as a positive control. Figure 21 ( Figure 21 The results showed the indel levels of GEBx0308 targeting endogenous genes after PHH cells were transfected with a liposome complex containing a fixed amount (20 ng) of sgRNA and different proportions of mRNA. Figure 22 ( Figure 22 The image shows the Guide-seq inserts of GEBx0305 for site 1 (CFTR-NGGAAAA-T5) and site 2 (EMX1-NGGAAAA-T5); Figure 23 ( Figure 23 The image shows the Guide-seq inserts of GEBx0308 for site 1 (CD34-NGGTACT-T4) and site 2 (POLQ-NGGTACT-T1); Figure 24 ( Figure 24 The image shows Guide-seq inserts of GEBx0328 for (CFTR-NGGCCT-T3) and site 2 (CFTR-NGGCCT-T5); Figure 25 ( Figure 25The results show the base editing efficiency of GEBx0305-ABE for adenine A-to-G conversion at four sites in HEK293T cells; Figure 26 ( Figure 26 The results show the base editing efficiency of GEBx0308-ABE for adenine A-to-G conversion at five sites in HEK293T cells; Figure 27 ( Figure 27 The results showed the base editing efficiency of GEBx0328-ABE for adenine A-to-G conversion at five sites in HEK293T cells.
[0081] Figure 28 ( Figure 28 This study demonstrates 4-allele-specific editing of GEBx0308 at the RHO-P23H pathogenic site. Invention Details
[0083] The following examples further illustrate the content of this disclosure, but this disclosure is not limited thereto.
[0084] It is necessary to point out that the singular forms or terms such as “a,” “an,” “this,” and “the,” as well as similar terms used in the context of this disclosure (particularly in the context of the claims), should be interpreted to cover both the singular and the plural, unless otherwise stated herein or clearly contradicted by the context. In some embodiments, the foregoing terms may reasonably be understood as “a” or “one or more.” Furthermore, unless the context requires otherwise, singular terms should include the plural, and plural terms should include the singular.
[0085] Unless otherwise stated, all singular / plural terms also include the active and past voice forms of the terms and should be interpreted according to the context of the text.
[0086] Please note that in this disclosure and particularly in the claims and / or paragraphs, terms such as “comprising,” “consisting of,” “including,” etc., may have the meanings given by U.S. Patent Law; for example, they may mean “included,” “contained,” “including,” etc.; while terms such as “substantially composed of” and “substantially composed of” have the meanings given by U.S. Patent Law. The term “a group of” refers to a specific set or collection of elements, components, or features. It may include one or more of the specified elements, components, or features. For example, “a group of A, B, or C” may refer to a collection that includes one or more of any specified elements A, B, or C. The claims include the possibility that there may be any single element (A, B, or C), any combination of two elements (A and B, A and C, or B and C), or all three elements together (A, B, and C). This phrase defines the invention according to the specified options, allowing for different combinations of the listed elements while still maintaining the scope of the claim.
[0087] When “t” or “T” appears as a nucleotide in the RNA sequence of this invention, it should be understood as “u” or “U”.
[0088] In the context of two or more nucleic acid or polypeptide sequences, the term “identity” refers to two or more sequences or subsequences being identical, or having a specific percentage of identical amino acid residues or nucleotides, using sequence comparison algorithms such as BLAST, BLAST 2.0, or FASTA, with default parameters as described below.
[0089] The word “exemplary” is used herein to mean as an example, instance, or illustration. Any aspect or design described as “exemplary” is not necessarily to be construed as being more preferred or advantageous than other aspects, embodiments, or designs.
[0090] As stated in the text, the words “optional” or “or” mean that the event, situation, or alternative described below may or may not occur, and the description includes instances where the event or situation occurs and instances where it does not occur.
[0091] The use of “or” or “ / ” is inclusive, meaning “and / or” unless otherwise stated; or interpreted according to the context. The term “and / or” is used here as in phrases such as “A and / or B” to include A and B; A or B; A alone; and B alone. Similarly, the term “and / or” is used here as in phrases such as “A, B and / or C” to cover each of the following combinations: A, B and C; A, B or C; A or C; A or B; B or C; A and C; A and B; B and C; A alone; B alone; and C alone.
[0092] When used in this document, terms such as “approximately” and “~” refer to the range of variation of measurable values, such as parameters, quantities, durations of time, etc., and should be understood to include variations from the specified value. It should be understood that the value referred to by the modifiers “approximately” or “~” is itself specific and explicitly disclosed.
[0093] The word “exemplary” is used herein to mean as an example, instance, or illustration. Any aspect or design described as “exemplary” is not necessarily to be construed as being more preferred or advantageous than other aspects, embodiments, or designs.
[0094] The terms “subject,” “individual,” and “patient” are used interchangeably herein and refer to vertebrates, preferably mammals, and more preferably humans. Mammals include, but are not limited to, mice, primates, humans, farm animals, sporting animals, and pets. Tissues, cells, and their offspring of biological entities, whether obtained in vivo or cultured in vitro, are also included.
[0095] "Variations" of the sequences disclosed herein include sequences that have one or more additions, deletions, stop positions, or substitutions compared to the sequences disclosed herein.
[0096] "Encoding" refers to the property of a specific sequence of nucleotides in a gene, such as cDNA or mRNA, to serve as a template for the synthesis of other macromolecules, such as a defined amino acid sequence. Therefore, a gene encodes a protein if the transcription and translation of the mRNA corresponding to that gene in a cell or other biological system produces a protein. Polynucleotides that encode proteins include all nucleotide sequences that are degenerate versions of each other and encode the same amino acid sequence or amino acid sequences that are substantially similar in form and function.
[0097] The terms “non-spontaneous” or “engineered” are used interchangeably to indicate human involvement. When referring to nucleic acid molecules or peptides, these terms imply that the nucleic acid molecule or peptide is substantially unrelated in nature to at least one of its naturally associated components. In all aspects and embodiments, whether or not these terms are included, it should be understood that they are preferably optional, and therefore preferably included or not. Furthermore, the terms “non-spontaneous” and “engineered” are used interchangeably and can therefore be used alone or in combination, with one replacing a mention of both. In particular, “engineered” is preferred and can replace “non-spontaneous” or “non-spontaneous and / or engineered” or “engineered, non-spontaneous”.
[0098] As described in this invention, the term "cleavage event" refers to a DNA break created on a target nucleic acid by a type II Cas nuclease in the CRISPR system described in this invention. In some embodiments, the cleavage event is a double-stranded DNA break. In some embodiments, the cleavage event is a single-stranded DNA break.
[0099] As described in this invention, the term "targeted" refers to the ability of a complex including CRISPR-related proteins and RNA guides to preferentially or specifically bind, for example, to a target nucleic acid, compared to other nucleic acids that do not have the same or similar sequences as the target nucleic acid.
[0100] According to this disclosure, the term "GEBx" followed by a numerical suffix is used as a generic code to represent nucleic acids or proteins. It is important to note that using the same code for nucleic acids and proteins or their derivatives does not mean that the substances represented by these codes are the same. In other words, GEBx-0305 may refer to a specific nucleic acid sequence in one context and a different protein in another. Some embodiments may demonstrate a direct correspondence between nucleic acids and proteins represented by the same or derived codes. Therefore, the "GEBx" code serves as an indexing system for organizing and referencing the various biomolecules described in this disclosure, and the meaning of the code will be understood based on the provided context.
[0101] Unless otherwise defined herein, scientific and technical terms relating to this disclosure should have the meaning commonly understood by a person of ordinary technical skill. The meaning and scope of terms should be unambiguous; however, in the event of any potential ambiguity, the definitions provided herein shall prevail over any dictionary or external definitions.
[0102] Various embodiments are described below. It should be noted that the specific embodiments are not intended as an exhaustive description or a limitation on the broader schemes discussed herein. An embodiment described in conjunction with a particular embodiment is not necessarily limited to that embodiment and can be practiced with any other embodiment. References throughout the specification to “in some embodiments,” “in some particular embodiments,” “in some preferred embodiments,” “in some typical embodiments,” or similar expressions mean that a particular feature, structure, or characteristic associated with that embodiment is included in at least one embodiment. Furthermore, specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments according to this disclosure. Moreover, although some embodiments described herein include features not included in other embodiments, combinations of features from different embodiments should be within the scope of the disclosure. For example, in the dependent claims, embodiments of any claim can be used in any combination.
[0103] The range of values referenced by an endpoint includes all numbers and decimals within their respective ranges, as well as the endpoint being referenced.
[0104] Various embodiments are described below. It should be noted that the specific embodiments are not intended as an exhaustive description or a limitation on the broader schemes discussed herein. An aspect described in conjunction with a particular embodiment is not necessarily limited to that embodiment and can be practiced in conjunction with any other embodiment. The phrases “in a particular embodiment,” “in some embodiments,” or “in certain specific embodiments” throughout the specification mean that a particular feature, structure, or characteristic associated with that embodiment is included in at least one embodiment. Therefore, the phrases “in a particular embodiment,” “in one embodiment,” or “in certain specific embodiments” appearing in different places throughout the specification do not necessarily refer to the same embodiment, but may refer to them. Furthermore, specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments according to this disclosure. Moreover, although some embodiments described herein include features not included in other embodiments, combinations of features from different embodiments should be within the scope of the disclosure. For example, in the appended claims, embodiments of any claim can be used in any combination.
[0105] All publications, published patent documents, and patent applications cited in this document are cited to the same extent that each individual publication, published patent document, or patent application is specifically and individually indicated as being cited.
[0106] This invention provides an engineered, non-naturally occurring type II CRISPR-associated (Cas) protein or a variant thereof, wherein the protein shares at least 70% identity with any amino acid sequence in SEQ ID NO: 1-71. In some embodiments, the Cas protein shares at least 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 1-71.
[0107] As described herein, "M" stands for the starting amino acid methionine, which is typically the starting point for the synthesis of many proteins. Naturally occurring Cas proteins usually begin with methionine as the first amino acid in their sequence. However, when scientists modify these proteins, for example by fusing them with nuclear localization signals (NLS) or other domains, this initial amino acid "M" may be replaced or altered to introduce new functions or properties.
[0108] This disclosure focuses particularly on the flexibility and functionality of the engineered Cas proteins. Aside from the intentional modification of the starting amino acid methionine (M) at position 1, the engineered Cas proteins exhibit significant sequence identity with the reference sequence, ranging from 70% to 100%. This strategic alteration not only aligns with our goal of customizing protein properties but also ensures the preservation of the Cas protein's inherent or intended functions. By manipulating the initial amino acid without compromising overall sequence similarity, specific properties are enhanced, such as improved cellular localization or the introduction of other advantageous features, while maintaining the fundamental characteristics that make Cas proteins indispensable tools in genome manipulation.
[0109] The present invention also provides an engineered, non-naturally occurring type II CRISPR-associated (Cas) protein, except for amino acid "M" at position 1 in the sequence, wherein the Cas protein is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% identical to any amino acid sequence in SEQ ID NO: 1-71.
[0110] As described in this disclosure, the terms "Cas protein," "CRISPR-related protein," or other similar terms refer to components of the CRISPR-Cas system, specifically proteins that are integral parts of the CRISPR-Cas system. These proteins may possess inherent nuclease activity, enabling them to cleave double-stranded DNA or RNA molecules in a sequence-specific manner, guided by complementary RNA molecules, such as Cas9 and Cas12, which have been widely used in genome editing applications. Furthermore, some Cas proteins may be engineered to retain only one of the two active sites required for double-strand cleavage, resulting in cleavage enzyme activity that allows them to introduce single-strand breaks in the target nucleic acid sequence to control the cleavage event. Additionally, some Cas proteins may be engineered to lack any inherent nuclease activity, referred to as inactivated Cas or nuclease-inactive Cas proteins. Despite lacking enzymatic function, these inactivated Cas proteins retain their specific binding ability to target nucleic acid sequences and are often used in conjunction with other effector domains (or functional domains) for gene regulation, epigenetics editing, and as components of advanced imaging systems. In some embodiments, Cas proteins can be used to reduce non-target effects. In some embodiments, an active Cas nuclease, a nicking enzyme, or an inactivated Cas may also be part of a fusion protein containing another effector domain. Fusion proteins containing these other effector domains (or functional domains) also fall within the scope of Cas proteins. In some embodiments, the Cas protein may be a cleaved form. In some embodiments, the Cas protein may also be an inducible Cas protein. In some embodiments, type II Cas proteins may be part of a self-inactivating system (SIN); in some embodiments, type II Cas nucleases may also be part of a co-activating system (SAM), as defined elsewhere in this document.
[0111] In some embodiments, the domain arrangement of the type II Cas protein includes a RuvC domain, a BH (bridged helix) domain, a REC domain, an HNH domain, and / or a CTD (C-terminal domain). The RuvC domain is a key catalytic site responsible for cleaving the target DNA strand. It contains three split RuvC subdomains that fold in a complex manner to form the active site where DNA cleavage occurs. These subdomains work in coordination, guided by guide RNA, to recognize and cleave DNA at specific locations. The BH domain, or bridged helix domain, acts as a structural link between the different domains of the Cas protein. It is recognized as an arginine-rich region containing multiple arginine amino acids. The abundance of these arginine residues is crucial for interaction with the phosphate backbone of the target DNA strand. Arginine residues can form hydrogen bonds with phosphate groups, helping to correctly locate and align the target DNA for cleavage. The REC domain, or recognition leaf, participates in the recognition of the target DNA sequence. It contributes to the specificity of the Cas protein by distinguishing between target and non-target sequences, ensuring that only the intended DNA fragment is cleaved. The HNH domain is another catalytic site that works in conjunction with the RuvC domain to cleave the complementary strand of the target DNA. Named for its characteristic histidine-aspartic-histidine sequence, it is crucial for the nuclease activity of Cas proteins. The CTD, or C-terminal domain, is typically involved in interactions with other proteins or cellular structures, contributing to the localization and regulation of Cas proteins within the cell. It may also play a role in the stability and overall conformation of Cas proteins, ensuring their functionality and specificity at their target. This complex domain arrangement enables type II Cas proteins to perform their precise and critical functions in CRISPR systems, making them invaluable tools for genome editing and manipulation.
[0112] In some embodiments, the Cas protein further includes an effector domain (or functional domain). Such an effector domain may have one or more types of enzymatic activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain includes nuclease, nickase, deaminase, reverse transcriptase, recombinase, methyltransferase, methyltransferase, acetyltransferase, transcription activator, transcription repressor domain, cryptochrome, photoinducible / controllable domain, or chemically inducible / controllable domain. In some embodiments, the Cas protein further includes one or more nuclear localization signal sequences, nuclear export signal sequences, cell-penetrating peptide sequences, and affinity tags. Type II Cas proteins include one or more nuclear localization signals (NLS). The NLS may be located at the end of the peptide chain or elsewhere. The NLS located at both ends or other parts of the Cas9 amino acid sequence may be the same or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are the same. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains an NLS, and the C-terminus contains an NLS. The amino acid sequences of the NLS are fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence, respectively. The NLS may be an SV40 (monkey virus 40) NLS, a c-Myc NLS, or other suitable monomeric NLS. The NLS may be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, the Cas protein is purified by affinity chromatography using an affinity tag (such as a GST, FLAG, or a six-histidine sequence). In some embodiments, the amino acid sequence of the C-terminal NLS is given in SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is given in SEQ ID NO: 883. Other available sequences and different combinations may also be selected for the NLS and FLAG sequences.
[0113] In some embodiments, the Cas protein comprises an amino acid sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any of the amino acid sequences in SEQ ID NO: 1-71.
[0114] In some preferred embodiments, the Cas protein is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any of the amino acid sequences in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, 68. In some other preferred embodiments, except for amino acid "M" at position 1 in the sequence, the Cas protein is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any of the amino acid sequences in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, 68.
[0115] One of the key elements of the CRISPR-Cas system's modification process is the protospacer adjacent motif (PAM), a short DNA sequence immediately adjacent to the target DNA sequence. The PAM sequence is crucial for the binding and cleavage activity of the Cas protein, ensuring that only the predetermined DNA sequence is targeted for editing. The Cas protein's ability to recognize these specific PAM sequences is attributed to their unique structural features and binding interactions with DNA. Each PAM sequence provides a distinct motif that the Cas protein can recognize, ensuring accurate and efficient targeting. For example, the NRRANH sequence may provide a specific nucleotide arrangement that the Cas protein can bind to with high affinity. This diversity of PAM sequences expands the potential applications of the CRISPR-Cas system. It enables researchers and scientists to target a wider range of DNA sequences for editing, increasing the versatility and effectiveness of this powerful gene-editing tool. Furthermore, understanding these PAM sequences can help develop more advanced Cas proteins with higher precision targeting capabilities, further advancing the field of gene editing and its potential benefits for research, medicine, and biotechnology. The Cas protein disclosed in this invention possesses the unique ability to recognize a variety of PAM sequences. These sequences include NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. The specific recognition of these sequences by these Cas proteins allows for greater flexibility in selecting target DNA sequences for editing.
[0116] As described in this invention, in the provided PAM sequence, “N” represents any one of the four standard DNA nucleotides: adenine (A), thymine (T), cytosine (C), or guanine (G). This code facilitates the inclusion of any nucleotide at a given position without requiring individual specification; “R” specifically represents a purine nucleotide, namely adenine (A) or guanine (G). Purines are essential in DNA structure and function, and this code simplifies the process of incorporating these larger nucleotides into the PAM sequence; “H” represents any nucleotide other than guanine (G), and therefore can represent adenine (A), thymine (T), or cytosine (C). This code helps exclude guanine (G) at specific positions, as its larger size may affect structural or functional aspects of the sequence. “Y” represents a pyrimidine nucleotide, namely thymine (T) or cytosine (C). Pyrimidines commonly appear in specific regions of DNA and RNA, and this code facilitates their inclusion without requiring specification of the exact nucleotide. “W” represents a weak base, which can be adenine (A) or thymine (T). This distinction is biochemically relevant because adenine and thymine share similar properties in certain situations, such as hydrogen bonding. "V" reflects any nucleotide other than thymine (T), thus including adenine (A), cytosine (C), or guanine (G). This code is helpful when thymine is not needed or preferred due to its unique chemical properties among pyrimidines. "M" represents either adenine (A) or cytosine (C).
[0117] In some embodiments, the Cas protein as described in this invention is capable of recognizing at least one protospacer adjacent motif (PAM) having or comprising the following sequences: NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 9, and is capable of recognizing protospacer adjacent motifs (PAMs) with the NRRANH sequence; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 100% amino acid sequence identity with SEQ ID NO: 9, and is capable of recognizing protospacer adjacent motifs (PAMs) with the NRRANH sequence; The amino acid sequence identity of 12 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRHACT; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 19 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing protospacer adjacent motifs (PAMs) with sequence NRAAR; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 21 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is able to recognize the original spacer adjacent motif (PAM) with the sequence NNNCCY.In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 22, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NNRYYYY; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 100%, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NNRYYYY; The amino acid sequence identity of 24 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NGG; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 25 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NNNCAA; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 27 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRNACN; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 29 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is able to recognize the original spacer adjacent motif (PAM) with the sequence NNGR;In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 30, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NGGNR; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 30, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NGGNR; The amino acid sequence identity of 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NNNCCH; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 36 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRRAAG; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 37 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRHRAC; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 38 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is able to recognize the original spacer adjacent motif (PAM) with the sequence NRYART.In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 43, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with NRHACC sequences; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 100% amino acid sequence identity with SEQ ID NO: 43, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with NRHACC sequences; The amino acid sequence identity of 44 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing protospacer adjacent motifs (PAMs) with NRAAR sequences; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 51 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRNVHH; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 56 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence YMACAW; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 59 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is able to recognize protospacer adjacent motifs (PAMs) with the sequence NAHAA.In some embodiments, the Cas protein has 97%, 98%, 99%, or 100% amino acid identity with SEQ ID NO: 60 and is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NNNCAA; in some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence NRNACN and is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRNACN; in some embodiments, the Cas protein has 97%, 98%, 99%, or 100% identity with the amino acid sequence NNNCAA; The amino acid sequence identity of 29 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing protospacer adjacent motifs (PAMs) with sequence NNGR; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 30 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NGGNR; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is able to recognize the original spacer adjacent motif (PAM) with the sequence NNNCCH.In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 36, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRRAAG; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 100% amino acid sequence identity with SEQ ID NO: 36, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRRAAG; The amino acid sequence identity of 37 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRHRAC; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 38 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRYART; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 43 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRHACC; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 44 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is able to recognize protospacer adjacent motifs (PAMs) with NRAAR sequence.In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 51, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRNVHH; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 100% amino acid sequence identity with SEQ ID NO: 51, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRNVHH; The amino acid sequence identity of 56 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence YMACAW; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 59 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NHAA; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 60 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRHAYY; in some embodiments, the Cas protein is associated with SEQ ID NO: The amino acid sequence identity of 68 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing proto-spacer adjacent motifs (PAMs) with the NGGHA sequence.
[0118] As described herein, the terms "recognizing," "recognizing," or "recognizing" refer to the ability of the Cas protein to form a functional complex with sgRNA at the DNA target site where the sgRNA hybridizes (i.e., the location where the spacer sequence of the sgRNA hybridizes), and this site is surrounded by a PAM sequence, wherein the Cas protein is able to perform its natural function, namely DNA cleavage or DNA binding. In this context, it is important to note that such DNA cleavage precludes the possibility that the type II Cas protein is a type II Cas nuclease with lost catalytic activity. For example, in the case of an inactivated type II Cas nuclease (e.g., an inactivated type II Cas nuclease), a complex can still be formed between the type II Cas nuclease, sgRNA, and the corresponding target if the desired PAM sequence is present, but such a complex will not lead to DNA cleavage.
[0119] In some embodiments, the Cas protein is a cleaving enzyme or an inactivated Cas protein. The DNA cleavage domain of the active Cas protein of the present invention comprises two subdomains: an HNH nuclease subdomain and a RuvC subdomain. Mutations within these subdomains can inhibit the nuclease activity of the Cas protein. In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 9, and includes mutations at residues D11 or H859; or with SEQ ID NO: The amino acid sequence identity of SEQ ID NO: 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes a mutation at residue D10 or H862; or the amino acid sequence identity with SEQ ID NO: 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes a mutation at residue D12 or H903.In some embodiments, except for amino acid "M" at position 1 in the sequence, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 9, and includes mutations at residues D11 or H859; or except for amino acid "M" at position 1 in the sequence, the amino acid sequence of Cas protein is identical to that of SEQ ID NO: 9. The amino acid sequence identity of 12 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes mutations at residue D10 or H862; or, except for amino acid "M" at position 1 in the sequence, it is identical to SEQ ID NO: The amino acid sequence identity of 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes a mutation at residue D12 or H903. In some embodiments, the mutation at residue D11 or H859 in SEQ ID NO: 9 is D11A or H859A; the mutation at residue D10 or H862 in SEQ ID NO: 12 is D10A or H862A; and the mutation at residue D12 or H903 in SEQ ID NO: 31 is D12A or H903A. In some embodiments, the Cas protein is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any of the amino acid sequences in SEQ ID NO: 869-877.
[0120] The present invention also provides an engineered, non-naturally occurring polynucleotide that encodes the type II CRISPR-associated (Cas) protein disclosed in this invention.
[0121] According to the description of the invention, the term "polynucleotide" refers to a polymeric form of nucleotides of any length, whether deoxyribonucleotides or ribonucleotides, or the like. Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. Thus, this term includes, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers composed of purine and pyrimidine bases or other naturally, chemically, or biochemically modified non-natural or derived nucleotide bases. In some embodiments, the polynucleotide encoding the type II Cas protein described in this invention can be operatively linked to each other and to associated regulatory sequences (such as promoters, enhancers, and termination regions). For example, a functional link between a regulatory sequence and a foreign nucleic acid sequence can result in the expression of the latter. In other embodiments, it can be said that the first nucleic acid sequence is operatively linked to the second nucleic acid sequence when the first nucleic acid sequence is in a functional relationship with the second nucleic acid sequence. For example, if a promoter affects the transcription or expression of a coding sequence, then the promoter is operatively linked to the coding sequence. Typically, the operatively linked DNA sequences are contiguous, and the coding regions are linked into the same reading frame when necessary or helpful. In some embodiments, the promoter is a constitutive promoter, a tissue-specific promoter, or an inducible promoter. In some embodiments, the term may also include all introns and other DNA sequences spliced from mRNA transcripts, as well as variants resulting from alternative splicing sites. These nucleic acid sequences may be DNA strand sequences transcribed into RNA or RNA sequences translated into proteins. Nucleic acid sequences include full-length nucleic acid sequences and non-full-length sequences derived from full-length proteins. Sequences may also include degenerate codons of the original sequence or sequences introduced to provide codon preference in a particular cell type.
[0122] In some embodiments, the polynucleotide encoding the Cas protein is operatively linked to a promoter and presented in a vector; alternatively, the vector is selected from the group consisting of: retroviral vectors, lentiviral vectors, phage vectors, adenovirus vectors, adeno-associated virus vectors, herpes simplex virus vectors, and plasmid vectors.
[0123] In some embodiments, the polynucleotide is a ribonucleotide sequence or a deoxyribonucleotide sequence, or an analogue thereof; optionally, the polynucleotide is codon-optimized for expression in the cells of interest; in some embodiments, the polynucleotide is codon-optimized for expression in eukaryotic cells. In some embodiments, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cells are mammalian cells, preferably human cells.
[0124] In some embodiments, the polynucleotide is mRNA and further comprises a 5' cap sequence and / or a poly-A tail sequence. In some embodiments of the invention, the mRNA used may be modified to enhance its functional properties and stability. Specifically, in some embodiments, the modification process involves replacing uridine (represented by the letter "U") with N1-methylpseudouridine or pseudouridine. This substitution aims to improve the mRNA's resistance to ribonuclease degradation, potentially increasing its intracellular half-life and translation efficiency. Incorporating N1-methylpseudouridine or pseudouridine into the mRNA structure can also positively influence the immunogenicity profile, as these modifications have been shown to reduce the immunogenicity of the mRNA molecule compared to unmodified mRNA molecules. This is particularly critical for the development of mRNA therapeutics and vaccines, where minimizing adverse immune responses is essential.
[0125] In some embodiments, the polynucleotides of the present invention are codon-optimized for expression in eukaryotic cells; alternatively, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, single-celled eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, and non-human primate cells.
[0126] In some embodiments, the polynucleotide shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the nucleotide sequences in SEQ ID NO: 161-231, 241-311, 851-852, 861-863.
[0127] The present invention also provides an engineered, non-naturally occurring CRISPR-Cas system comprising: a) a type II Cas protein or a polynucleotide encoding a Cas protein as described herein; b) at least one engineered guide RNA or at least one engineered nucleic acid encoding a guide RNA, wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein binding segment that interacts with the Cas protein, wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer.
[0128] According to the description of the invention, the term "complementarity" describes the ability of two nucleic acid strands to pair with each other through their bases. This complementarity can occur through perfect base pairing, where each base on one strand forms a specific hydrogen bond with its complementary base on the opposite strand, following the standard Watson-Crick base pairing rules: adenine (A) pairs with thymine (T) or uracil (U), and cytosine (C) pairs with guanine (G). This perfect base pairing allows for precise recognition and binding between the two strands. Furthermore, complementarity can also involve imperfect base pairing, including mismatches, insertions, or deletions that may result in non-standard base pairing or reduced inter-strand affinity. Despite these imperfections, sufficient complementarity is maintained between the strands to interact, although this may result in reduced specificity or stability in a duplex.
[0129] According to the description of the present invention, the direct repeat (DR) sequence originates from short repetitive DNA sequence elements in a CRISPR array. This sequence is typically scattered among spacer sequences derived from exogenous genetic material, such as bacteriophage or plasmid DNA. Post-transcriptionally, the DR sequence is part of the precrRNA transcript and is processed into mature CRISPR RNA (crRNA). The DR sequence in crRNA is a key component in pairing with its complementary sequence in transcriptionally activated CRISPR RNA (tracrRNA) to form a double-stranded RNA (dsRNA) dimer. This is essential for the function of the guide RNA (gRNA) complex in the CRISPR / Cas system. In the context of the present invention, the DR sequence includes the portion of gRNA that pairs with the tracrRNA sequence to form a dsRNA dimer, thereby enabling the formation of the Cas protein-binding segment. The term "hybridization to form a double-stranded RNA (dsRNA) dimer," or similar expressions, refers in this invention to the process by which a direct repeat (DR) sequence in CRISPR RNA (crRNA) pairs with a complementary sequence in transcriptionally activated CRISPR RNA (tracrRNA) to form a stable double-stranded RNA structure. This dsRNA dimer is a key component of the guide RNA (gRNA) complex, which facilitates Cas protein binding. Specifically, the DR sequence in crRNA and its complementary sequence in tracrRNA pair base-pair, forming hydrogen bonds between their complementary bases, leading to the formation of the dsRNA dimer. This dimer is essential for the function of the Cas protein binding segment, which consists of the tracrRNA and DR sequences that hybridize to form the dsRNA dimer.
[0130] According to the description of the present invention, the term "target nucleic acid" refers to a specific nucleic acid substrate comprising a nucleic acid sequence that is wholly or partially complementary to an RNA guide sequence. In some embodiments, the target nucleic acid comprises a gene or a sequence within a gene. In some embodiments, the target nucleic acid comprises a non-coding region (e.g., a promoter). In certain embodiments, the target nucleic acid is single-stranded. In certain embodiments, the target nucleic acid is double-stranded. The terms "target nucleic acid" or "target sequence" should be understood according to the context of the disclosure.
[0131] In some embodiments, the guide RNA is a dual guide RNA. In some embodiments, the guide RNA is a single guide RNA. In these embodiments, the guide RNA further includes a linker sequence connecting the tracrRNA sequence and the DR sequence. In some typical embodiments, the linker sequence comprises a short GAAA sequence. In some embodiments, the linker sequence is an artificial loop. In some embodiments, the sgRNA comprises the following sequence: a) a spacer sequence capable of hybridizing with the target nucleic acid sequence to be manipulated; b) the DR sequence; c) the linker sequence; and d) the tracrRNA sequence. The spacer sequence, DR sequence, linker sequence, and tracrRNA sequence are tandemly arranged in a 5' to 3' direction or a 3' to 5' direction; in some embodiments, the sgRNA backbone includes a sequence that has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the sequences in SEQ ID NO:431-562.
[0132] In some embodiments, the sgRNA comprises a spacer sequence (such as any of the sequences in SEQ ID NO: 571-835, 885, 901) and a backbone sequence, wherein the spacer sequence is located at the 5' end of the backbone sequence (such as the sequence in SEQ ID NO: 903). In some embodiments, the sgRNA comprises a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the sequence in SEQ ID NO: 903.
[0133] In some embodiments, the spacer sequence hybridizes with one or more nucleic acids in a prokaryotic or eukaryotic cell. In some embodiments, the eukaryotic cell is selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the eukaryotic cell includes mammalian cells. In some embodiments, the mammalian cell includes human cells. In some embodiments, the eukaryotic cell includes plant cells.
[0134] In some embodiments, the system also includes a donor template nucleic acid.
[0135] As described in this disclosure, the term "donor template nucleic acid" refers to a nucleic acid molecule that one or more cellular proteins can use to alter the structure of a target nucleic acid after a Cas protein has altered the structure of that target nucleic acid. In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid. In some embodiments, the donor template nucleic acid is a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear. In some embodiments, the donor template nucleic acid is circular (e.g., a plasmid). In some embodiments, the donor template nucleic acid is a foreign nucleic acid molecule. In some embodiments, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, the donor template nucleic acid is DNA or RNA or a DNA-RNA hybrid.
[0136] The present invention also provides an engineered vector comprising the polynucleotides described in this disclosure.
[0137] As described in this disclosure, a “vector” is a tool that allows or facilitates the transfer of an entity from one environment to another. It is a replicon, such as a plasmid, bacteriophage, or cosmosome, that can insert another DNA fragment to achieve replication of the inserted fragment. Typically, a vector is capable of replication when associated with appropriate control elements. Generally, a “vector” refers to a nucleic acid molecule capable of transporting another nucleic acid molecule linked to it. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends or without free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other polynucleotide variants known in the art. One type of vector is a “plasmid,” which refers to a circular, double-stranded DNA loop into which additional DNA fragments can be inserted, as through standard molecular cloning techniques. Another type of vector is a viral vector, in which a virus-derived DNA or RNA sequence is present within the vector for packaging into a virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses (AAVs)). Viral vectors also include polynucleotides carried by the virus for transfection into host cells. Some vectors are capable of autonomous replication in the host cells in which they are introduced (e.g., bacterial vectors with bacterial origins of replication and circular mammalian vectors). Other vectors (e.g., non-circular mammalian vectors) integrate into the host genome upon introduction into the host cell and replicate along with the host genome. Furthermore, some vectors are capable of directing the expression of genes operatively linked to them. These vectors are referred to herein as “expression vectors.” Expression vectors commonly used in recombinant DNA technology are typically in plasmid form. Recombinant expression vectors may contain the nucleic acids of the present invention in a form suitable for expression in host cells, meaning that the recombinant expression vector includes one or more regulatory elements selected based on the host cell used, operatively linked to the nucleic acid sequence to be expressed. In recombinant expression vectors, “operatively linked” means that the nucleotide sequence of interest is linked to the regulatory element in a manner that allows the nucleotide sequence to be expressed (e.g., in an in vitro transcription / translation system or when the vector is introduced into a host cell). In some embodiments, the vector is an expression vector. In some embodiments, the vector is an inducible, conditional, or constitutive expression vector. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA are located on the same vector or on different vectors.
[0138] The present invention also provides a vector system comprising one or more polynucleotides described herein and one or more polynucleotides encoding guide RNA; wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein-binding segment that interacts with the Cas protein, wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA are located on the same vector or on different vectors.
[0139] The present invention also provides an engineered, non-naturally occurring cell comprising: the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, or the vector system described herein.
[0140] The present invention also provides a cell modified by utilizing the Cas protein, the polynucleotide, the CRISPR-Cas system, the vector, or the vector system described in this disclosure.
[0141] In some embodiments, the cells are eukaryotic or prokaryotic cells. In some embodiments, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cells are mammalian cells, human cells, or plant cells.
[0142] In some embodiments, the cell is a vertebrate, mammal, rodent, goat, pig, bird, chicken, turkey, cow, horse, sheep, fish, primate, or human cell. In some embodiments, the cell is a mammalian cell. In one embodiment, the cell is a human cell. In some embodiments, the cell is a somatic cell, germ cell, or fetal cell. In some embodiments, the cell is a zygote, blastocyst, embryonic cell, stem cell, mitotically capable cell, or meiotically capable cell. In some embodiments, the cell is not part of a human embryo. In some embodiments, the cell is a somatic cell. In one embodiment, the cell is a T cell, CD8+ T cell, CD8+ naive T cell, central memory T cell, effector memory T cell, or CD4+ T cell. T cells, stem cell memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricted progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythrocyte progenitor cells, megakaryocytic erythrocyte progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, aqueous reticulum cells, cochlear hair cells, outer hair cells, inner hair cells, lung epithelial cells, bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, nerve cells. The cells may include: neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells (e.g., progenitor B cells, pre-B cells, memory B cells), plasma cells, gastrointestinal epithelial cells, biliary epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, preadipocytes, pancreatic islet cells (e.g., β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0143] In some embodiments, this disclosure provides an isolated eukaryotic cell comprising a modified target site, wherein the target site has been modified using the methods described in this invention, or using the systems described in this invention, or using the Cas protein described in this invention, or using the polynucleotide described in this invention, or using the CRISPR-Cas system described in this invention, or using the vector described in this invention, or using the vector system described in this invention, or using the kit described in this invention, or using the pharmaceutical composition described in this invention.
[0144] In some embodiments, the cells are eukaryotic or prokaryotic cells. In some embodiments, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cells are mammalian cells, human cells, or plant cells.
[0145] In some embodiments, the cell is a vertebrate, mammal, rodent, goat, pig, bird, chicken, turkey, cow, horse, sheep, fish, primate, or human cell. In some embodiments, the cell is a mammalian cell. In one embodiment, the cell is a human cell. In some embodiments, the cell is a somatic cell, germ cell, or fetal cell. In some embodiments, the cell is a zygote, blastocyst, embryonic cell, stem cell, mitotically capable cell, or meiotically capable cell. In some embodiments, the cell is not part of a human embryo. In some embodiments, the cell is a somatic cell. In one embodiment, the cell is a T cell, CD8+ T cell, CD8+ naive T cell, central memory T cell, effector memory T cell, or CD4+ T cell. T cells, stem cell memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricted progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythrocyte progenitor cells, megakaryocytic erythrocyte progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, aqueous reticulum cells, cochlear hair cells, outer hair cells, inner hair cells, lung epithelial cells, bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, nerve cells. The cells may include: neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells (e.g., progenitor B cells, pre-B cells, memory B cells), plasma cells, gastrointestinal epithelial cells, biliary epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, preadipocytes, pancreatic islet cells (e.g., β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0146] In some embodiments, the vector, such as a plasmid or viral vector, is delivered to the target tissue via intramuscular injection, intravenous administration, transdermal administration, intranasal administration, oral administration, or mucosal administration. Such administration can be a single dose or multiple doses. Those skilled in the art will understand that the actual dose can vary considerably due to a variety of factors, such as the choice of vector, target cells, organism, tissue, general condition of the subject, the required degree of transformation / modification, route of administration, manner of administration, and type of transformation / modification required.
[0147] In some embodiments, delivery is made via adeno-associated virus (AAV), such as AAV2, AAV8, or AAV9, which may contain at least 1 × 10^5 adenovirus or adeno-associated virus particles (also referred to as particle units, pu) in a single dose. In some embodiments, the dose contains at least approximately 1 × 10^6 particles, at least approximately 1 × 10^7 particles, at least approximately 1 × 10^8 particles, or at least approximately 1 × 10^9 particles of adeno-associated virus. Due to the limited genomic payload of recombinant AAV, the smaller size of the type II Cas nucleases described in this invention allows for greater flexibility in packaging effectors and RNA guidelines, as well as the appropriate control sequences (e.g., promoters) required to achieve efficient and cell-type-specific expression.
[0148] In some embodiments, delivery is performed via a recombinant adeno-associated virus (rAAV) vector. For example, in some embodiments, a modified AAV vector may be used for delivery. The modified AAV vector may be based on one or more capsid types, including AAV1, AAV2, AAV5, AAV6, AAV8, AAV8.2, AAV9, AAV rh10, modified AAV vectors (e.g., modified AAV2, modified AAV3, modified AAV6), and pseudotyped AAVs (e.g., AAV2 / 8, AAV2 / 5, and AAV2 / 6).
[0149] In some embodiments, delivery is made via plasmid. The dosage may be sufficient to elicit a response. In some embodiments, the appropriate amount of plasmid DNA in the plasmid composition may range from about 0.1 to about 2 mg. A plasmid typically comprises (i) a promoter; (ii) a sequence encoding a CRISPR enzyme targeting a nucleic acid, operatively linked to the promoter; (iii) a selectivity marker; (iv) an origin of replication; and (v) a transcription terminator, downstream of (ii) and operatively linked to (ii). The plasmid may also encode RNA components of the CRISPR-Cas system, but one or more of these may be encoded on different vectors. The frequency of administration is determined by a medical or veterinary practitioner (e.g., a physician, veterinarian) or skilled technician.
[0150] The present invention also provides a kit comprising the Cas protein described in this disclosure, the polynucleotide described in this disclosure, the CRISPR-Cas system described in this disclosure, the vector described in this disclosure, the vector system described in this disclosure, or the cell described in this disclosure.
[0151] The kits described in this disclosure may include one or more containers containing components necessary to perform the methods of this disclosure, and may include instructions for use. Any described kit may also include auxiliary components required to perform the editing methods. Each component in the kit may be provided in liquid form (e.g., dissolved in solution) or solid form (e.g., lyophilized powder), where applicable. In certain embodiments, some components may need to be reformulated or otherwise treated (e.g., activated state) after the addition of suitable solvents or other substances (such as water or buffers), which may or may not be provided with the kit. In some embodiments, the kit may further include other suitable excipients, such as buffers or reagents, to facilitate the application of the kit. The kit can be used in a variety of applications, such as medical applications, including therapeutic and diagnostic, research, etc. Therefore, the type II Cas nuclease and kit of the present invention can be used to prepare pharmaceuticals or reagents for therapeutic and / or research purposes.
[0152] Cas proteins, the CRISPR-Cas system, and polynucleotides, as described herein, can be delivered via various delivery systems, such as vectors (e.g., plasmids), viral delivery vectors (e.g., adeno-associated virus (AAV), lentivirus, adenovirus, or other viral vectors), or methods (e.g., ribo-electroporation or electroporation of ribonucleoprotein complexes composed of type V-VI effectors and their corresponding RNA guides or guides). Proteins and one or more RNA guides can be packaged into one or more vectors, such as plasmids or viral vectors. For bacterial applications, nucleic acids encoding any of the CRISPR system components described herein can be delivered to bacteria using bacteriophages. Exemplary bacteriophages include, but are not limited to, T4 phage, Mu, λ phage, T5 phage, T7 phage, T3 phage, Φ29, M13, MS2, Qβ, and ΦX174.
[0153] The present invention also provides a pharmaceutical composition comprising the Cas protein described in this disclosure, the polynucleotide described in this disclosure, the CRISPR-Cas system described in this disclosure, the vector described in this disclosure, the vector system described in this disclosure, or the cell described in this disclosure.
[0154] As described in this disclosure, a "pharmaceutical composition" refers to a formulation intended for pharmaceutical use. In some embodiments, the pharmaceutical composition further includes an acceptable pharmaceutical excipient. In some embodiments, the pharmaceutical composition may contain other therapeutic agents. In some embodiments, the pharmaceutical composition is prepared according to standard procedures and can be administered to a subject, such as a human patient, via intravenous, intramuscular, intradermal, intra-articular, intralesional, intraperitoneal, intracardiac, intracerebrospinal fluid, intraventricular, epidural, topical, subconjunctival, periocular, intraocular, vitreous, posterior sclera, penetrating sclera, suprascleral, subretinal, retroretinal, fundus, intranasal inhalation, pressurized inhalation, oral, subcutaneous, or topical routes. For example, compositions for injection may be provided as a sterile isotonic aqueous solution. If necessary, the pharmaceutical composition may also contain a solvent and a local anesthetic, such as lidocaine, to minimize discomfort at the injection site. Typically, components may be provided individually or as a mixture of unit doses, for example as a lyophilized powder or anhydrous concentrated solution, in a sealed container indicating the amount of active ingredient. If the pharmaceutical composition is intended for infusion, it can be combined with an infusion bottle containing sterile pharmaceutical-grade water or physiological saline. If the pharmaceutical composition is intended for injection, sterile water for injection or physiological saline may be included to mix the components before administration. Furthermore, wetting agents, colorants, release agents, coating agents, sweeteners, flavoring agents, preservatives, and antioxidants may also be incorporated into the formulation as needed.
[0155] In some embodiments, the pharmaceutical composition further includes a delivery system selected from: AAV (adeno-associated virus), adenovirus, retrovirus, HSV (herpes simplex virus), gamma retrovirus, lentivirus, eCIS (extracellular contractile injection system), eVLPs (engineered virus-like particles), VLPs (virus-like particles), liposomes, plasmids, LNPs (lipid nanoparticles), exosomes, microvesicles, nucleic acid nanoassemblies, gene guns, and / or implantable devices.
[0156] The present invention also provides methods for treating, preventing, diagnosing or detecting diseases using the Cas proteins, polynucleotides, CRISPR-Cas systems, vectors, vector systems, cells, kits or pharmaceutical compositions described in this disclosure.
[0157] The present invention also provides a method for modifying or targeting a target DNA site, the method comprising delivering a Cas protein, a polynucleotide, a CRISPR-Cas system, a vector, a vector system, a kit, or a pharmaceutical composition described herein to the site.
[0158] In some embodiments, the disclosure also provides a method for targeting and cleaving target DNA, the method comprising: contacting the target DNA with a Cas protein, a polynucleotide, a CRISPR-Cas system, a vector, a vector system, a kit, or a pharmaceutical composition as described in this disclosure.
[0159] In some embodiments, modifying or targeting a target site includes inducing DNA strand breaks. In some embodiments, modifying or targeting a target site includes inducing DNA double-strand breaks or DNA single-strand breaks. In some embodiments, modifying or targeting a target site includes altering the gene expression of one or more genes. In some embodiments, modifying or targeting a target site includes epigenetic modification of a target DNA site. In some embodiments, the method is a method of modifying a cell, cell line, or organism by manipulating one or more target sequences at a site of interest in the genome.
[0160] In some embodiments, cleaving the target DNA or target sequence results in the formation of an insertion or deletion (indel) or the insertion of a nucleotide sequence. In some embodiments, cleaving the target DNA or target nucleotide includes cleaving the target DNA or target sequence at two sites, resulting in deletion or inversion of the sequence between the two sites. In some embodiments, the target DNA is double-stranded DNA or single-stranded DNA, or a DNA-RNA hybrid.
[0161] In some embodiments, modifying or targeting a target site includes inducing DNA strand breaks, altering the gene expression of one or more genes, or epigenetic modification of the target DNA site; optionally, DNA strand breaks include DNA double-strand breaks or DNA single-strand breaks.
[0162] In some embodiments, the method is performed in vitro or in vivo.
[0163] The present invention also provides an isolated eukaryotic cell comprising a modified target site, wherein the target site has been modified by the method described in the present invention, or by using the system described in the present invention, or by using the Cas protein described in the present invention, or by using the polynucleotide described in the present invention, or by using the CRISPR-Cas system described in the present invention, or by using the vector described in the present invention, or by using the vector system described in the present invention, or by using the kit described in the present invention, or by using the pharmaceutical composition described in the present invention.
[0164] The present invention also provides a system for detecting the presence of a nucleic acid target sequence in an in vitro sample, comprising: a) the Cas protein described in this disclosure; b) at least one guide polynucleotide comprising a guide sequence capable of binding to the target sequence and designed to form a complex with the Cas protein; and c) a nucleic acid-based masking structure comprising a non-target sequence, wherein the Cas protein exhibits incidental cleavage activity against RNA and / or ssDNA and cleaves the non-target sequence in the nucleic acid-based masking structure activated by the target sequence.
[0165] The present invention also provides a method for detecting target nucleic acids in a sample, comprising: contacting one or more samples with a) the Cas protein described in this disclosure; b) at least one guide polynucleotide comprising a guide sequence designed to be complementary to a target sequence and designed to form a complex with the Cas protein; and c) a nucleic acid-based masking structure comprising a non-target sequence, wherein the Cas protein exhibits incidental cleavage activity against RNA and / or ssDNA and cleaves the non-target sequence in the nucleic acid-based masking structure activated by the target sequence; and detecting a signal of non-target sequence cleavage to detect one or more target sequences in the sample.
[0166] As described in this disclosure, a “sample” may contain whole cells and / or live cells and / or cell debris. A sample may contain (or be derived from) “body fluids.” This disclosure includes examples of body fluids selected from amniotic fluid, aqueous humor, vitreous humor, bile, serum, breast milk, cerebrospinal fluid, earwax, lymph, gastric juice, mucus (including nasal secretions and sputum), ascites, pericardial fluid, pleural fluid, pus, tears, saliva, sebum (skin oil), semen, sputum, synovial fluid, sweat, urine, vaginal secretions, vomit, and one or more mixtures thereof. Samples include cell cultures, body fluids, and cell cultures derived from body fluids. Body fluids can be obtained from mammals, for example, through puncture or other collection or sampling procedures.
[0167] The present invention also provides a guide RNA (gRNA) comprising: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA further comprises a Cas protein-binding segment; wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence, which hybridize to form a double-stranded RNA (dsRNA) dimer. In some embodiments, the gRNA is a dual guide RNA. In some embodiments, the gRNA is a single guide RNA. In some embodiments, the gRNA is modified. In some embodiments, at least three nucleotides of the gRNA are modified. In one embodiment, the gRNA comprises a 5' end modification that contains at least two phosphothioester (PS) bonds in the first seven nucleotides of the 5' end. In some embodiments, the gRNA includes a 3' end modification that contains at least two phosphothioester (PS) bonds in the first seven nucleotides of the 3' end. In some embodiments, the gRNA exhibits at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any nucleotide sequence of SEQ ID NO: 903.
[0168] The present invention also provides a polynucleotide encoding the gRNA described in this document; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises a sequence that has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the sequences in SEQ ID NO: 903.
[0169] An engineered, non-naturally occurring CRISPR-Cas system comprising: a) a Cas protein or a polynucleotide encoding a Cas protein; b) at least one gRNA or at least one engineered nucleic acid encoding a gRNA as described in this disclosure, wherein the gRNA further comprises a Cas protein binding segment that interacts with the Cas protein; wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 903.
[0170] In some embodiments, the Cas protein comprises a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to any of the sequences in SEQ ID NO: 12; or, except for the amino acid "M" at position 1 in the sequence, the Cas protein comprises a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 98%, 99%, or 100% identity to any of the sequences in SEQ ID NO: 12; 12. Any sequence identity is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0171] In some embodiments, the Cas protein further includes an effector domain (or functional domain). Such an effector domain may have one or more types of enzymatic activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain includes nuclease, nickase, deaminase, reverse transcriptase, recombinase, methyltransferase, methyltransferase, acetyltransferase, transcription activator, transcription repressor domain, cryptochrome, photoinducible / controllable domain, or chemically inducible / controllable domain.
[0172] In some embodiments, the Cas protein further comprises one or more nuclear localization signal sequences, nuclear export signal sequences, cell-penetrating peptide sequences, and affinity tags. Type II Cas proteins comprise one or more nuclear localization signals (NLS). The NLS may be located at the ends of the peptide chain or elsewhere. NLS located at both ends or other parts of the Cas9 amino acid sequence may be the same or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are the same. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains one NLS, and the C-terminus contains one NLS. The amino acid sequences of the NLS are fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence, respectively. The NLS may be SV40 (monkey virus 40) NLS, c-MycNLS, or other suitable monomeric NLS. The NLS may be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, the Cas protein is purified by affinity chromatography using an affinity tag (such as GST, FLAG, or a six-histidine sequence). In some embodiments, the amino acid sequence of the C-terminal NLS is given in SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is given in SEQ ID NO: 883. Other available sequences and different combinations may also be selected for the NLS and FLAG sequences.
[0173] In some embodiments, the Cas protein comprises an amino acid sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 12.
[0174] In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence NRHACT, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRHACT.
[0175] In some embodiments, the Cas protein is a cleaving enzyme or an inactivated Cas protein. The DNA cleavage domain of the active Cas protein of the present invention comprises two subdomains, namely the HNH nuclease subdomain and the RuvC subdomain. Mutations within these subdomains can inhibit the nuclease activity of the Cas protein. In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 12, and includes mutations at residues D10 or H862; in some embodiments, except for the amino acid "M" at position 1 in the sequence, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 78%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 12. The amino acid sequence identity of 12 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes mutations at residues D10 or H862.
[0176] In some embodiments, the mutation at residue D10 or H862 in SEQ ID NO: 12 is D10A or H862A; in some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 872-874.
[0177] The present invention also provides an engineered vector comprising a polynucleotide encoding the gRNA disclosed herein; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the sequence of SEQ ID NO: 903.
[0178] In some embodiments, the vector may be an inducible, conditional, or constitutive expression vector.
[0179] The present invention also provides a vector system comprising one or more polynucleotides encoding the gRNA disclosed herein; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 871 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the sequence of SEQ ID NO: 903.
[0180] The present invention also provides a pharmaceutical composition comprising the gRNA disclosed herein; a polynucleotide encoding the gRNA; a CRISPR-Cas system comprising the gRNA or polynucleotide; a vector comprising the gRNA coding sequence; and a vector system comprising the gRNA coding sequence; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the sequence of SEQ ID NO: 903.
[0181] The present invention also provides a method for treating, preventing, or diagnosing diseases associated with the RHO gene locus, comprising administering a composition to a subject in need, wherein the composition comprises: a) a guide RNA comprising a spacer sequence of SEQ ID NO: 901; b) a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c) a guide RNA comprising a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence.
[0182] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a) a guide RNA comprising a spacer sequence of SEQ ID NO: 901; b) a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c) a guide RNA comprising a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence.
[0183] The present invention also provides a method for modifying RHO gene sites, comprising delivering a composition to cells, wherein the composition comprises: a) sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b) sgRNA comprising an sgRNA sequence having at least 90% identity with the sequence of SEQ ID NO: 903; or c) sgRNA comprising an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 903.
[0184] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a) an sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b) an sgRNA comprising an sgRNA sequence having at least 90% identity with the sequence of SEQ ID NO: 903; or c) an sgRNA comprising an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 903.
[0185] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a) a guide RNA comprising a spacer sequence of SEQ ID NO: 901; b) a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c) a guide RNA comprising a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence.
[0186] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a. a guide RNA comprising a guide sequence of SEQ ID NO: 901; b. a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c. a guide RNA comprising a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence.
[0187] The present invention also provides a method for modifying RHO gene sites, comprising delivering a composition to cells, wherein the composition comprises: a. sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b. sgRNA comprising an sgRNA sequence having at least 90% identity with the sequence of SEQ ID NO: 903; or c. sgRNA comprising an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 903.
[0188] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a. an sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b. an sgRNA comprising an sgRNA sequence having at least 90% identity with the sequence of SEQ ID NO: 903; or c. an sgRNA comprising an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 903.
[0189] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a. a guide RNA comprising a spacer sequence of SEQ ID NO: 901; b. a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c. a guide RNA comprising a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence.
[0190] The present invention also provides a method for modifying RHO gene sites, comprising delivering a composition to a cell, wherein the composition comprises: (i) a Cas protein, wherein: a. the Cas protein contains a sequence that is at least 90% identical to SEQ ID NO: 12 or 92; and / or b. an RNA-guided DNA binder that contains a sequence that is at least 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 12 or 92; and / or (ii) a guide RNA or a vector encoding the guide RNA, wherein the guide RNA contains a spacer sequence of SEQ ID NO: 901.
[0191] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: (i) an RNA-guided DNA binder, wherein: a. the RNA-guided DNA binder contains a sequence that is at least 90% identical to SEQ ID NO: 12 or 92; and / or b. the RNA-guided DNA binder contains a sequence that is at least 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 12 or 92; and / or (ii) an sgRNA or a vector encoding sgRNA, wherein the sgRNA contains the sequence of SEQ ID NO: 903.
[0192] Table 1: Exemplary amino acid sequences involved in this disclosure.
[0193]
[0194] Table 2: Exemplary polynucleotides encoding Cas proteins described in this disclosure.
[0195]
[0196] Table 3: Exemplary direct repeat sequences described in this disclosure.
[0197]
[0198] In some embodiments, this disclosure provides an engineered, non-naturally occurring crRNA or a variant thereof, wherein the crRNA comprises a nucleotide sequence having at least 90% identity with any of the sequences in SEQ ID NO: 341-417 (Table 3). In some embodiments, the crRNA comprises a nucleotide sequence having at least 95% or 98% identity with any of the sequences in SEQ ID NO: 341-417. In some embodiments, the crRNA comprises any of the sequences in SEQ ID NO: 341-417.
[0199] The following non-limiting examples provide further illustration of embodiments of this disclosure. Those skilled in the art should understand that the techniques disclosed in the following examples represent effective methods found in practicing this disclosure and can therefore be considered examples of their modes of practice. Those skilled in the art should understand, in consideration of this disclosure, that many modifications can be made to the specific embodiments disclosed and similar or identical results can still be obtained without departing from the spirit and scope of this disclosure.
[0200] Example 1: Metagenomic Analysis Methods for Proteins
[0201] Hidden Markov models based on known Cas protein sequences (including type II Cas effector proteins) were used to search metagenomic sequence data from publicly available databases. Potential active sites were identified by comparing the retrieved CRISPR-Cas proteins with known proteins. After screening hundreds of potential sequences, this metagenomic workflow ultimately identified the type II Cas proteins, as detailed in Table 1.
[0202] Phylogenetic trees were generated using MUSCLE 3 (Veen et al., 2020) to explore the relationships of orthologs at the primary amino acid level. This study utilized hundreds of Class 2 Type II-A / B / C sequences from the National Center for Biotechnology Information (NCBI) and various publications and patents. Notably, the phylogenetic tree indicates that the Cas protein described in this invention is distinct from previously known Cas proteins. The Type II Cas protein detailed in this disclosure shows low similarity to other known Cas proteins.
[0203] The structure of the type II Cas protein was modeled using AlphaFold2. Annotation of its domain arrangement revealed that the Cas protein disclosed in this invention comprises a RuvC domain, a BH (bridged helix) domain, a REC domain, an HNH domain, and a CTD (C-terminal domain). Notably, the RuvC domain contains three distinct RuvC subdomains and the BH domain. Figure 1 Visualizations of some example protein structures are shown.
[0204] Example 2: Experimental Protocol for Predicting RNA Folding
[0205] The predicted RNA folding of the putative guide RNA sequence matching the Cas protein was calculated using the RNAfold network server developed by Lorenz et al. in 2011.
[0206] Example 3: Identification of PAM in mammalian cell lines
[0207] In one set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 450 µL of cells at a density of 120,000 cells / well were mixed with 50 μL of a mixture containing Lipofectamine™ 3000 (ThermoFisher Scientific, catalog number L3000008), Opti-Mem (volume to 50 μL), 1 μL dsODN (2.5 pM), 100 ng (approximately 1 μL) of psgRNA carrying an sgRNA scaffold (Table 6) and the Humanspacer3 spacer sequence (SEQ ID NO: 885, Table 4), and 400 ng (approximately 1 μL) of pCas protein particles carrying the Cas9 CDS, NLS, and FLAG coding sequences (Table 2) according to the manufacturer's protocol. The cell mixture was then seeded into 24-well plates and cultured at 37°C and 5% CO2. 10 pM dsODN was prepared by annealing dsODN-Top and dsODN-BoT oligonucleotides before transfection.
[0208] 72 hours after transfection, the supernatant was removed and the cell layer was washed with PBS. Genomic DNA was then extracted from each well of a 24-well plate using DNA extraction solution (Denogen (Beijing) Biotechnology Co., Ltd., catalog number DNS033-48) according to the manufacturer's protocol. All DNA samples (500 ng, 260 / 280 value: 1.8–2.0) were analyzed by Guide-Seq NGS.
[0209] The basic method for Guide-Seq library preparation was performed according to Nikolay et al. (Nat. Protoc. 2021). Extracted DNA samples were first cleaved using a KAPA Frag Kit (Catalog No. KK8602, Roche). The cleaved DNA was purified and then phosphorylated using T4 polynucleotide kinase (Catalog No. M0201S, NEB). SS5 adapters (generated by annealing 10 μM SS5TOP oligonucleotides and 10 μM SS5BTM oligonucleotides) were ligated to the cleaved DNA using a Quick Ligation™ Kit (Catalog No. M2200S, NEB), followed by two-step off-target PCR to add the chemical modifications required for sequencing.
[0210] Off-target PCR 1 was performed using Platinum™ Taq DNA polymerase (catalog number 15966005, Invitrogen) with GSP1 (a mixture of GSP1-Top and GSP1-BoT) and Y_XX oligonucleotides. Off-target PCR 2 was performed using Platinum™ Taq DNA polymerase with GSP2 (a mixture of GSP2-TopA / B / C and GSP1-BoTA / B / C), Y_XX (the same as PCR 1), and i753_XX oligonucleotides. DNA products from each step were purified using SPRI Select (catalog number B23318, Beckman Coulter). The final library was quantified by qPCR and sequenced on an Illumina NextSeq 1000. Reads were aligned to a reference genome after eliminating low-quality fractions. The Q30 rate was greater than 0.9. Read lengths were between 130 bp and 140 bp. The generated file containing the read data is mapped to a reference genome (BAM file), where reads that overlap with the target region are selected.
[0211] Table 4: Nucleotide sequences mentioned in the examples.
[0212]
[0213] Note: "p" indicates phosphorylation modification; " " indicates a phosphate thioester (PS) bond; "N" indicates any natural or non-natural nucleotide.
[0214] Figure 2 and Table 5 show the PAM preferences of some example Cas proteins when using the corresponding sgRNA scaffolds in the HEK293 cell line.
[0215] Table 5: PAM Preferences of Example Cas Proteins
[0216] Example 4: Screening the in vitro editing efficiency of the CRISPR-Cas system in mammalian cell lines
[0217] In one set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 24 hours prior to transfection, 250 μL of cells at a density of 50,000 cells / well were seeded into 48-well plates. Cells were transfected using a lipid complex containing Lipofectamine™ 3000 (0.4 μL / well), P3000 (2 μL / well), pgRNA / pCas protein particles (125 ng / well and 375 ng / well, respectively), and Opti-Mem (to a final volume of 25 μL / well), following the manufacturer's protocol. Seeded cells were then statically and adherently cultured in a tissue culture incubator at 37°C and 5% CO2 for 72 hours. The nucleotide sequences of the pgRNAs used in this embodiment include sequences encoding the corresponding Cas sgRNA scaffolds (Table 6; SEQ ID NO: 465 (GEBx0305), SEQ ID NO: 451 (GEBx0308), SEQ ID NO: 454 (GEBx0308-Rq-V3), SEQ ID NO: 478 (spCas9), SEQ ID NO: 491 and SEQ ID NO: 548) and the corresponding spacer sequences (Table 7).
[0218] 72 hours after transfection, the supernatant was removed and the cell layer was washed with PBS. Genomic DNA was then extracted from each well of a 24-well plate using DNA extraction solution (Denogen (Beijing) Biotechnology Co., Ltd., catalog number DNS033-48) according to the manufacturer's protocol. All DNA samples (500 ng, 260 / 280 value: 1.8–2.0) were subjected to amplicon NGS analysis.
[0219] To quantitatively determine the editing efficiency at target locations in the genome, NGS was used to identify insertions and deletions introduced by gene editing. Primers for NGS were designed to be positioned around the target region of the endogenous gene. Additional PCR was performed according to the manufacturer's protocol (Illumina) to add the chemical modifications required for sequencing. Amplicon sequencing was performed on an Illumina iSeq 100. Reads were aligned to a reference genome after removing low-quality fractions. The Q30 rate was greater than 0.9. Read lengths were between 130 bp and 140 bp. The resulting file containing read data was mapped to a reference genome (BAM file), where reads overlapping with the target region were selected, and the number of wild-type reads and reads containing insertions, substitutions, or deletions were calculated. The number of reads mapped to the reference genome exceeded 1000.
[0220] Table 6: sgRNA scaffold sequences of corresponding Cas proteins
[0221] The spacer sequence is located at the 5' end of the stent. The 3' end of each spacer sequence is directly connected to the 5' end of the subsequent stent sequence, forming a typical repeat-spacer pattern.
[0222] Table 7: Example Interval Sequences
[0223] Figure 3 The levels of insertion / deletion (indel) of GEBx0305 targeting 16 sites with GGAAAA-PAM were shown in the HEK293T cell line. The sgRNA sequences used for GEBx0305 in this experiment included the GEBx0305-HPT-V1 scaffold (WT, SEQ ID NO: 465) and 20nt spacer sequences (SEQ ID NO: 571, 573, 575, 577, 579, 581, 583, 585, 587, 589, 591, 593, 595, 597, 599, and 601). SpCas9 targeting the corresponding sites (SEQ ID NO: 572, 574, 576, 578, 580, 582, 584, 586, 588, 590, 592, 594, 596, 598, 600, and 602) served as positive controls. Figure 4The insertion and deletion levels of GEBx0308 targeting 19 sites with GGTACT-PAM were shown in the HEK293T cell line. The sgRNA sequences used for GEBx0308 in this experiment contained the GEBx0308-PT-V1 scaffold (WT, SEQ ID NO: 451) and 20nt spacer sequences (SEQ ID NO: 611, 613, 615, 617, 619, 621, 623, 625, 627, 629, 631, 633, 635, 637, 639, 641, 643, 645, 647, and 649). SpCas9 targets the corresponding sites (SEQ ID NO: 612, 614, 616, 618, 620, 622, 624, 626, 628, 630, 632, 634, 636, 638, 640, 642, 644, 646, 648 and 650) as positive controls.
[0224] Figure 5 The insertion and deletion levels of GEBx0308 at 15 CATACT-PAM targets in the HEK293T cell line were shown. The psgRNA sequence used for GEBx0308 in this experiment contained the GEBx0308-Rq-V3 scaffold (M0, SEQ ID NO:454) and a 20nt spacer sequence (SEQ ID NO: 659-674).
[0225] Figure 6 The insertion / deletion levels of GEBx0328 at 20 target sites with NNGCCT-PAM were shown in the HEK293T cell line. The sgRNA sequence used for GEBx0328 in this experiment contained the GEBx0328-PT-V1 scaffold (WT, SEQ ID NO:491) and a 20nt spacer sequence (SEQ ID NO:765-784). GEBx0328 exhibited moderate insertion / deletion activity at the 20 target sites.
[0226] Figure 7 and Figure 8 The insertion / deletion levels of GEBx0361 at 19 target sites with GGTACC or TGTACC PAM were shown in the HEK293T cell line. The sgRNA sequence used for GEBx0361 in this experiment contained the GEBx0361-HPT-V2 scaffold (SEQ ID NO: 548) and a 20nt spacer sequence (SEQ ID NO: 817-835). GEBx0361 exhibited moderate insertion / deletion activity at the 19 target sites.
[0227] Example 5: Optimizing an Example CRISPR / Cas System
[0228] To further improve the editing efficiency of the Cas protein, guide sequences of different lengths and sgRNA scaffolds were tested. In one set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 24 hours before transfection, 250 μL of cells at a density of 50,000 cells / well were seeded into 48-well plates. Cells were transfected using a lipid complex containing Lipofectamine™ 3000 (0.4 μL / well), P3000 (2 μL / well), pgRNA / pCas protein particles (125 ng / well and 375 ng / well, respectively), and Opti-Mem (volume adjusted to 25 μL / well) according to the manufacturer's protocol. Seeded cells were then statically and adherently cultured in a tissue culture incubator at 37°C and 5% CO2 for 72 hours. The nucleotide sequence of the pgRNA used in this embodiment includes the sequence encoding the corresponding sgRNA scaffold (SEQ ID NO: 465 (GEBx0305), SEQ ID NO: 451 (GEBx0308) or SEQ ID NO: 491 (GEBx0328)) and the corresponding spacer sequence (SEQ ID NO: 575, 599 or 695-707 for GEBx0305; SEQ ID NO: 625, 676 or 708-721 for GEBx0308; SEQ ID NO: 769, 782, 785-798 for GEBx0328).
[0229] like Figure 9 As shown, the insertion / deletion level of GEBx0305 varied with the length of the guide sequence. The sgRNA sequences used in this experiment contained a GEBx0305-HPT-V1 scaffold (WT, SEQ ID NO: 465) and either a CFTR-NGGAAAA-T3 or POLQ-NGGAAAA-T5 spacer sequence, ranging in length from 18 nt to 25 nt (SEQ ID NO: 575, 599, or 695-707). The CFTR-NGGAAAA-T3-23 nt and POLQ-NGGAAAA-T5-25 nt spacer sequences showed the highest insertion / deletion levels in each group.
[0230] Figure 10(A and B) show that the insertion / deletion level of GEBx0308 varies with the length of the guide sequence. The sgRNA sequences used in this experiment contained the GEBx0308-PT-V1 scaffold (WT, SEQ ID NO: 451) and either the CFTR-NGGTACT-T4 or CD34-TATACT-T2 spacer sequences, ranging in length from 18 nt to 25 nt (SEQ ID NO: 625, 676, or 708-721). The CFTR-NGGTACT-T4-21 nt spacer sequence showed a higher insertion / deletion level than the spacer sequences of other lengths.
[0231] Figure 11 (A and B) show that the insertion / deletion levels of GEBx0328 varied with the length of the guide sequence. The sgRNA sequences used in this experiment contained the GEBx0328-PT-V1 scaffold (WT, SEQ ID NO: 491) and either the TTR-NGGCCT-T1 or CFTR-NGGCCT-T5 spacer sequences, ranging in length from 18 nt to 25 nt (SEQ ID NO: 769, 782, 785-798). The TTR-NGGCCT-T1-24 nt and CFTR-NGGCCT-T5-24 nt spacer sequences showed the highest insertion / deletion levels in each group.
[0232] In another set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 24 hours prior to transfection, 100 μL of cells at a density of 25,000 cells / well were seeded into 96-well plates. Cells were transfected using a lipid complex containing Lipofectamine™ 3000 (0.4 μL / well), P3000 (2 μL / well), pCas protein-gRNA plasmid (300 ng / well), and Opti-Mem (to a final volume of 25 μL / well) according to the manufacturer's protocol. The seeded cells were then statically and adherently cultured in a tissue culture incubator at 37°C and 5% CO2 for 72 hours. The nucleotide sequence of the Cas protein-gRNA used in this embodiment consists of Cas CDS, Cas sgRNA scaffold (Table 6), and corresponding spacer sequences (Table 7).
[0233] In another set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 24 hours prior to transfection, 100 μL of cells at a density of 25,000 cells / well were seeded into 96-well plates. Cells were transfected using a lipid complex containing Lipofectamine™ 3000 (0.4 μL / well), P3000 (2 μL / well), pCas protein-gRNA plasmid (300 ng / well), and Opti-Mem (to a final volume of 25 μL / well) according to the manufacturer's protocol. The seeded cells were then statically and adherently cultured in a tissue culture incubator at 37°C and 5% CO2 for 72 hours. The nucleotide sequence of the Cas protein-gRNA used in this example consists of Cas CDS, a Cas sgRNA scaffold, and corresponding spacer sequences.
[0234] Figure 12 (A and B) show the insertion / deletion levels of GEBx0305 using modified RNA scaffolds when targeting endogenous genes. The sgRNA sequences used in this experiment contained scaffolds of GEBx0305-HPT-V1 (WT, SEQ ID NO: 465), GEBx0305-Rq-V2 (M0, SEQ ID NO: 466), or GEBx0305-M1 to M6 (SEQ ID NO: 467-472), along with 20nt spacer sequences targeting eight endogenous gene sites (SEQ ID NO: 575, 579, 585, 587, 589, 591, 599, and 601). The GEBx0305-M5 and M6 scaffolds showed the highest mean insertion / deletion levels.
[0235] Figure 13 (A and B) show the insertion / deletion levels of GEBx0308 using modified RNA scaffolds when targeting endogenous genes. The sgRNA sequences used in this experiment contained scaffolds of GEBx0308-PT-V1 (WT, SEQ ID NO: 451), GEBx0308-Rq-V3 (M0, SEQ ID NO: 454), or GEBx0308-M1 to M6 (SEQ ID NO: 455-460), along with a 20nt spacer sequence, targeting seven endogenous gene sites (SEQ ID NO: 621, 625, 627, 631, 637, 639, 641). The GEBx0308-M3 and M4 scaffolds showed the highest mean insertion / deletion levels.
[0236] Figure 14(A and B) show the insertion and deletion levels of GEBx0305 targeting endogenous genes under optimized conditions. The optimized sgRNA sequence used in this experiment contained a GEBx0305-M5 (SEQ ID NO: 471) scaffold and a 21nt spacer sequence, targeting 27 endogenous gene sites (SEQ ID NO: 652, 654, 655, 658, 696, 703, 722-725, 729-745). Compared with WT sgRNA (GEBx0305-HPT-V1 scaffold (SEQ ID NO: 465) and 20nt spacer sequences (SEQ ID NO: 571, 573, 575, 577, 579, 581, 583, 585, 587, 589, 591, 593, 595, 597, 599, 651-658, 746-749)), optimized sgRNA significantly improved the insertion and deletion levels of GEBx0305 (P<0.0001).
[0237] Figure 15 (A and B) show the insertion and deletion levels of GEBx0308 targeting endogenous genes under optimized conditions. The optimized sgRNA sequence used in this experiment contained a GEBx0308-M4 (SEQ ID NO: 458) scaffold and a 21nt spacer sequence, targeting 25 endogenous gene sites (SEQ ID NO: 710, 659, 665, 667, 668, 672, 674, 726, 727, 728, 750-764). Compared with WT sgRNA (GEBx0308-PT-V1 scaffold (SEQ ID NO: 451) and 20nt spacer sequences (SEQ ID NO: 611, 613, 615, 617, 619, 621, 623, 625, 627, 629, 631, 633, 635, 637, 639, 641, 645, 647, 659, 662, 663, 665, 667, 668, 672, 674)), optimized sgRNA slightly improved the insertion and deletion levels of GEBx0308 (P=0.03).
[0238] Figure 16(A and B) show the insertion / deletion levels of GEBx0328 using modified RNA scaffolds when targeting endogenous genes. The sgRNA sequences used in this experiment contained scaffolds of GEBx0328-PT-V1 (WT, SEQ ID NO: 491), GEBx0328-Rq-V1 (M0, SEQ ID NO: 492), or GEBx0328-M1 to M8 (SEQ ID NO: 493-500), along with 20nt spacer sequences, targeting eight endogenous gene sites (SEQ ID NO: 765, 767, 769, 770, 771, 776, 782, and 783). The GEBx0328-M6 scaffold showed the highest mean insertion / deletion level.
[0239] Figure 17 (A and B) show the insertion and deletion levels of GEBx0328 targeting endogenous genes under optimized conditions. The optimized sgRNA sequence used in this experiment contained a GEBx0328-M6 (SEQ ID NO: 498) scaffold and a 24nt spacer sequence, targeting 19 endogenous gene sites (SEQ ID NO: 790, 797, 799-807, 809-816). Compared with WT sgRNA (GEBx0328-PT-V1 scaffold and 20nt spacer sequence), the optimized sgRNA significantly improved the insertion and deletion levels of GEBx0328 (P=0.0002).
[0240] Example 6: In vitro gene editing using RNA. In one set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 24 hours prior to transfection, 100 μL of cells at a density of 25,000 cells / well were seeded into 96-well plates. Following the manufacturer's protocol, cells were transfected using a lipid complex containing Lipofectamine™ RNAiMAX (Invitrogen™) and RNA (20 ng sgRNA, sgRNA:mRNA = 1:1 - 1:8 w / w), and Opti-Mem (volume supplemented to 25 μL / well). The seeded cells were statically and adherently cultured in a tissue culture incubator at 37°C and 5% CO2 for 72 hours. The sgRNA and mRNA sequences are shown in Table 8. The mRNA used in this example was N1-methyl-pseudouridine-modified uracil.
[0241] Primary human hepatocytes (PHH) were thawed and resuspended in hepatocyte thawing medium (Lonza, catalog number MCHT50) supplemented with additives, followed by centrifugation at 100 g for 10 minutes. The supernatant was discarded, and the cell pellet was resuspended in hepatocyte plating medium (Lonza, catalog number MP100) with 10% fetal bovine serum added. After cell counting, the cells were seeded at a density of 40,000 cells / well in 96-well ultra-low adsorption cell culture plates (Liver Biotech, catalog number LV-ULA002-96W). The seeded cells were incubated statically and adherently in a tissue culture incubator at 37°C and 5% CO2 for 24 hours. After incubation, the cell monolayer formation was checked, and the medium was replaced with hepatocyte culture medium (Lonza, catalog number CC-3198) with 10% fetal bovine serum added. Following the manufacturer's protocol, cells were transfected using a lipid complex containing Lipofectamine™ RNAiMAX (Invitrogen™) and RNA (30 ng gRNA, gRNA:mRNA = 1:1 - 1:16 w / w) and Opti-Mem (volume adjusted to 25 μL / well). The seeded cells were then statically incubated in a 37°C, 5% CO2 tissue culture incubator for 72 hours.
[0242] The genomic DNA extraction and NGS methods are the same as those described in Example 4.
[0243] Figure 18 This study demonstrates the insertion / deletion levels of GEBx0305 targeting endogenous genes after transfection of a lipid complex containing a fixed amount (20 ng) of sgRNA targeting the POLQ-NGGAAAA-T5 site (SEQ ID NO: 855) and varying proportions of GEBx0305 mRNA (SEQ ID NO: 851) into HEK293T cells. Equal doses of SpCas9 mRNA and sgRNA (SEQ ID NO: 853 and 854) served as positive controls. GEBx0305 exhibited insertion / deletion efficiency comparable to SpCas9.
[0244] Figure 19 The insertion and deletion levels of GEBx0308 targeting endogenous genes were shown after transfection of a lipid complex containing a fixed amount (20 ng) of sgRNA targeting CFTR-NGGTACT-T4 or POLQ-NGGTACT-T1 sites (SEQ ID NO: 856 / 857) and different proportions of GEBx0308 mRNA (SEQ ID NO: 852) in HEK293T cells.
[0245] Figure 20This study demonstrates the insertion / deletion levels of GEBx0305 targeting endogenous genes after transfection of a lipid complex containing a fixed amount (20 ng) of sgRNA targeting the POLQ-NGGAAAA-T5 site (SEQ ID NO: 855) and different proportions of GEBx0305 mRNA (SEQ ID NO: 851) into primary human hepatocytes (PHH). Equal doses of SpCas9 mRNA and sgRNA (SEQ ID NO: 853, 854, Table 9) served as positive controls.
[0246] Figure 21 The insertion and deletion levels of GEBx0308 targeting endogenous genes were shown after transfection of a lipid complex containing a fixed amount (20 ng) of sgRNA targeting the POLQ-NGGTACT-T1 site (SEQ ID NO: 857) and different proportions of GEBx0308 mRNA (SEQ ID NO: 852, Table 8) into primary human hepatocytes (PHH).
[0247] Table 8: Example mRNA and gRNA sequences used for gene editing
[0248] Example 7: Off-target analysis in cell lines using GUIDE-Seq
[0249] GUIDE-Seq utilizes dsODN to insert double-strand break sites generated by CRISPR / Cas. HEK293T cells were cultured in advanced DMEM medium supplemented with 5% fetal bovine serum (Gibco™). Twenty-four hours prior to transfection, cells were seeded into 24-well plates at a density of 100,000 cells / well. Following the manufacturer's protocol, 400 ng pCas protein grains, 150 ng pgRNA plasmids, and 2.5 pmol dsODN were transfected using Lipofectamine 3000 (Invitrogen™), and the cells were cultured at 37°C and 5% CO2. Cells were harvested on day 3 post-transfection.
[0250] For GUIDE-Seq library construction, 500 ng of genomic DNA was used. In short, DNA was fragmented using the KAPA Frag Kit (KAPA Biosystems), followed by aptamer ligation and two rounds of semi-nested PCR enrichment of dsODN-integrated fragments. The final sequencing library was quantified using KAPA Library Quantification Kits and sequenced on an Illumina NextSeq 1000 system. Index 1 data demultiplexing was performed using bcl2fq (version 2.19), followed by Index 2 demultiplexing using a custom script, aptamer trimming using the BBduk tool, and analysis using GUIDE-seq software. In short, FASTQ files with unique molecular indexes (UMIs) were integrated to generate UMI-consistent sequences and aligned to the human reference genome (hg19) using BWA MEM. High-quality alignments (MAPQ ≥50) were used to identify genomic loci containing dsODNs as potential off-target sites. Up to six candidate sites that mismatched with the corresponding target original spacer sequence were identified as true off-target sites.
[0251] Figure 22 The GUIDE-seq insertion sites of GEBx0305 are shown. No detectable off-target sites were detected at site 1 (CFTR-NGGAAAA-T5) and site 2 (EMX1-NGGAAAA-T5).
[0252] Figure 23 The GUIDE-seq insertion sites of GEBx0308 are shown. No off-target sites were detected at site 1 (CD34-NGGTACT-T4) and site 2 (POLQ-NGGTACT-T1).
[0253] Figure 24 The GUIDE-seq insertion sites of GEBx0328 are shown. No detectable off-target sites were detected at site 2 (CFTR-NGGCCT-T5), while only one off-target site was detected at site 1 (CFTR-NGGCCT-T3).
[0254] Example 8: Determination of base editing efficiency of nCas protein
[0255] To generate the Cas base editor, *E. coli* tRNA adenosine deaminase (TadA-8e) was fused to the N-terminus of GEBx0305, GEBx0308, and GEB328 nickases (GEBx0305-D11A, GEBx0308-D10A, and GEB328-D12A) and ligated using an ABE-connector (SEQ ID NO: 868). HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 24 hours prior to transfection, 100 μL of cells at a density of 25,000 cells / well were seeded into 96-well plates. Following the manufacturer's protocol, cells were transfected using a lipid complex containing Lipofectamine™ 3000 (0.4 μL / well), P3000 (2 μL / well), pABE-gRNA plasmid (300 ng / well), and Opti-Mem (volume adjusted to 25 μL / well). Seeded cells were incubated statically and adherently in a tissue culture incubator at 37°C and 5% CO2 for 72 hours. The nucleotide sequence of the pABE / gRNA plasmid used in this example contained Cas-ABE CDS (SEQ ID NO: 861, 862, 863, Table 9), a Cass gRNA scaffold, and corresponding spacer sequences. After 72 hours of transfection, the supernatant was removed, and the cell layer was washed with PBS. Genomic DNA was then extracted from each well of a 24-well plate using DNA extraction solution (Denogen (Beijing) Biotechnology Co., Ltd., catalog number DNS033-48) according to the manufacturer's protocol. All DNA samples were subjected to amplicon NGS analysis.
[0256] Table 9: CDS sequences of example ABE sequences.
[0257]
[0258] Figure 25This study demonstrated the base editing efficiency of GEBx0305-ABE in performing A-to-G conversions of adenine at five endogenous gene sites in HEK293T cells. The sgRNA sequences used in this experiment contained the GEBx0305-M5 (SEQ ID NO: 471) scaffold and 21nt spacer sequences (SEQ ID NO: 703, 722, 723, 725). GEBx0305-ABE exhibited highly efficient A-to-G conversions at these sites.
[0259] Figure 26 This study demonstrates the base editing efficiency of GEBx0308-ABE in performing A-to-G conversions of adenine at five endogenous gene sites in HEK293T cells. The sgRNA sequences used in this experiment comprise the GEBx0308-M4 (SEQ ID NO: 458) scaffold and 21nt spacer sequences (SEQ ID NO: 710, 726, 727, 728, 667). GEBx0308-ABE exhibited highly efficient A-to-G conversions at these sites.
[0260] Figure 27 This study demonstrates the base editing efficiency of GEBx0328-ABE in performing A-to-G conversions of adenine at five endogenous gene sites in HEK293T cells. The sgRNA sequences used in this experiment comprise the GEBx0328-M6 (SEQ ID NO: 498) scaffold and 24nt spacer sequences (SEQ ID NO: 790, 797, 801, 809, 815). GEBx0328-ABE exhibited highly efficient A-to-G conversions at these sites.
[0261] Example 8: Evaluation of the specific knockout effect of GEBx0308 on the RHO-P23H site in HEK293T cells. Editor plasmid and target plasmid design: An editor plasmid containing the encoding gene of SaCas9 (SEQ ID NO: 906) or GEBx0308 (SEQ ID NO: 252) and the corresponding guide sequences (RHO-P23H sgRNA (SaCas9), i.e., the single guide RNA of SaCas9, sequence SEQ ID NO: 904; and RHO-P23H sgRNA, i.e., the single guide RNA of GEBx0308, sequence SEQ ID NO: 902) was designed and generated. A target plasmid containing RHO-WT (approximately 170 bp), RHO-P23H (approximately 170 bp, c.68C>A), and a 500 bp irrelevant sequence (used to separate RHO-WT and RHO-P23H) was designed and generated.
[0262] To evaluate the specific knockout effect of GEBx0308 on the RHO-P23H site in HEK293T cells, HEK293T cells were cultured in Dulbecco modified Eagle medium (DMEM, CORNING) containing 2 mM L-glutamine (GlutaMAX™-l, Gibco) and 10% fetal bovine serum (FETAL BOVINE SERUM, GEMINI) at 37°C in a 5% CO2 buffer incubator.
[0263] To evaluate the specific knockout efficiency of SaCas9 (SEQ ID NO: 906, Table 10) and GEBx0308 (SEQ ID NO: 252) at the RHO-P23H site, editor and target plasmids were co-transfected into HEK293T cells. 1.0 × 10^5 cells were seeded into each well of a 24-well plate. Following the manufacturer's instructions (Life Technologies), editor and target plasmids were co-transfected into HEK293T cells at 2:1 and 3:1 (w / w) ratios using Lipofectamine 3000 reagent. After 72 hours of incubation, cells were harvested and lysed to release the target plasmid. The editing efficiency of the RHO-WT and RHO-P23H sites was analyzed by NGS using target plasmid-specific primer pairs and the potentially editable target plasmid as a template. Results are as follows: Figure 28 As shown.
[0264] Table 10: Example sequences associated with RHO-P23H site-specific knockout.
[0265]
[0266] While preferred embodiments of the invention have been shown and described herein, these embodiments are provided by way of example only and will be apparent to those skilled in the art. The invention is not intended to be limited to the specific embodiments provided in the specification. Although the invention has been described with reference to the foregoing specification, the description and illustration of embodiments herein should not be construed as limiting. Various changes, modifications, and substitutions can now be made by those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the invention are not limited to the specific descriptions, configurations, or relative proportions presented herein, as these depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in carrying out the invention. Therefore, the invention should also cover any such alternatives, modifications, variations, or equivalents. The scope of the invention is intended to be defined by the following claims and to cover all methods and structures within the scope of these claims and their equivalents.
Claims
1. An engineered, non-naturally occurring type II CRISPR-associated (Cas) protein or a variant thereof, having at least 70% sequence identity with any amino acid sequence in SEQ ID NO:1-71.
2. The Cas protein according to claim 1, wherein the sequence identity of the Cas protein with any amino acid sequence in SEQ ID NO: 1-71 is at least 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%.
3. An engineered, non-naturally occurring type II CRISPR-associated (Cas) protein, wherein, except for amino acid "M" at position 1 of the sequence, the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with any amino acid sequence in SEQ ID NO: 1-71.
4. The Cas protein according to any one of claims 1-3, wherein the sequence identity of the Cas protein with any amino acid sequence of SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60 and 68 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%.
5. The Cas protein according to any one of claims 1-3, wherein, Except for amino acid "M" at position 1 in the sequence, the Cas protein has sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with any of the amino acid sequences in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, and 68.
6. The Cas protein according to any one of claims 1-5, wherein the Cas protein is capable of recognizing a group consisting of the following sequence adjacent motifs (PAMs): NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, and NGGHA.
7. The Cas protein according to any one of claims 1-6, wherein: (1) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 9, and is able to recognize PAM with the sequence NRRANH; (2) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 12, and is able to recognize PAM with the sequence NRHACT; (3) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 19, and is able to recognize PAM with the sequence NRAAR; (4) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 9, and is able to recognize PAM with the sequence NRAAR; The amino acid sequence identity of 21 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NNNCCY; (5) The amino acid sequence identity of Cas protein with SEQ ID NO: 22 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NNRYYYY; (6) The amino acid sequence identity of Cas protein with SEQ ID NO: 24 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NGG; (7) The amino acid sequence identity of Cas protein with SEQ ID NO: The amino acid sequence identity of 25 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NNNCAA; (8) The amino acid sequence identity of Cas protein with SEQ ID NO: 27 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NRNACN; (9) The amino acid sequence identity of Cas protein with SEQ ID NO: 29 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NNGR; (10) The amino acid sequence identity of Cas protein with SEQ ID NO: The sequence identity of the amino acid sequence of 30 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%, and it is able to recognize PAM with the sequence NGGNR.(11) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 31, and is capable of recognizing PAM with the sequence NNNCCH; (12) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 36, and is capable of recognizing PAM with the sequence NRRAAG; (13) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 37, and is capable of recognizing PAM with the sequence NRHRAC; (14) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 37, and is capable of recognizing PAM with the sequence NRHRAC; The amino acid sequence identity of 38 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the NRYART sequence; (15) The amino acid sequence identity of the Cas protein with SEQ ID NO: 43 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the NRHACC sequence; (16) The amino acid sequence identity of the Cas protein with SEQ ID NO: 44 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the NRAAR sequence; (17) The amino acid sequence identity of the Cas protein with SEQ ID NO: The amino acid sequence identity of 51 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NRNVHH; (18) The amino acid sequence identity of Cas protein with SEQ ID NO: 56 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence YMACAW; (19) The amino acid sequence identity of Cas protein with SEQ ID NO: 59 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NAHAA; (20) The amino acid sequence identity of Cas protein with SEQ ID NO: The amino acid sequence of 60 has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and is able to recognize PAM with the sequence NRHAYY.Or (21) the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 68, and is able to recognize PAM with the sequence NGGHA.
8. The Cas protein according to any one of claims 1-7, wherein the Cas protein is a cleavage enzyme or an inactivated Cas protein.
9. The Cas protein according to claim 8, wherein the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 9, and has a mutation at residue D11 or H859; or has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 12, and has a mutation at residue D10 or H862; or has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 31, and has a mutation at residue D12 or H903.
10. The Cas protein according to claim 8, wherein, except for amino acid "M" at position 1 in the sequence, the sequence identity of the Cas protein with the amino acid sequence of SEQ ID NO: 9 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and there is a mutation at residue D11 or H859; or except for amino acid "M" at position 1 in the sequence, the sequence identity with the amino acid sequence of SEQ ID NO: 12 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and there is a mutation at residue D10 or H862; or except for amino acid "M" at position 1 in the sequence, the sequence identity with the amino acid sequence of SEQ ID NO: 31 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and there is a mutation at residue D12 or H903.
11. The Cas protein according to any one of claims 8-10, wherein the mutation at residue D11 or H859 of SEQ ID NO: 9 is D11A or H859A; the mutation at residue D10 or H862 of SEQ ID NO: 12 is D10A or H862A; or the mutation at residue D12 or H903 of SEQ ID NO: 31 is D12A or H903A.
12. The Cas protein according to any one of claims 8-11, wherein the sequence identity of the Cas protein with any amino acid sequence in SEQ ID NO: 869-877 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%.
13. The Cas protein according to any one of claims 8-12, wherein the Cas protein further comprises a sequence selected from the following sequence group: nuclear localization signal sequence, nuclear exit signal sequence, cell penetration peptide sequence, affinity tag sequence, deaminase sequence, reverse transcriptase sequence, recombinase sequence, methyltransferase sequence, methyltransferase sequence, acetyltransferase sequence, acetyltransferase sequence, transcription activator sequence, transcription repressor domain sequence, cryptochrome sequence, photoinducible / controllable domain sequence, and chemically inducible / controllable domain sequence.
14. The Cas protein according to any one of claims 1-13, wherein the Cas protein comprises an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the amino acid sequences in SEQ ID NO: 81-151 and 864-866.
15. An engineered, non-naturally occurring polynucleotide encoding a type II CRISPR-associated (Cas) protein as described in any one of claims 1-14.
16. The polynucleotide of claim 15, wherein the polynucleotide is a ribonucleotide sequence or a deoxyribonucleotide sequence, or an analogue thereof; optionally, the polynucleotide is codon-optimized for expression in cells of interest; preferably, the polynucleotide is mRNA and further comprises a 5' cap sequence and / or a poly-A tail sequence.
17. The polynucleotide of claim 16, wherein the polynucleotide is codon-optimized for expression in eukaryotic cells; optionally, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and / or human cells.
18. The polynucleotide according to any one of claims 15-17, wherein the sequence identity of the polynucleotide with any of the nucleotide sequences in SEQ ID NO: 161-231, 241-311, 851-852 and 861-863 is at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99% or 100%.
19. According to an engineered, non-naturally occurring CRISPR-Cas system, comprising: a) The type II Cas protein or the polynucleotide encoding the Cas protein as described in any one of claims 1-14; b) at least one engineered guide RNA or at least one engineered nucleic acid encoding the guide RNA, wherein the guide RNA comprises a spacer sequence complementary to the target nucleic acid and a Cas protein binding segment that interacts with the Cas protein, wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer.
20. The system of claim 19, wherein the guide RNA further comprises a linker sequence connecting the tracrRNA sequence and the DR sequence to form the sgRNA backbone.
21. The system of claim 20, wherein the sgRNA backbone comprises a sequence that is at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identical to any one of the sequences in SEQ ID NO: 431-562.
22. The system according to any one of claims 19-21, wherein the polynucleotide encoding the Cas protein is operatively linked to a promoter; optionally, the promoter is a constitutive promoter, a tissue-specific promoter, or an inducible promoter.
23. The system of claim 22, wherein the polynucleotide encoding the Cas protein is operatively linked to a promoter and is present in a vector; optionally, the vector is selected from the group consisting of: retroviral vectors, lentiviral vectors, phage vectors, adenovirus vectors, adeno-associated virus vectors, herpes simplex virus vectors, and plasmid vectors.
24. An engineered vector comprising the polynucleotide of any one of claims 15-18, wherein the vector may be an inducible, conditional, or constitutive expression vector.
25. A vector system comprising one or more polynucleotides according to any one of claims 15-18 and one or more polynucleotides encoding a guide RNA; wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein binding segment that interacts with the Cas protein, wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer.
26. An engineered, non-naturally occurring cell, comprising: The Cas protein as described in any one of claims 1-14, the polynucleotide as described in any one of claims 15-18, the CRISPR-Cas system as described in any one of claims 19-23, the vector as described in claim 24, and the vector system as described in claim 25.
27. A cell modified with any one of the Cas protein as claimed in claims 1-14, any one of the polynucleotides as claimed in claims 15-18, any one of the CRISPR-Cas systems as claimed in claims 19-23, the vector as claimed in claim 24, or the vector system as claimed in claim 25.
28. A kit comprising: the Cas protein as described in any one of claims 1-14, the polynucleotide as described in any one of claims 15-18, the CRISPR-Cas system as described in any one of claims 19-23, the vector as described in claim 24, the vector system as described in claim 25, or the cell as described in claim 26 or 27.
29. A pharmaceutical composition comprising: the Cas protein as described in any one of claims 1-14, the polynucleotide as described in any one of claims 15-18, the CRISPR-Cas system as described in any one of claims 19-23, the vector as described in claim 24, the vector system as described in claim 25, or the cell as described in claim 26 or 27.
30. The pharmaceutical composition of claim 29, wherein the pharmaceutical composition further comprises a delivery system selected from the following options: AAV (adeno-associated virus), adenovirus, retrovirus, HSV (herpes simplex virus), Gammaretrovirus, LV (lentivirus), eCIS (extracellular contractile injection system), eVLPs (engineered virus-like particles), VLPs (virus-like particles), liposomes, plasmids, LNPs (lipid nanoparticles), exosomes, microvesicles, nucleic acid nanoassemblies, gene guns, and / or implantable devices.
31. A method for treating, preventing, diagnosing, or detecting a disease, comprising a Cas protein as claimed in any one of claims 1-14, a polynucleotide as claimed in any one of claims 15-18, a CRISPR-Cas system as claimed in any one of claims 19-23, a vector as claimed in claim 24, a vector system as claimed in claim 25, a cell as claimed in claim 26 or 27, a kit as claimed in claim 28, or a pharmaceutical composition as claimed in claim 29 or 30.
32. A method for modifying or targeting a target DNA site, comprising providing the site with the Cas protein as described in any one of claims 1-14, the polynucleotide as described in any one of claims 15-18, the CRISPR-Cas system as described in any one of claims 19-23, the vector as described in claim 24, or the vector system as described in claim 25.
33. The method of claim 32, wherein modifying or targeting the target DNA site comprises inducing DNA strand breaks, altering the gene expression of one or more genes, or epigenetically modifying the target DNA site; optionally, the DNA strand breaks comprise DNA double-strand breaks or DNA single-strand breaks.
34. The method according to any one of claims 32 or 33, wherein the method is performed in vitro or in vivo.
35. A method for targeting and cleaving double-stranded target DNA, comprising: The double-stranded target DNA is contacted with the Cas protein as described in any one of claims 1-14, the polynucleotide as described in any one of claims 15-18, the CRISPR-Cas system as described in any one of claims 19-23, or the pharmaceutical composition as described in claim 29 or 30.
36. An isolated eukaryotic cell comprising a modified target site, wherein the target site has been modified by the method of any one of claims 32-35, or using a pharmaceutical composition as described in claims 29 or 30, or using a CRISPR-Cas system as described in any one of claims 19-23.
37. A system for detecting the presence of a target nucleic acid sequence in an in vitro sample, comprising: a) The Cas protein as described in any one of claims 1-14; b) At least one guide polynucleotide containing a guide sequence capable of binding to the target sequence, designed to form a complex with the Cas protein; c) A nucleic acid-based masking structure containing a non-target sequence; The Cas protein exhibits incidental cleavage activity on RNA and / or ssDNA, and cleaves non-target sequences in nucleic acid-based masking structures activated by the target sequence.
38. A method for detecting a target nucleic acid in a sample, comprising: The sample is subjected to one or more samples with a) the Cas protein as described in any one of claims 1-14; b) At least one guide polynucleotide containing a guide sequence designed to be complementary to the target sequence to form a complex with the Cas protein; c) A nucleic acid-based masking structure contact containing a non-target sequence; The Cas protein exhibits incidental cleavage activity of RNA and / or ssDNA, and cleaves non-target sequences in nucleic acid-based masking structures activated by the target sequence; and detects signals from non-target sequence cleavage, thereby detecting one or more target sequences in the sample.
39. A guide RNA (gRNA) comprising: a) a spacer sequence of SEQ ID NO: 901; a spacer sequence having at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the sequence SEQ ID NO: 901; or a spacer sequence having at least 99%, at least 98%, at least 97%, at least 96%, at least 95%, at least 94%, at least 93%, at least 92%, at least 91% or at least 90% sequence identity with the sequence SEQ ID NO: 901; and b) a Cas protein binding segment; wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence, the DR sequence hybridizing to form a double-stranded RNA (dsRNA) duplex.
40. The gRNA according to claim 39, wherein the gRNA is a double-stranded gRNA or a single-stranded gRNA.
41. The gRNA of claim 40, wherein the gRNA has at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 98%, at least 99%, or 100% sequence identity with the nucleotide sequence of SEQ ID NO:
903.
42. The gRNA of claim 41, wherein at least three nucleotides of the gRNA are modified.
43. A polynucleotide encoding the gRNA of any one of claims 39-42.
44. An engineered, non-naturally occurring CRISPR-Cas system, comprising: a) a Cas protein or a polynucleotide encoding the Cas protein; b) at least one guide RNA (gRNA) or at least one engineered nucleic acid encoding the guide RNA as described in any one of claims 39-42, wherein the gRNA further comprises a Cas protein binding segment that interacts with the Cas protein; wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer.
45. The CRISPR-Cas system of claim 44, wherein the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence having at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identity with the sequence of SEQ ID NO:
903.
46. The CRISPR-Cas system according to claim 44 or 45, wherein the Cas protein comprises a sequence having at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identity with SEQ ID NO: 12; or except for the amino acid "M" at position 1 in the sequence, the Cas protein comprises a sequence having at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identity with SEQ ID NO:
12.
47. An engineered vector comprising the polynucleotide of claim 43, wherein the vector is optionally an inducible, conditional, or constitutive expression vector.
48. A carrier system comprising one or more polynucleotides as described in claim 43.
49. A pharmaceutical composition comprising gRNA as claimed in any one of claims 39-42, a polynucleotide as claimed in claim 43, a CRISPR-Cas system as claimed in any one of claims 44-46, a vector as claimed in claim 47, or a vector system as claimed in claim 48.
50. A method for treating, preventing, or diagnosing diseases associated with the RHO locus, comprising contacting target cells of a subject requiring treatment with a gRNA as described in any one of claims 39-42, a polynucleotide as described in claim 43, a CRISPR-Cas system as described in any one of claims 44-46, or a pharmaceutical composition as described in claim 49.