Type ii cas protein, crispr-cas system and uses thereof
Patent Information
- Application Number
- HK62026125860
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-06
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-08-15
Smart Images

Figure 00000444_0000 
Figure 00000445_0000 
Figure 00000445_0001
Abstract
Description
(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202480054010.5 (22) Application Date 2024.08.16 (66) Domestic Priority Data PCT / CN2023 / 113355 2023.08.16 CN PCT / CN2023 / 116757 2023.09.04 CN PCT / CN2023 / 136724 2023.12.06 CN PCT / CN2024 / 091198 2024.05.06 CN PCT / CN2024 / 091203 2024.05.06 CN PCT / CN2024 / 091211 2024.05.06 CN (85) PCT International Application Enters National Phase Date 2026.02.14 (86) PCT International Application Application Data PCT / CN2024 / 112724 2024.08.16 (87) PCT International Application Publication Data WO2025 / 036482 EN 2025.02.20 (71) Applicant Zhengji Gene Technology Co., Ltd. Address Room 101-102, 1 / F, 6W, Science Avenue West, Pak Shek Kok Science Park, New Territories, Hong Kong, China (72) Inventor Wang Bang (74) Patent Agency Shanghai Bisheng Intellectual Property Agency Co., Ltd. 31347 Patent Attorney Yuan Hong (51) Int.Cl. C12N 9 / 22 (2006.01) C12N 15 / 113 (2006.01) C12Q 1 / 682 (2006.01) C12Q 1 / 70 (2006.01) (54) Title of Invention Type II Cas Protein, CRISPR-Cas System and Uses Thereof (57) Abstract This disclosure relates to type II Cas protein, CRISPR-Cas system and various applications thereof. The type II Cas protein described in this disclosure expands the utility of the CRISPR-Cas system in targeting or modifying genes. Claims 5 pages Description 316 pages Sequence Listing (electronic publication) Drawings 31 pages CN 121729488 A 2026.03.24 CN 1 21 72 94 88 A 1. An engineered, non-naturally occurring type II CRISPR-related (Cas) protein or a variant thereof, having at least 70% sequence identity with any amino acid sequence in SEQ ID NO: 1-71. 2. The Cas protein according to claim 1, wherein the sequence identity of the Cas protein with any amino acid sequence in SEQ ID NO: 1-71 is at least 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%.3. An engineered, non-naturally occurring type II CRISPR-associated (Cas) protein, wherein, except for amino acid "M" at position 1 in the sequence, the Cas protein shares at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% sequence identity with any amino acid sequence in SEQ ID NO: 1-71. 4. The Cas protein according to any one of claims 1-3, wherein the Cas protein shares at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% sequence identity with any amino acid sequence in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, and 68. 5. The Cas protein according to any one of claims 1-3, wherein, except for amino acid "M" at position 1 in the sequence, the Cas protein has sequence identity with any of the amino acid sequences in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60 and 68 at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%. 6. The Cas protein according to any one of claims 1-5, wherein the Cas protein is capable of recognizing a group consisting of the following sequence adjacent motifs (PAMs): NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, and NGGHA. 7. The Cas protein according to any one of claims 1-6, wherein: (1) the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 9, and is capable of recognizing PAM with the sequence NRRANH; (2) the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 12, and is capable of recognizing PAM with the sequence NRHACT; (3) the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, or 92% with the amino acid sequence of SEQ ID NO: 19.95%, 98%, 99% or 100%, and capable of recognizing PAM with the NRAAR sequence; (4) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 21, and is capable of recognizing PAM with the NNNCCY sequence; (5) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 22, and is capable of recognizing PAM with the NNRYYYY sequence; (6) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 24, and is capable of recognizing PAM with the NGG sequence; (7) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 24, and is capable of recognizing PAM with the NGG sequence; (7) The amino acid sequence identity of 25 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NNNCAA; (8) The amino acid sequence identity of Cas protein with SEQ ID NO: 27 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NRNACN; (9) The amino acid sequence identity of Cas protein with SEQ ID NO: 29 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NNGR; (10) The amino acid sequence identity of Cas protein with SEQ ID NO: 30 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NNGR; The amino acid sequence identity of the Cas protein is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%, and it is able to recognize PAM with the sequence NGGNR; (11) The amino acid sequence identity of the Cas protein with SEQ ID NO: 31 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%, and it is able to recognize PAM with the sequence NNNCCH; (12) The amino acid sequence identity of the Cas protein with SEQ ID NO: 36 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%, and it is able to recognize PAM with the sequence NRRAAG; (13) The amino acid sequence identity of the Cas protein with SEQ ID NO: 36 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%, and it is able to recognize PAM with the sequence NRRAAG; (14) The amino acid sequence identity of the Cas protein with SEQ ID NO: 36 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%, and it is able to recognize PAM with the sequence NRRAAG;(14) The amino acid sequence identity of the Cas protein with SEQ ID NO: 38 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize the PAM with the sequence NRHRAC; (15) The amino acid sequence identity of the Cas protein with SEQ ID NO: 43 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize the PAM with the sequence NRYART; (16) The amino acid sequence identity of the Cas protein with SEQ ID NO: 44 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize the PAM with the sequence NRHACC; 99% or 100%, and capable of recognizing PAM with the NRAAR sequence; (17) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 51, and is capable of recognizing PAM with the NRNVHH sequence; (18) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 56, and is capable of recognizing PAM with the YMACAW sequence; (19) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 59, and is capable of recognizing PAM with the NAHAA sequence; (20) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 59, and is capable of recognizing PAM with the NAHAA sequence; The amino acid sequence identity of 60 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is capable of recognizing PAM with the sequence NRHAYY; or (21) the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 68, and it is capable of recognizing PAM with the sequence NGGHA. 8. The Cas protein according to any one of claims 1-7, wherein the Cas protein is a nicking enzyme or an inactivated Cas protein. 9. The Cas protein according to claim 8, wherein the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 9.The sequence identity with the amino acid sequence of SEQ ID NO: 12 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and there is a mutation at residue D11 or H859; or the sequence identity with the amino acid sequence of SEQ ID NO: 31 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and there is a mutation at residue D10 or H862; or the sequence identity with the amino acid sequence of SEQ ID NO: 31 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and there is a mutation at residue D12 or H903. 10. The Cas protein according to claim 8, wherein, except for amino acid "M" at position 1 in the sequence, the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 9, and has a mutation at residue D11 or H859; or except for amino acid "M" at position 1 in the sequence, the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 12, and has a mutation at residue D10 or H862; or except for amino acid "M" at position 1 in the sequence, the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 31, and has a mutation at residue D12 or H903. 11. The Cas protein according to any one of claims 8-10, wherein the mutation at residue D11 or H859 of SEQ ID NO: 9 is D11A or H859A; the mutation at residue D10 or H862 of SEQ ID NO: 12 is D10A or H862A; or the mutation at residue D12 or H903 of SEQ ID NO: 31 is D12A or H903A. 12. The Cas protein according to any one of claims 8-11, wherein the sequence identity of the Cas protein with any amino acid sequence in SEQ ID NO: 869-877 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%. 13. The Cas protein according to any one of claims 8-12, wherein the Cas protein further comprises a sequence selected from the following sequence group: nuclear localization signal sequence, nuclear exit signal sequence, cell penetration peptide sequence, affinity tag sequence, deaminase sequence, reverse transcriptase sequence, recombinase sequence, methyltransferase sequence, methyltransferase sequence, acetyl ...Transferase sequence, transcription activator sequence, transcription repressor domain sequence, cryptochrome sequence, photoinducible / controllable domain sequence, and chemically inducible / controllable domain sequence. Claims 2 / 5 pages 3 CN 121729488 A 14. The Cas protein according to any one of claims 1-13, wherein the Cas protein comprises an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identical to any one of the amino acid sequences in SEQ ID NO: 81-151 and 864-866. 15. An engineered, non-naturally occurring polynucleotide encoding a type II CRISPR-associated (Cas) protein according to any one of claims 1-14. 16. The polynucleotide of claim 15, wherein the polynucleotide is a ribonucleotide sequence or a deoxyribonucleotide sequence, or an analogue thereof; optionally, the polynucleotide is codon-optimized for expression in cells of interest; preferably, the polynucleotide is mRNA and further comprises a 5' cap sequence and / or a poly-A tail sequence. 17. The polynucleotide of claim 16, wherein the polynucleotide is codon-optimized for expression in eukaryotic cells; optionally, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and / or human cells. 18. The polynucleotide according to any one of claims 15-17, wherein the sequence identity of the polynucleotide with any nucleotide sequence of SEQ ID NO: 161-231, 241-311, 851-852, and 861-863 is at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100%. 19. According to an engineered, non-naturally occurring CRISPR-Cas system, comprising: a) a type II Cas protein or a polynucleotide encoding the Cas protein according to any one of claims 1-14; b) at least one engineered guide RNA or at least one engineered nucleic acid encoding the guide RNA, wherein the guide RNA comprises a spacer sequence complementary to the target nucleic acid and a Cas protein binding segment interacting with the Cas protein, wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer. 20. The system of claim 19, wherein the guide RNA further comprises a linker between the tracrRNA sequence and the DR sequence.21. The system of claim 20, wherein the sgRNA backbone comprises a sequence that is at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identical to any sequence in SEQ ID NO: 431-562. 22. The system of any one of claims 19-21, wherein the polynucleotide encoding the Cas protein is operatively linked to a promoter; optionally, the promoter is a constitutive promoter, a tissue-specific promoter, or an inducible promoter. 23. The system of claim 22, wherein the polynucleotide encoding the Cas protein is operatively linked to a promoter and is present in a vector; optionally, the vector is selected from the group consisting of retroviral vectors, lentiviral vectors, phage vectors, adenovirus vectors, adeno-associated virus vectors, herpes simplex virus vectors, and plasmid vectors. 24. An engineered vector comprising the polynucleotide of any one of claims 15-18, wherein the vector may be an inducible, conditional, or constitutive expression vector. 25. A vector system comprising one or more polynucleotides as described in any one of claims 15-18 and one or more polynucleotides encoding guide RNA; wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein-binding segment that interacts with the Cas protein, wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer. 26. An engineered, non-naturally occurring cell comprising: a Cas protein as described in any one of claims 1-14, a polynucleotide as described in any one of claims 15-18, a CRISPR-Cas system as described in any one of claims 19-23, a vector as described in claim 24, and a vector system as described in claim 25. Claims 3 / 5 pages 4 CN 121729488 A 27. A cell modified using a Cas protein as described in any one of claims 1-14, a polynucleotide as described in any one of claims 15-18, a CRISPR-Cas system as described in any one of claims 19-23, a vector as described in claim 24, or a vector system as described in claim 25. 28. A kit comprising: the Cas protein as described in any one of claims 1-14, the polynucleotide as described in any one of claims 15-18, the CRISPR-Cas system as described in any one of claims 19-23, the vector as described in claim 24, the vector system as described in claim 25, or the cell as described in claim 26 or 27.29. A pharmaceutical composition comprising: the Cas protein of any one of claims 1-14, the polynucleotide of any one of claims 15-18, the CRISPR-Cas system of any one of claims 19-23, the vector of claim 24, the vector system of claim 25, or the cell of claim 26 or 27. 30. The pharmaceutical composition of claim 29, wherein the pharmaceutical composition further comprises a delivery system selected from the following options: AAV (adeno-associated virus), adenovirus, retrovirus, HSV (herpes simplex virus), Gammaretrovirus, LV (lentivirus), eCIS (extracellular contractile injection system), eVLPs (engineered virus-like particles), VLPs (virus-like particles), liposomes, plasmids, LNPs (lipid nanoparticles), exosomes, microvesicles, nucleic acid nanoassemblies, gene guns, and / or implantable devices. 31. A method for treating, preventing, diagnosing, or detecting a disease using a Cas protein as claimed in any one of claims 1-14, a polynucleotide as claimed in any one of claims 15-18, a CRISPR-Cas system as claimed in any one of claims 19-23, a vector as claimed in claim 24, a vector system as claimed in claim 25, a cell as claimed in claim 26 or 27, a kit as claimed in claim 28, or a pharmaceutical composition as claimed in claim 29 or 30. 32. A method for modifying or targeting a target DNA site, comprising providing the site with a Cas protein as claimed in any one of claims 1-14, a polynucleotide as claimed in any one of claims 15-18, a CRISPR-Cas system as claimed in any one of claims 19-23, a vector as claimed in claim 24, or a vector system as claimed in claim 25. 33. The method of claim 32, wherein modifying or targeting the target DNA site comprises inducing DNA strand breaks, altering gene expression of one or more genes, or epigenetically modifying the target DNA site; optionally, the DNA strand breaks comprise DNA double-strand breaks or DNA single-strand breaks. 34. The method of any one of claims 32 or 33, wherein the method is performed in vitro or in vivo. 35. A method for targeting and cleaving double-stranded target DNA, comprising: contacting the double-stranded target DNA with a Cas protein as described in any one of claims 1-14, a polynucleotide as described in any one of claims 15-18, a CRISPR-Cas system as described in any one of claims 19-23, or a pharmaceutical composition as described in claim 29 or 30. 36. An isolated eukaryotic cell comprising a modified target site, wherein the target site has been modified according to claim 32-35. The method of any one of claims 35, or modified using the pharmaceutical composition of any one of claims 29 or 30, or using the CRISPR-Cas system of any one of claims 19-23. 37. A system for detecting the presence of a nucleic acid target sequence in an in vitro sample, comprising: a) the Cas protein of any one of claims 1-14; b) at least one guide polynucleotide comprising a guide sequence capable of binding the target sequence, designed to form a complex with the Cas protein; c) a nucleic acid-based masking structure comprising a non-target sequence; wherein the Cas protein exhibits incidental cleavage activity of RNA and / or ssDNA and cleaves the non-target sequence in the nucleic acid-based masking structure activated by the target sequence. 38. A method for detecting a target nucleic acid in a sample, comprising: contacting one or more samples with a) a Cas protein as described in any one of claims 1-14; b) at least one guide polynucleotide comprising a guide sequence designed to be complementary to a target sequence, designed to form a complex with the Cas protein; c) contacting a nucleic acid-based masking structure comprising a non-target sequence; wherein the Cas protein exhibits incidental cleavage activity of RNA and / or ssDNA and cleaves the non-target sequence in the nucleic acid-based masking structure activated by the target sequence; and detecting a signal from the cleavage of the non-target sequence, thereby detecting one or more target sequences in the sample. 39. A guide RNA (gRNA) comprising: a) a spacer sequence of SEQ ID NO: 901; a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence SEQ ID NO: 901; or a spacer sequence having at least 99%, at least 98%, at least 97%, at least 96%, at least 95%, at least 94%, at least 93%, at least 92%, at least 91%, or at least 90% sequence identity with the sequence SEQ ID NO: 901; and b) a Cas protein-binding segment; wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence, said DR sequence hybridizing to form a double-stranded RNA (dsRNA) duplex. 40. The gRNA of claim 39, wherein said gRNA is a double-stranded gRNA or a single-stranded gRNA. 41. The gRNA of claim 40, wherein the gRNA has at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 98%, at least 99%, or 100% sequence identity with the nucleotide sequence of SEQ ID NO: 903. 42. The gRNA of claim 41, wherein the gRNA is modified with at least three nucleotides.43. A polynucleotide encoding the gRNA of any one of claims 39-42. 44. An engineered, non-naturally occurring CRISPR-Cas system comprising: a) a Cas protein or a polynucleotide encoding the Cas protein; b) at least one guide RNA (gRNA) or at least one engineered nucleic acid encoding the guide RNA as described in any one of claims 39-42, wherein the gRNA further comprises a Cas protein binding segment that interacts with the Cas protein; wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer. 45. The CRISPR-Cas system of claim 44, wherein the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence having at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identity with the sequence of SEQ ID NO: 903. 46. The CRISPR-Cas system of claim 44 or 45, wherein the Cas protein comprises a sequence having at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identity with SEQ ID NO: 12; or, except for the amino acid "M" at position 1 in the sequence, the Cas protein comprises a sequence having at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identity with SEQ ID NO: 12. 47. An engineered vector comprising the polynucleotide of claim 43, wherein the vector is optionally an inducible, conditional, or constitutive expression vector. 48. A vector system comprising one or more polynucleotides of claim 43. 49. A pharmaceutical composition comprising the gRNA of any one of claims 39-42, the polynucleotide of claim 43, the CRISPR-Cas system of any one of claims 44-46, the vector of claim 47, or the vector system of claim 48. 50. A method of treating, preventing, or diagnosing a disease associated with the RHO locus, comprising contacting the gRNA of any one of claims 39-42, the polynucleotide of claim 43, the CRISPR-Cas system of any one of claims 44-46, or the pharmaceutical composition of claim 49 with target cells of a subject requiring treatment. Claims 5 / 5 Page 6 CN 121729488 AType II Cas protein, CRISPR-Cas system and its uses Technical Field
[0001] This invention relates to type II Cas protein, CRISPR-Cas system and its uses. In particular, type II Cas protein and CRISPR-Cas system are used for gene targeting or gene editing. This application claims priority to PCT applications PCT / CN2023 / 113355, PCT / CN2023 / 116757, PCT / CN2023 / 136724, PCT / CN2024 / 091203, PCT / CN2024 / 091198, and PCT / CN2024 / 091211. The entire contents of the above applications are incorporated herein by reference. Background Art
[0002] Targeted genome editing or modification is rapidly becoming an important tool in basic and applied research, among which clustered regularly spaced short palindromic repeats and CRISPR-related protein (CRISPR-Cas) systems show great promise due to their easy and specific targeting through engineered related guide RNAs. Recent advances in genome sequencing technology and analytical methods have significantly accelerated the ability to catalog and locate genetic factors associated with a variety of biological functions and diseases. Precise genome targeting technologies are needed to systematically reverse engineer causal genetic variations by allowing selective interference with individual genetic elements, and to advance synthetic biology, biotechnology, and medical applications.
[0003] Various CRISPR-Cas systems have been explored in the prior art, and different CRISPR-Cas systems exhibit different characteristics. For example, the CRISPR-Cas9 system, belonging to the class 2 CRISPR-Cas system, has been used for genome editing and has shown great promise in biomedical research. Summary of the Invention
[0004] This invention provides an engineered, non-naturally occurring type II CRISPR-associated (Cas) protein or a variant thereof, said protein having at least 70% sequence identity with any amino acid sequence in SEQ ID NO: 1-71. In some embodiments, the sequence identity of the Cas protein with any amino acid sequence in SEQ ID NO: 1-71 is at least 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%.
[0005] The present invention also provides an engineered, non-naturally occurring type II CRISPR-related (Cas) protein, wherein, except for the amino acid "M" at position 1 in the sequence, the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with any amino acid sequence in SEQ ID NO: 1-71.
[0006] In some embodiments, the Cas protein further comprises an effector domain (or functional domain). Such an effector domain canIt possesses one or more types of enzyme activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain includes nuclease, nickase, deaminase, reverse transcriptase, recombinase, methyltransferase, methyltransferase, acetyltransferase, transcription activator, transcription repressor domain, cryptochrome, photoinducible / controllable domain, or chemically inducible / controllable domain.
[0007] In some embodiments, the Cas protein further includes one or more nuclear localization signal sequences, nuclear export signal sequences, cell-penetrating peptide sequences, and affinity tags. Type II Cas proteins include one or more nuclear localization signals (NLS). The NLS may be located at the end of the peptide chain or other parts. The NLS located at both ends or other parts of the Cas9 amino acid sequence may be the same or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are the same. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains an NLS, and the C-terminus of the Cas9 amino acid sequence contains an NLS. NLS are fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence. The NLS can be SV40 (monkey virus 40) NLS, c-Myc NLS, or other suitable monomeric NLS. The NLS can be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, the Cas protein is purified by affinity chromatography using an affinity tag (such as GST, FLAG, or a six-histidine sequence). In some embodiments, the amino acid sequence of the C-terminal NLS is given in SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is given in SEQ ID NO: 883. The NLS and FLAG sequences can also be selected from other available sequences and different combinations.
[0008] In some embodiments, the Cas protein comprises an amino acid sequence that shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the amino acid sequences in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, 68.73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%. In some other preferred embodiments, except for amino acid "M" at position 1 in the sequence, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the amino acid sequences in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, 68.
[0010] A key element in the CRISPR-Cas system-related modification process is the adjacent motif (PAM), which is a short DNA sequence immediately adjacent to the target DNA sequence. PAM sequences are crucial for the binding and cleavage activity of Cas proteins, ensuring that only predetermined DNA sequences are targeted for editing. The ability of Cas proteins to recognize these specific PAM sequences is attributed to their unique structural features and binding interactions with DNA. Each PAM sequence provides a distinct motif that Cas proteins can recognize, ensuring accurate and efficient targeting. For example, the NRRANH sequence may provide a specific nucleotide arrangement that Cas proteins can bind with high affinity. This diversity of PAM sequences expands the potential applications of the CRISPR-Cas system. It enables researchers and scientists to target a wider range of DNA sequences for editing, increasing the versatility and effectiveness of this powerful gene-editing tool. Furthermore, understanding these PAM sequences can help develop more advanced Cas proteins with higher precision targeting capabilities, further advancing the field of gene editing and its potential benefits for research, medicine, and biotechnology. The Cas protein disclosed in this invention possesses the unique ability to recognize a variety of PAM sequences. These sequences include NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. The specific recognition of these sequences by these Cas proteins allows for greater flexibility in selecting target DNA sequences for editing.
[0011] In some embodiments, the Cas protein disclosed in this invention is capable of recognizing at least one protospacer adjacent motif (PAM) having or containing the following sequences: NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. In some embodiments, the Cas protein exhibits at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 9, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the NRRANH sequence; in some embodiments, the Cas protein exhibits at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, or 82% identity with any amino acid sequence in SEQ ID NO: 12. 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and the specification page 2 / 316 8 CN 121729488 A and capable of recognizing the original spacer adjacent motif (PAM) with the sequence NRHACT; in some embodiments, the Cas protein has an identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% with any amino acid sequence in SEQ ID NO: 19, and is capable of recognizing the original spacer adjacent motif (PAM) with the sequence NRAAR ... The identity of any amino acid sequence in 21 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, and 93%.94%, 95%, 96%, 97%, 98%, 99%, or 100%, and capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NNNCCY; in some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 22, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NNRYY ...94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: The identity of any amino acid sequence in SEQ ID NO: 24 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NGG; in some embodiments, the identity of the Cas protein with any amino acid sequence in SEQ ID NO: 25 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, or 81%. 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NNNCAA; in some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with any amino acid sequence in SEQ ID NO: 27, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRNACN ... The identity of any amino acid sequence in 29 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, and 91%.92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NNGR; in some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 30, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NGGNR ...99%, or 100% identity with any amino acid sequence in SEQ ID NO: 30, and is capable of recogni The identity of any amino acid sequence in SEQ ID NO: 36 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NNNCCH; in some embodiments, the identity of the Cas protein with any amino acid sequence in SEQ ID NO: 36 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, or 81%. 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRRAAG; in some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with any amino acid sequence in SEQ ID NO: 37, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRHRAC ... The identity of any amino acid sequence in 38 is at least 70%, 71%, 72%, and 73%.74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and capable of recognizing protospacer adjacent motifs with the sequence NRYART (PAM); in some embodiments, the Cas protein has an identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88% with any amino acid sequence in SEQ ID NO: 43. 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRHACC; in some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 44, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRAAR ...98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: The identity of any amino acid sequence in SEQ ID NO: 51 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRNVHH; in some embodiments, the identity of the Cas protein with any amino acid sequence in SEQ ID NO: 56 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, or 96%. 97%, 98%, 99%, or 100%, and capable of recognizing protospacer adjacent motifs (PAMs) with the sequence YMACAW; in some embodiments, the Cas protein is associated with SEQ ID NO:The identity of any amino acid sequence in SEQ ID NO: 60 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NHAA; in some embodiments, the identity of the Cas protein with any amino acid sequence in SEQ ID NO: 60 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, or 88%. 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRHAYY; in some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 68, and is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NGGHA.
[0012] In some embodiments, the Cas protein is a nicking enzyme active or inactive Cas protein. The DNA cleavage domain of the active Cas protein in this invention comprises two subdomains: the HNH nuclease subdomain and the RuvC subdomain. Mutations within these subdomains can inhibit the nuclease activity of the Cas protein. In some embodiments, the Cas protein exhibits at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 12, and includes mutations at residues D11 or H859; or at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, and 92% amino acid sequence identity with SEQ ID NO: 12.93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and including mutations at residue D10 or H862; or the amino acid sequence identity with SEQ ID NO: 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and including mutations at residue D12 or H903. In some embodiments, except for amino acid "M" at position 1 in the sequence, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 12, and includes mutations at residues D11 or H859; or except for amino acid "M" at position 1 in the sequence, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, or 94% amino acid sequence identity with SEQ ID NO: 12. Specification 4 / 316 pages 10 CN 121729488 A 95%, 96%, 97%, 98%, 99% or 100%, and including mutations at residue D10 or H862; or except for amino acid "M" at position 1 in the sequence, the amino acid sequence identity with the amino acid sequence of SEQ ID NO: 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and including mutations at residue D12 or H903. In some embodiments, the mutation at residue D11 or H859 in SEQ ID NO: 9 is D11A or H859A; the mutation at residue D10 or H862 in SEQ ID NO: 12 is D10A or H862A; and the mutation at residue D12 or H903 in SEQ ID NO: 31 is D12A or H903A. In some embodiments, the Cas protein is related to SEQ ID NO:The identity of any amino acid sequence in 869-877 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0013] The present invention also provides an engineered, non-naturally occurring polynucleotide encoding a type II CRISPR-associated (Cas) protein.
[0014] In some embodiments, the polynucleotide encoding the Cas protein is operatively linked to a promoter and presented in a vector; optionally, the vector is selected from the group consisting of retroviral vectors, lentiviral vectors, phage vectors, adenovirus vectors, adeno-associated virus vectors, herpes simplex virus vectors, and plasmid vectors.
[0015] In some embodiments, the polynucleotide is a ribonucleotide sequence or a deoxyribonucleotide sequence, or an analogue thereof; optionally, the polynucleotide is codon-optimized for expression in cells of interest; in some embodiments, the polynucleotide is codon-optimized for expression in eukaryotic cells. In some embodiments, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cells are mammalian cells, preferably human cells.
[0016] In some embodiments, the polynucleotide is mRNA and further comprises a 5' cap sequence and / or a poly-A tail sequence. In some embodiments of the invention, the mRNA used may be modified to enhance its functional properties and stability. Specifically, in some embodiments, the modification process involves replacing uridine (represented by the letter "U") with N1-methylpseuuridine or pseudouridine. This substitution aims to improve the mRNA's resistance to ribonuclease degradation, potentially increasing its intracellular half-life and translation efficiency. Incorporating N1-methylpseuuridine or pseudouridine into the mRNA structure can also positively influence the immunogenicity profile, as these modifications have been shown to reduce the immunogenicity of mRNA molecules compared to unmodified mRNA molecules. This is particularly critical for the development of mRNA-based therapeutics and vaccines, where minimizing adverse immune responses is essential.
[0017] In some embodiments, the polynucleotides of the present invention are codon-optimized for expression in eukaryotic cells; optionally, the eukaryotic cells are selected from the group consisting of plant cells, fungal cells, unicellular eukaryotes, mammalian cells, etc.Cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and / or human cells.
[0018] In some embodiments, the polynucleotide shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the nucleotide sequences in SEQ ID NO: 161-231, 241-311, 851-852, 861-863.
[0019] The present invention also provides an engineered, non-naturally occurring CRISPR-Cas system comprising: a) a type II Cas protein or a polynucleotide encoding a Cas protein as described herein; b) at least one engineered guide RNA or at least an engineered nucleic acid encoding a guide RNA, wherein the guide RNA comprises a spacer sequence complementary to the target nucleic acid and a Cas protein-binding segment that interacts with the Cas protein, wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer.
[0020] In some embodiments, the guide RNA is a dual guide RNA. In some embodiments, the guide RNA is a single guide RNA. In these embodiments, the guide RNA further comprises a linker sequence connecting the tracrRNA sequence and the DR sequence. In some typical embodiments, the linker sequence comprises a short GAAA sequence. In some embodiments, the linker sequence is an artificial loop. In some embodiments, the sgRNA comprises the following sequence: a) a spacer sequence capable of hybridizing with the target nucleic acid sequence to be manipulated; b) a DR sequence; c) a linker sequence; and d) a tracrRNA sequence. The spacer sequence, DR sequence, linker sequence, and tracrRNA sequence are arranged in a 5' to 3' orientation, or a 3' to 5' orientation; in some embodiments, the sgRNA backbone comprises a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to any of the sequences in SEQ ID NO: 431-562.
[0021] In some embodiments, the sgRNA comprises a spacer sequence (such as any one of SEQ ID NO: 571-835, 885, 901) and a backbone sequence, wherein the spacer sequence is located at the 5' end of the backbone sequence (such as the sequence in SEQ ID NO: 903). In some embodiments, the sgRNA comprises a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any one of the sequences in SEQ ID NO: 903.
[0022] In some embodiments, the spacer sequence hybridizes with one or more nucleic acids in prokaryotic or eukaryotic cells. In some embodiments, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the eukaryotic cells include mammalian cells. In some embodiments, mammalian cells include human cells. In some embodiments, the eukaryotic cells include plant cells.
[0023] In some embodiments, the system further includes a donor template nucleic acid.
[0024] In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid. In some embodiments, the donor template nucleic acid is a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear. In some embodiments, the donor template nucleic acid is circular (e.g., a plasmid). In some embodiments, the donor template nucleic acid is an exogenous nucleic acid molecule. In some embodiments, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, the donor template nucleic acid is DNA or RNA or a DNA-RNA hybrid.
[0025] The present invention also provides an engineered vector comprising the polynucleotides described in this disclosure.
[0026] In some embodiments, the vector is an expression vector. In some embodiments, the vector is an inducible, conditional, or constitutive expression vector. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA are located on the same vector or on different vectors.
[0027] The present invention also provides a vector system comprising one or more polynucleotides described in this disclosure and one or more polynucleotides encoding guide RNA; wherein the guide RNA comprises a spacer sequence complementary to the target nucleic acid and a spacer sequence complementary to the target nucleic acid.The Cas protein-binding segment of the Cas protein interaction comprises a tracrRNA sequence and a direct repeat (DR) sequence, which hybridize to form a double-stranded RNA (dsRNA) dimer. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA are located on the same vector or on different vectors.
[0028] In some embodiments, the vector, such as a plasmid or viral vector, is delivered to the target tissue by, for example, intramuscular injection, intravenous administration, transdermal administration, intranasal administration, oral administration, or mucosal administration. Such administration may be a single dose or multiple doses. Those skilled in the art will understand that the actual dose may vary greatly due to a variety of factors, such as vector selection, target cells, target organism, target tissue, general condition of the subject, required degree of transformation / modification, route of administration, manner of administration, required type of transformation / modification, etc.
[0029] The present invention also provides an engineered, non-naturally occurring cell comprising: the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, or the vector system described herein.
[0030] The present invention also provides a cell modified by utilizing the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, or the vector system described herein.
[0031] In some embodiments, the cell is a eukaryotic cell or a prokaryotic cell. In some embodiments, the eukaryotic cell is selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cell is a mammalian cell, a human cell, or a plant cell.
[0032] In some embodiments, the cell is a vertebrate, mammal, rodent, goat, pig, bird, chicken, turkey, cow, horse, sheep, fish, primate, or human cell. In some embodiments, the cell is a mammalian cell. In one embodiment, the cell is a human cell. In some embodiments, the cell is a somatic cell, germ cell, or fetal cell. In some embodiments, the cell is a zygote, blastocyst, embryonic cell, stem cell, mitotically capable cell, or meiotically capable cell. In some embodiments, the cell is not part of a human embryo. In some embodiments, the cell is a somatic cell. In one embodiment, the cell is a T cell, CD8+ T cell, CD8+ naive T cell, central memory T cell, effector memory T cell, or CD4+ T cell.T cells, stem cell memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricted progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythrocyte progenitor cells, megakaryocytic erythrocyte progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, aqueous reticulum cells, cochlear hair cells, outer hair cells, inner hair cells, lung epithelial cells, bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, neurons, neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells. Monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells (e.g., progenitor B cells, pre-B cells, original B cells, memory B cells, plasma cells), gastrointestinal epithelial cells, biliary epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, preadipocytes, pancreatic islet cells (e.g., β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0033] In some embodiments, this disclosure provides an isolated eukaryotic cell comprising a modified target site, wherein the target site has been modified according to the method described in the invention or using the composition or system described in the invention.
[0034] In some embodiments, the cell is a eukaryotic cell or a prokaryotic cell. In some embodiments, the eukaryotic cell is selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cell is a mammalian cell, a human cell, or a plant cell.
[0035] In some embodiments, the cell is a vertebrate, mammalian, rodent, goat, pig, bird, chicken, turkey, cow, horse, sheep, fish, primate, or human cell. In some embodiments, the cell is a mammalian cell. In one embodiment, the cell is a human cell. In some embodiments, the cell is a somatic cell, germ cell, or fetal cell. In some embodimentsIn this example, the cell is a zygote cell, a blastocyst cell, an embryonic cell, a stem cell, a cell capable of mitosis, or a cell capable of meiosis. In some embodiments, the cell is not part of a human embryo. In some embodiments, the cell is a somatic cell. In one embodiment, the cells are T cells, CD8+ T cells, CD8+ naive T cells, central memory T cells, effector memory T cells, CD4+ T cells, stem cell memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricted progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythrocyte progenitor cells, megakaryocytic erythrocyte progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, aqueous reticulum cells, cochlear hair cells, outer hair cells, inner hair cells, lung epithelial cells, bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, neurons, neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, and B cells. Examples include progenitor B cells, pre-B cells, original B cells, memory B cells, plasma cells, gastrointestinal epithelial cells, biliary epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, preadipocytes, pancreatic islet cells (e.g., β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0036] The present invention also provides a kit comprising: the Cas protein described in this disclosure, the polynucleotide described in this disclosure, the CRISPR-Cas system described in this disclosure, the vector described in this disclosure, the vector system described in this disclosure, or the cells described in this disclosure.
[0037] The kits described in this disclosure may include one or more containers containing the components necessary to perform the methods of this disclosure, and may include instructions for use. Any run-through kit may also include auxiliary components necessary to perform the editing methods. Each component in the kit may be provided in liquid form (e.g., dissolved in solution) or solid form (e.g., lyophilized powder), where applicable. In certain embodiments, some components may need to be reformulated or otherwise treated (e.g., activated state) after the addition of suitable solvents or other substances (such as water or buffers), which may be used to re-formulate or otherwise treat these components.The substance may or may not be provided with the kit. In some embodiments, the kit may further include other suitable excipients, such as buffers or reagents, to facilitate the application of the kit. The kit can be used in a variety of applications, such as medical applications, including therapeutic and diagnostic, research, etc. Therefore, the type II Cas nuclease and kit of the present invention can be used to prepare pharmaceuticals or reagents for therapeutic and / or research purposes.
[0038] Cas proteins, the CRISPR-Cas system, and the polynucleotides described herein can be delivered by various delivery systems, such as vectors, such as plasmids, viral delivery vectors, such as adeno-associated virus (AAV), lentivirus, adenovirus and other viral vectors, or methods, such as ribo-electroporation or electroporation of ribonucleoprotein complexes composed of type V-VI effectors and their corresponding guide RNAs. Proteins and one or more guide RNAs can be packaged into one or more vectors, such as plasmids or viral vectors. For bacterial applications, nucleic acids encoding any component of the CRISPR system described in the present invention can be delivered into bacteria using bacteriophages. Exemplary phages include, but are not limited to, T4 phage, Mu, λ phage, T5 phage, T7 phage, T3 phage, Φ29, M13, MS2, Qβ, and ΦX174. Specification 8 / 316 pages 14 CN 121729488 A
[0039] The present invention also provides a pharmaceutical composition comprising: the Cas protein described in this disclosure, the polynucleotide described in this disclosure, the CRISPR-Cas system described in this disclosure, the vector described in this disclosure, the vector system described in this disclosure, or the cell described in this disclosure.
[0040] As stated in this disclosure, a “pharmaceutical composition” refers to a formulation intended for pharmaceutical use. In some embodiments, the pharmaceutical composition further includes an acceptable pharmaceutical excipient. In some embodiments, the pharmaceutical composition may contain other therapeutic agents. In some embodiments, the pharmaceutical composition is prepared according to standard procedures and can be administered to a subject, such as a human patient, via intravenous, intramuscular, intradermal, intra-articular, intralesional, intraperitoneal, intracardiac, intracerebrospinal fluid, intravenous, epidural, local, subconjunctival, periocular, intraocular, vitreous, posterior sclera, penetrating sclera, suprascleral, subretinal, retroretinal, fundus, intranasal inhalation, pressurized inhalation, oral, subcutaneous, or local routes. For example, compositions for injection may be provided as sterile isotonic aqueous solutions. If necessary, the pharmaceutical composition may also contain a resolving agent and a local anesthetic, such as lidocaine, to minimize discomfort at the injection site. Typically, components may be provided individually or as mixtures in unit doses, such as as lyophilized powders or anhydrous concentrated solutions, in sealed containers indicating the amount of active ingredient. If the pharmaceutical composition is intended for infusion, it may be combined with an infusion bottle containing sterile pharmaceutical-grade water or physiological saline. If the pharmaceutical composition is for injection, sterile water for injection or...Physiological saline may be included to mix the components prior to administration. Additionally, wetting agents, colorants, release agents, coating agents, sweeteners, flavoring agents, preservatives, and antioxidants may also be incorporated into the formulation as needed.
[0041] In some embodiments, the pharmaceutical composition further includes a delivery system selected from: AAV (adeno-associated virus), adenovirus, retrovirus, HSV (herpes simplex virus), γ-retrovirus, lentivirus, eCIS (extracellular contractile injection system), eVLPs (engineered virus-like particles), VLPs (virus-like particles), liposomes, plasmids, LNPs (lipid nanoparticles), exosomes, microvesicles, nucleic acid nanoassemblies, gene guns, and / or implantable devices.
[0042] In some embodiments, delivery is carried out via adeno-associated virus (AAV), such as AAV2, AAV8, or AAV9, which may contain at least 1 × 10^5 adenovirus or adeno-associated virus particles (also known as particle units, pu) in a single dose.
[0043] In some embodiments, delivery is performed via a recombinant adeno-associated virus (rAAV) vector. For example, in some embodiments, a modified AAV vector may be used for delivery. The modified AAV vector may be based on one or more capsid types, including AAV1, AAV2, AAV5, AAV6, AAV8, AAV8.2, AAV9, AAV rh10, modified AAV vectors (e.g., modified AAV2, modified AAV3, modified AAV6), and pseudotyped AAVs (e.g., AAV2 / 8, AAV2 / 5, and AAV2 / 6).
[0044] In some embodiments, delivery is performed via plasmids. The dosage may be a sufficient quantity of plasmid to elicit a response. In some embodiments, a suitable amount of plasmid DNA in the plasmid composition may range from about 0.1 mg to about 2 mg. Plasmids typically include (i) a promoter; (ii) a sequence encoding a CRISPR enzyme targeting a nucleic acid, operatively linked to the promoter; (iii) a selectivity marker; (iv) an origin of replication; and (v) a transcription terminator, located downstream of (ii) and operatively linked to (ii). Plasmids may also encode RNA components of the CRISPR-Cas system, but one or more of these may be encoded on different vectors. Dosing frequency is determined by a medical or veterinary practitioner (e.g., a physician, veterinarian) or skilled technician.
[0045] The present invention also provides methods for treating, preventing, diagnosing, or detecting diseases using the Cas protein, polynucleotide, CRISPR-Cas system, vector, vector system, cell, kit, or pharmaceutical composition described in this disclosure.
[0046] The present invention also provides a method for modifying or targeting a target DNA site, the method comprising using the DNA described in this disclosure…The Cas protein, polynucleotide, CRISPR-Cas system, vector, vector system, kit, or pharmaceutical composition described herein are delivered to the site.
[0047] In some embodiments, the disclosure also provides a method for targeting and cleaving target DNA, the method comprising: contacting the target DNA with the Cas protein, polynucleotide, CRISPR-Cas system, vector, vector system, kit, or pharmaceutical composition described herein.
[0048] In some embodiments, modifying or targeting the target site includes inducing DNA strand breaks. In some embodiments, modifying or targeting the target site includes inducing DNA double-strand breaks or DNA single-strand breaks. In some embodiments, modifying or targeting the target site includes altering the gene expression of one or more genes. In some embodiments, modifying or targeting the target site includes epigenetic modification of the target DNA site. In some embodiments, the method is a method of modifying a cell, cell line, or organism by manipulating one or more target sequences at a site of interest in the genome.
[0049] In some embodiments, cleaving the target DNA or target sequence results in the formation of an insertion or deletion (indel) or the insertion of a nucleotide sequence. In some embodiments, cleaving the target DNA or target nucleotide includes cleaving the target DNA or target sequence at two sites, resulting in deletion or inversion of the sequence between the two sites. In some embodiments, the target DNA is double-stranded DNA or single-stranded DNA, or a DNA-RNA hybrid.
[0050] In some embodiments, modifying or targeting the target site includes inducing DNA strand breaks, altering gene expression of one or more genes, or epigenetic modification of the target DNA site; optionally, DNA strand breaks include DNA double-strand breaks or DNA single-strand breaks.
[0051] In some embodiments, the method is performed in vitro or in vivo.
[0052] The present invention also provides an isolated eukaryotic cell comprising the modified target site, wherein the target site has been modified according to the methods described in this disclosure, or using the systems described in this disclosure, or using the Cas protein described in this disclosure, or using the polynucleotides described in this disclosure, or using the CRISPR-Cas system described in this disclosure, or using the vectors described in this disclosure, or using the vector systems described in this disclosure, or using the kits described in this disclosure, or using the pharmaceutical compositions described in this disclosure.
[0053] The present invention also provides a system for detecting the presence of a nucleic acid target sequence in an in vitro sample, comprising: a) the Cas protein described in this disclosure; b) at least one guide polynucleotide comprising a guide sequence capable of binding to the target sequence and designed to form a complex with the Cas protein; and c) a nucleic acid-based masking structure comprising a non-target sequence, wherein the Cas proteinIt exhibits incidental cleavage activity against RNA and / or ssDNA, and cleaves non-target sequences in a nucleic acid-based masking structure activated by the target sequence.
[0054] The present invention also provides a method for detecting target nucleic acids in a sample, comprising: contacting one or more samples with a) the Cas protein described in this disclosure; b) at least one guide polynucleotide comprising a guide sequence designed to be complementary to a target sequence and designed to form a complex with the Cas protein; and c) a nucleic acid-based masking structure comprising a non-target sequence, wherein the Cas protein exhibits incidental cleavage activity against RNA and / or ssDNA, and cleaves non-target sequences in a nucleic acid-based masking structure activated by the target sequence; and detecting a signal of non-target sequence cleavage, thereby detecting one or more target sequences in the sample.
[0055] The present invention also provides a guide RNA (gRNA) comprising: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence having at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA further comprises a Cas protein-binding segment; wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence, which hybridize to form a double-stranded RNA (dsRNA) dimer. In some embodiments, the gRNA is a dual guide RNA. In some embodiments, the gRNA is a single guide RNA. In some embodiments, the gRNA is modified. In some embodiments, at least three nucleotides of the gRNA are modified. In some embodiments, the gRNA includes a 5' end modification that contains at least two phosphothioester (PS) bonds in the first seven nucleotides at the 5' end. In some embodiments, the gRNA includes a 3' end modification that contains at least two phosphothioester (PS) bonds in the first seven nucleotides at the 3' end. In some embodiments, the gRNA has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any nucleotide sequence of SEQ ID NO: 903.
[0056] The present invention also provides a polynucleotide encoding the gRNA described in this document; wherein the guide RNA (gRNA) comprises: a)a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence consisting of at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the sequence of SEQ ID NO: 901; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 901. In some embodiments, the gRNA comprises a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any sequence of SEQ ID NO: 903.
[0057] The present invention also provides an engineered, non-naturally occurring CRISPR-Cas system comprising: a) a Cas protein or a polynucleotide encoding a Cas protein; b) at least one gRNA or at least one engineered nucleic acid encoding a gRNA as described in this disclosure, wherein the gRNA further comprises a Cas protein binding segment that interacts with the Cas protein; wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence, the sequences hybridizing to form a double-stranded RNA (dsRNA) dimer; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the sequence of SEQ ID NO: 12.
[0058] In some embodiments, except for the amino acid "M" at position 1 in the sequence, the Cas protein comprises a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, or 81% identity to any one of the sequences of SEQ ID NO: 12.82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the sequence; or the Cas protein contains a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to any of the sequences in SEQ ID NO: 12.
[0059] In some embodiments, the Cas protein further comprises an effector domain (or functional domain). Such an effector domain may have one or more types of enzymatic activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain includes nuclease, nickase, deaminase, reverse transcriptase, recombinase, methyltransferase, methyltransferase, acetyltransferase, transcription activator, transcription repressor domain, cryptochrome, photoinducible / controllable domain, or chemically inducible / controllable domain.
[0060] In some embodiments, the Cas protein further includes one or more nuclear localization signal sequences, nuclear export signal sequences, cell-penetrating peptide sequences, and affinity tags. Type II Cas proteins include one or more nuclear localization signals (NLS). The NLS may be located at the end of the peptide chain or other parts. The NLS located at both ends or other parts of the Cas9 amino acid sequence may be the same or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are the same. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains an NLS, and the C-terminus of the Cas9 amino acid sequence contains an NLS. The amino acid sequence of the NLS is fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence. The NLS can be SV40 (monkey virus 40) NLS, c-Myc NLS, or other suitable monomeric NLS. The NLS can be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, the Cas protein is purified by affinity chromatography using an affinity tag (such as GST, FLAG, or a hexahistidine sequence). In some embodiments, the amino acid sequence of the C-terminal NLS is selected from SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is selected from SEQ ID NO: 883. Other available sequences and different combinations of NLS and FLAG sequences can also be selected.
[0061] In some embodiments, the Cas protein comprises an amino acid sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 12.
[0062] In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the sequence NRHACT, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRHACT.
[0063] In some embodiments, the Cas protein is a cleaving enzyme or an inactivated Cas protein. The DNA cleavage domain of the active Cas protein of the present invention comprises two subdomains, namely the HNH nuclease subdomain and the RuvC subdomain. Mutations within these subdomains can inhibit the nuclease activity of the Cas protein. In some embodiments, the amino acid sequence identity of the Cas protein with SEQ ID NO: 12 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes mutations at residues D10 or H862; in some embodiments, except for the amino acid "M" at position 1 in the sequence, the amino acid sequence identity of the Cas protein with SEQ ID NO: 12 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, or 78%. 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and including mutations at residues D10 or H862.
[0064] In some embodiments, the mutation at residues D10 or H862 in SEQ ID NO: 12 is D10A or H862A; in some embodiments, the Cas protein is similar to SEQ ID NO:The identity of any amino acid sequence in 872-874 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0065] The present invention also provides an engineered vector comprising a polynucleotide encoding gRNA as described in this disclosure; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the sequence of SEQ ID NO: 903.
[0066] In some embodiments, the vector is optionally an inducible, conditional, or constitutive expression vector.
[0067] The present invention also provides a vector system comprising one or more polynucleotides encoding the gRNA described herein; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 871 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the sequence of SEQ ID NO: 903.
[0068] The present invention also provides a pharmaceutical composition comprising the gRNA described herein; a polynucleotide encoding such gRNA; a CRISPR-Cas system comprising such gRNA or polynucleotide; a vector comprising a gRNA encoding sequence; and a vector system comprising a gRNA encoding sequence as described on page 12 / 316 of the specification, CN 121729488 A; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO:901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO:901 sequence; and c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO:901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 903.
[0069] The present invention also provides a method for treating, preventing or diagnosing diseases related to the RHO gene locus, comprising: a) contacting target cells with the gRNA described in this disclosure; b) contacting target cells with a polynucleotide encoding such gRNA; c) contacting target cells with a CRISPR-Cas system comprising such gRNA or polynucleotide; or d) contacting target cells with a pharmaceutical composition described in this disclosure; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 903.
[0070] The present invention also provides a method for treating, preventing or diagnosing diseases related to gene loci, comprising administering a) the gRNA described in this disclosure; b) administering a target cell with a polynucleotide encoding such gRNA; c) administering a target cell with a CRISPR-Cas system containing such gRNA or polynucleotide; or d) administering a target cell with a pharmaceutical composition described in this disclosure; wherein the guide RNA (gRNA) comprises: i) a spacer sequence of SEQ ID NO: 901; ii) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or iii) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 903.
[0071] The present invention also provides a composition comprising: (i) a Cas protein, wherein: a. the Cas protein contains a sequence that is at least 90% identical to SEQ ID NO: 12 or 92; and / or b. the Cas protein contains a sequence that is at least 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 12 or 92; and / or (ii) an sgRNA or a vector encoding sgRNA, wherein the sgRNA contains the sgRNA sequence of SEQ ID NO: 903.
[0072] The present invention also provides a method for modifying RHO gene sites, comprising delivering a composition to cells, wherein the composition comprises: a. a guide RNA comprising a guide sequence of SEQ ID NO: 901; b. a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c. a guide RNA comprising a guide sequence that is at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identical to the SEQ ID NO: 901 sequence.
[0073] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a desired subject, wherein the composition comprises: a. a guide RNA comprising SEQ ID NO:a. A guide RNA containing at least 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c. A guide RNA containing a guide sequence that is at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identical to the SEQ ID NO: 901 sequence.
[0074] The present invention also provides a method for modifying RHO gene sites, comprising delivering a composition to cells, wherein the composition comprises: a. an sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b. an sgRNA comprising an sgRNA sequence having at least 90% identity with the SEQ ID NO: 903 sequence; or c. an sgRNA comprising an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 903 sequence.
[0075] The present invention also provides a method for treating, preventing or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a. an sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b. an sgRNA comprising an sgRNA sequence having at least 90% identity with the SEQ ID NO: 903 sequence; or c. an sgRNA comprising an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 903 sequence.
[0076] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a. a guide RNA comprising a spacer sequence of SEQ ID NO: 901; b. a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c. a guide RNA comprising a guide sequence that is at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identical to the SEQ ID NO: 901 sequence.
[0077] The present invention also provides a method for modifying an RHO gene site, comprising delivering a composition to a cell, wherein the composition comprises: (i) a Cas protein, wherein: a. the Cas protein comprises a sequence that is at least 90% identical to the sequence of SEQ ID NO: 12 or 92; and / or b. an RNA-guided DNA binder comprising at least 95%, 96%, 97%, 98%, or 99% identical to the sequence of SEQ ID NO: 12 or 92.100% identical sequence; and / or (ii) a guide RNA or a vector encoding a guide RNA, wherein the guide RNA comprises the spacer sequence of SEQ ID NO: 901.
[0078] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: (i) an RNA-guided DNA binder, wherein: a. the RNA-guided DNA binder comprises a sequence with at least 90% identity to SEQ ID NO: 12 or 92; and / or b. the RNA-guided DNA binder comprises a sequence with at least 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO: 12 or 92; and / or (ii) an sgRNA or a vector encoding an sgRNA, wherein the sgRNA comprises the sequence of SEQ ID NO: 903.
[0079] The advantages of the embodiments described herein, as well as other embodiments, objects, features, and exemplary embodiments, will be apparent to those skilled in the art in conjunction with the following detailed description.
[0080] The features and advantages of this disclosure can be understood by referring to the following detailed description, which describes exemplary embodiments that can utilize the principles of this disclosure, and is accompanied by corresponding figures: Figure 1 shows the domain arrangement of the GEBx type II Cas protein; Figures 2A to 2E show the PAM preference of Cas9 in the HEK293 cell line; Figure 3 shows the indel levels of GEBx0305 against 16 target sites with GGAAAA PAM in the HEK293T cell line (n=3); Figure 4 shows the indel levels of GEBx0308 against 19 target sites with GGTACT PAM in the HEK293T cell line (n=3); Figure 5 shows the indel levels of GEBx0308 against 15 target sites with CATACT PAM in the HEK293T cell line (n=3). Figure 6 shows the indel levels of GEBx0328 against 20 target sites and NGGCCTPAM in the HEK293T cell line (n=3); Figure 7 shows the indel levels of GEBx0361 against 11 target sites and GGTACC-PAM in the HEK293T cell line (n=2); Figure 8 shows the indel levels of GEBx0361 against 8 target sites and TGTACC-PAM in the HEK293T cell line (n=2); Figure 9: A and B show the indel levels of GEBx0305 as a function of guide sequence length (n=3).Figure 10: A and B (Figure 10: A and B) show the indel levels of GEBx0308 as a function of guide sequence length (n=3); Figure 11: A and B (Figure 11: A and B) show the indel levels of GEBx0328 as a function of guide sequence length (n=3); Figure 12: A and B (Figure 12: A and B) show the indel levels of GEBx0305 against endogenous genes, using a modified RNA backbone (n=3); Figure 13: A and B (Figure 13: A and B) show the indel levels of GEBx0308 against endogenous genes, using a modified RNA backbone (n=3); Figure 14: A and B (Figure 14: A and B) show the indel levels of GEBx0305 against endogenous genes under optimal conditions (n=3); Figure 15: A and B (Figure 15: A and B) show the indel levels of GEBx0308 against endogenous genes under optimal conditions (n=3). Figure 16: A and B show the indel levels of GEBx0328 against endogenous genes using a modified RNA backbone (n=3); Figure 17: A and B show the indel levels of GEBx0328 against endogenous genes under optimal conditions (n=3); Figure 18 shows the indel levels of GEBx0305 against endogenous genes after HEK293T cells were transfected with a liposome complex containing a fixed amount (20 ng) of sgRNA and different proportions of mRNA. SpCas9 was used as a positive control; Figure 19 shows the indel levels of GEBx0308 targeting endogenous genes after HEK293T cells were transfected with a liposome complex containing a fixed amount (20 ng) of sgRNA and different proportions of mRNA; Figure 20 shows the indel levels of GEBx0305 targeting endogenous genes after PHH cells were transfected with a liposome complex containing a fixed amount (20 ng) of sgRNA and different proportions of mRNA. SpCas9 was used as a positive control; Figure 21 shows the indel levels of GEBx0308 targeting endogenous genes after PHH cells were transfected with a liposome complex containing a fixed amount (20 ng) of sgRNA and different proportions of mRNA; Figure 22 shows the guide-seq insertions of GEBx0305 targeting site 1 (CFTR-NGGAAAA-T5) and site 2 (EMX1-NGGAAAA-T5); Figure 23 shows the guide-seq insertions of GEBx0308 targeting site 1 (CD34-NGGTACT-T4) and site 2 (POLQ-NGGTACT-T1);Figure 24 shows the Guide-seq insertion of GEBx0328 at (CFTR-NGGCCT-T3) and site 2 (CFTR-NGGCCT-T5); Figure 25 shows the base editing efficiency of GEBx0305-ABE for adenine A-to-G conversion at four sites in HEK293T cells; Figure 26 shows the base editing efficiency of GEBx0308-ABE for adenine A-to-G conversion at five sites in HEK293T cells; Figure 27 shows the base editing efficiency of GEBx0328-ABE for adenine A-to-G conversion at five sites in HEK293T cells.
[0081] Figure 28 shows the 4-allele-specific editing of GEBx0308 at the RHO-P23H pathogenic site. Specification 15 / 316 pages 21 CN 121729488 A
[0082] Detailed Description of the Invention
[0083] The following examples further illustrate the contents of this disclosure, but this disclosure is not limited thereto.
[0084] It is necessary to point out that the singular forms or terms such as “a,” “an,” “this,” and “said,” as used herein and in the appended claims, and similar terms used in the context of this disclosure (especially in the context of the claims), should be interpreted to cover both the singular and the plural, unless otherwise stated herein or clearly contradicted by the context. In some embodiments, the above terms may be reasonably understood as “a” or “one or more.” Furthermore, unless the context requires otherwise, singular terms should include the plural, and plural terms should include the singular.
[0085] Unless otherwise stated, all singular / plural terms also include the active and past voice forms of the terms and need to be understood according to the context in the article.
[0086] Please note that in this disclosure and particularly in the claims and / or paragraphs, terms such as “comprising,” “consisting of,” “including,” etc., may have the meanings given by U.S. Patent Law; for example, they may mean “included,” “contained,” “including,” etc.; while terms such as “substantially composed of” and “substantially composed of” have the meanings given by U.S. Patent Law. The term “a group of”, etc., refers to a particular group or set of elements, components, or features. It may include one or more of the specified elements, components, or features. For example, “a group of A, B, or C” may refer to a collection that includes one or more of any specified elements A, B, or C. The claims include the possibility that there may be any single element (A, B, or C), any combination of two elements (A and B, A and C, or B and C), or all three elements together (A, B, and C). This phrase defines the invention according to the specified options, allowing for different combinations of the listed elements while still maintaining the scope of the claim.
[0087] When “t” or “T” appears as a nucleotide in the RNA sequence of the present invention, it should be understood as “u” or “U”.
[0088] In the context of two or more nucleic acid or polypeptide sequences, the term “identity” refers to two or more sequences or subsequences being identical, or having a specific percentage of identical amino acid residues or nucleotides, using sequence comparison algorithms such as BLAST or BLAST 2.0 or FASTA, with default parameters as described below.
[0089] The term “exemplary” is used herein to mean as an example, instance, or illustration. Any aspect or design described as “exemplary” is not necessarily to be understood as more preferred or advantageous than other aspects, embodiments, or designs.
[0090] As stated herein, the terms “optional” or “or” mean that the events, situations, or alternatives described subsequently may or may not occur, and the description includes instances where the events or situations occur and instances where they do not occur.
[0091] The use of “or” or “ / ” is inclusive and means “and / or” unless otherwise stated; or interpreted according to the different contexts. The term “and / or” is used herein, as in phrases such as “A and / or B”, to include A and B; A or B; A alone; and B alone. Similarly, the term “and / or” is used herein, as in phrases such as “A, B and / or C”, to cover each of the following combinations: A, B and C; A, B or C; A or C; A or B; B or C; A and C; A and B; B and C; A alone; B alone; and C alone.
[0092] Terms such as “approximately” and “~” as used herein refer to a range of measurable values, such as parameters, quantities, durations of time, etc., and should be understood to include variations from the specified values. It should be understood that the values referred to by the modifiers “approximately” or “~” are themselves specifically and explicitly disclosed.
[0093] The term “exemplary” is used herein to mean as an example, instance, or illustration. Any aspect or design described as “exemplary” is not necessarily to be construed as more preferred or advantageous than other aspects, embodiments, or designs.
[0094] The terms “subject,” “individual,” and “patient” are used interchangeably herein and refer to vertebrates, preferably mammals, and more preferably humans. Mammals include, but are not limited to, mice, primates, humans, farm animals, sporting animals, and pets. Tissues, cells, and their offspring of biological entities, whether obtained in vivo or cultured in vitro, are also included.
[0095] A “variant” of the sequence disclosed herein includes a sequence having one or more additions, deletions, stop positions, or substitutions compared to the sequence disclosed herein. Specification 16 / 316 pages 22 CN 121729488 A
[0096] “Encoding” refers to a specific sequence of nucleotides in a gene, such as cDNA or mRNA, as a means of synthesizing other macromolecules.The properties of a template, such as the defined amino acid sequence. Therefore, if the transcription and translation of the mRNA corresponding to the gene produces a protein in a cell or other biological system, the gene encodes a protein. Polynucleotides encoding proteins include all nucleotide sequences that are degenerate versions of each other and encode the same amino acid sequence or amino acid sequences that are substantially similar in form and function.
[0097] Terms such as “non-spontaneous” or “engineered” are used interchangeably to indicate human involvement. These terms, when referring to nucleic acid molecules or polypeptides, mean that the nucleic acid molecule or polypeptide is substantially unrelated in nature to at least one of its naturally associated components. In all aspects and embodiments, whether or not these terms are included, it should be understood that they are preferably optional and therefore preferably included or not included. Furthermore, terms such as “non-spontaneous” and “engineered” are used interchangeably and can therefore be used alone or in combination, with one of them replacing a mention of both. In particular, “engineered” is preferred and can replace “non-spontaneous” or “non-spontaneous and / or engineered” or “engineered, non-spontaneous”.
[0098] As described in this invention, the term "cleavage event" refers to a DNA break created on a target nucleic acid by a type II Cas nuclease in the CRISPR system described in this invention. In some embodiments, the cleavage event is a double-strand DNA break. In some embodiments, the cleavage event is a single-strand DNA break.
[0099] As described in this invention, the term "target" refers to the ability of a complex comprising a CRISPR-associated protein and an RNA guide to preferentially or specifically bind, for example, to the target nucleic acid, compared to other nucleic acids that do not have the same or similar sequence as the target nucleic acid.
[0100] According to this disclosure, the term "GEBx" followed by a numerical suffix is used as a generic code representing a nucleic acid or protein. It should be noted that using the same code for nucleic acids and proteins or their derivatives does not mean that the substances represented by these codes are the same. In other words, GEBx-0305 may refer to a specific nucleic acid sequence in one case and a different protein in another. Some embodiments may demonstrate a direct correspondence between nucleic acids and proteins represented by the same or derived codes. Therefore, the “GEBx” code serves as an indexing system for organizing and referencing the various biomolecules described in this disclosure, and the meaning of the code will be understood in the context provided.
[0101] Unless otherwise defined herein, scientific and technical terms related to this disclosure should have the meaning commonly understood by a person of ordinary technical skill. The meaning and scope of terms should be clear; however, in the event of any potential ambiguity, the definitions provided herein take precedence over any dictionary or external definition.
[0102] Various embodiments are described below. It should be noted that the specific embodiments are not intended as an exhaustive description or an extension of this disclosure.The limitations of the broader schemes discussed herein. An scheme described in conjunction with a particular embodiment is not necessarily limited to that embodiment and may be practiced with any other embodiment. The use of phrases such as “in some embodiments,” “in some particular embodiments,” “in some preferred embodiments,” “in some typical embodiments,” or similar expressions throughout the specification means that a particular feature, structure, or characteristic associated with that embodiment is included in at least one embodiment. Furthermore, a particular feature, structure, or characteristic may be combined in one or more embodiments in any suitable manner according to this disclosure. In addition, although some embodiments described herein include features not included in other embodiments, combinations of features from different embodiments should be within the scope of the disclosure. For example, in the dependent claims, embodiments of any claim may be used in any combination.
[0103] Numerical ranges referenced by endpoints include all numbers and decimals within their respective ranges, as well as the endpoints referenced.
[0104] Various embodiments are described below. It should be noted that specific embodiments are not intended as exhaustive descriptions or limitations on the broader schemes discussed herein. An aspect described in conjunction with a particular embodiment is not necessarily limited to the example described in this specification 17 / 316 page 23 CN 121729488 A and may be practiced with any other embodiment. The phrases “in a particular embodiment,” “in some embodiments,” and “in some specific embodiments” used throughout the specification mean that a particular feature, structure, or characteristic associated with that embodiment is included in at least one embodiment. Therefore, phrases such as “in a particular embodiment,” “in one embodiment,” or “in some specific embodiments” appearing in different places throughout the specification do not necessarily refer to the same embodiment, but may refer to one or more. Furthermore, specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments according to this disclosure. Additionally, although some embodiments described herein include features not included in other embodiments, combinations of features from different embodiments should be within the scope of the disclosure. For example, in the appended claims, embodiments of any claim may be used in any combination.
[0105] All publications, published patent documents, and patent applications cited herein are cited to the same extent as each individual publication, published patent document, or patent application is specifically and individually indicated as being cited.
[0106] The present invention provides an engineered, non-naturally occurring type II CRISPR-associated (Cas) protein or a variant thereof, said protein having at least 70% identity with any amino acid sequence in SEQ ID NO: 1-71. In some embodiments, the Cas protein shares at least 75%, 80%, 85%, or 90% identity with any amino acid sequence in SEQ ID NO: 1-71.92%, 95%, 98%, 99%, or 100%.
[0107] As described herein, “M” refers to the starting amino acid methionine, which is typically the starting point for the synthesis of many proteins. Naturally occurring Cas proteins typically begin with methionine as the first amino acid in their sequence. However, when scientists engineer these proteins, for example by fusing them with nuclear localization signals (NLS) or other domains, this initial amino acid “M” may be replaced or altered to introduce new functions or properties.
[0108] In this disclosure, particular attention is paid to the flexibility and functionality of engineered Cas proteins. Apart from intentionally modifying the starting amino acid methionine (M) at position 1, engineered Cas proteins exhibit significant sequence identity, ranging from 70% to 100%, when compared to a reference sequence. This strategic alteration not only aligns with our goals of customizing protein properties but also ensures that the inherent or intended functions of the Cas protein are preserved. By manipulating the initial amino acid without compromising overall sequence similarity, certain properties are enhanced, such as improved cellular localization or the introduction of other advantageous features, while maintaining the essential characteristics that make Cas proteins indispensable tools in genome manipulation.
[0109] The present invention also provides an engineered, non-naturally occurring type II CRISPR-associated (Cas) protein, except for amino acid "M" at position 1 in the sequence, wherein the Cas protein has at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 1-71.
[0110] As described in this disclosure, the terms "Cas protein," "CRISPR-associated protein," or other similar terms refer to a component of the CRISPR-Cas system, and these proteins are components of the CRISPR-Cas system. These proteins may possess inherent nuclease activity, enabling them to cleave double-stranded DNA or RNA molecules in a sequence-specific manner, guided by complementary RNA molecules, such as Cas9 and Cas12, which have been widely used in genome editing applications. Furthermore, some Cas proteins may be engineered to retain only one of the two active sites required for double-strand cleavage, resulting in cleavage enzyme activity that allows them to introduce single-strand breaks in the target nucleic acid sequence to control the cleavage event. Furthermore, some Cas proteins may be engineered to lack any inherent nuclease activity, referred to as inactivated Cas or nuclease-inactive Cas proteins. Despite lacking enzymatic function, these inactivated Cas proteins retain their specific binding ability to target nucleic acid sequences and are often used in conjunction with other effector domains (or functional domains) for gene regulation, epigenetics editing, and as components of advanced imaging systems. In some embodiments, Cas proteins can be used to reduce non-target effects. In some embodiments, active Cas nucleases, nicking enzymes, or inactivated Cas nucleases...Live Cas may also be part of a fusion protein containing another effector domain. Fusion proteins containing these other effector domains (or functional domains) also fall under the category of Cas proteins. In some embodiments, the Cas protein may be a split form. In some embodiments, such as page 24 of CN 121729488 A on page 18 / 316 of this specification, the Cas protein may also be an inducible Cas protein. In some embodiments, the type II Cas protein may be part of a self-inactivating system (SIN); in some embodiments, the type II Cas nuclease may also be part of a co-activating system (SAM), as defined elsewhere in this document.
[0111] In some embodiments, the domain arrangement of the type II Cas protein includes a RuvC domain, a BH (bridged helix) domain, a REC domain, an HNH domain, and / or a CTD (C-terminal domain). The RuvC domain is a key catalytic site responsible for the cleavage of the target DNA strand. It contains three split RuvC subdomains that fold in a complex manner to form the active site where DNA cleavage occurs. These subdomains work in coordination, guided by guide RNA, to recognize and cleave DNA at specific locations. The BH domain, or bridge-helix domain, serves as a structural link between different domains of the Cas protein. It is recognized as an arginine-rich region containing multiple arginine amino acids. The abundance of these arginine residues is crucial for interaction with the phosphate backbone of the target DNA strand. Arginine residues can form hydrogen bonds with phosphate groups, helping to correctly locate and align the target DNA for cleavage. The REC domain, or recognition leaf, participates in the recognition of the target DNA sequence. It contributes to the specificity of the Cas protein by distinguishing between target and non-target sequences, ensuring that only the intended DNA fragment is cleaved. The HNH domain is another catalytic site, working in conjunction with the RuvC domain to cleave the complementary strand of the target DNA. Named for its characteristic histidine-aspartic-histidine sequence, it is essential for the nuclease activity of the Cas protein. The CTD, or C-terminal domain, is typically involved in interactions with other proteins or cellular structures, contributing to the localization and regulation of the Cas protein within the cell. It may also play a role in the stability and overall conformation of the Cas protein, ensuring its functionality and specificity at the target. This complex domain arrangement enables type II Cas proteins to perform their precise and critical functions in CRISPR systems, making them invaluable tools for genome editing and manipulation.
[0112] In some embodiments, the Cas protein further comprises an effector domain (or functional domain). Such an effector domain may have one or more types of enzymatic activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain includes nuclease, nickase, deaminase, reverse transcriptase, recombinase, methyltransferase, methyltransferase, acetyltransferase, transcription activator, transcription repressor domain, cryptoflora domain, etc.The Cas protein may contain a pigment, a photoinducible / controllable domain, or a chemically induced / controllable domain. In some embodiments, the Cas protein further includes one or more nuclear localization signal sequences, nuclear export signal sequences, cell-penetrating peptide sequences, and affinity tags. Type II Cas proteins contain one or more nuclear localization signals (NLS). The NLS may be located at the end of the peptide chain or at other locations. The NLS located at both ends or at other locations of the Cas9 amino acid sequence may be the same or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are the same. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains one NLS and the C-terminus contains one NLS. The amino acid sequences of the NLS are fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence, respectively. The NLS may be an SV40 (monkey virus 40) NLS, a c-Myc NLS, or other suitable monomeric NLS. The NLS may be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, the Cas protein is purified by affinity chromatography using an affinity tag (such as GST, FLAG, or a hexahistidine sequence). In some embodiments, the amino acid sequence of the C-terminal NLS is given in SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is given in SEQ ID NO: 883. Other available sequences and different combinations may also be selected for the NLS and FLAG sequences.
[0113] In some embodiments, the Cas protein comprises an amino acid sequence that has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the amino acid sequences in SEQ ID NO: 1-71.
[0114] In some preferred embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, 68. In some other preferred embodiments, except for position 1 in the sequenceThe amino acid "M" of the Cas protein is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any of the amino acid sequences in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, 68.
[0115] One of the key elements of the CRISPR-Cas system-related modification process is the protospacer adjacent motif (PAM), which is a short DNA sequence adjacent to the target DNA sequence. The PAM sequence is crucial for the binding and cleavage activity of the Cas protein, ensuring that only the predetermined DNA sequence is targeted for editing. The ability of Cas proteins to recognize these specific PAM sequences is attributed to their unique structural features and binding interactions with DNA. Each PAM sequence provides a distinct motif that the Cas protein can recognize, ensuring accurate and efficient targeting. For example, the NRRANH sequence may provide a specific nucleotide arrangement that the Cas protein can bind to with high affinity. This diversity of PAM sequences expands the potential applications of the CRISPR-Cas system. It enables researchers and scientists to target a wider range of DNA sequences for editing, increasing the versatility and effectiveness of this powerful gene-editing tool. Furthermore, understanding these PAM sequences can help develop more advanced Cas proteins with higher precision targeting capabilities, further advancing the field of gene editing and its potential benefits for research, medicine, and biotechnology. The Cas protein disclosed in this invention possesses the unique ability to recognize a variety of PAM sequences. These sequences include NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. The specific recognition of these sequences by these Cas proteins allows for greater flexibility in selecting target DNA sequences for editing.
[0116] As described in this invention, in the provided PAM sequence, “N” represents any one of the four standard DNA nucleotides: adenine (A), thymine (T), cytosine (C), or guanine (G). This code facilitates the inclusion of any nucleotide at a given position without separate specification; “R” specifically represents a purine nucleotide, namely adenine (A) or guanine (G). Purines are present in DNA structures...Essential for both structure and function, this code simplifies the process of incorporating these larger nucleotides into the PAM sequence; "H" represents any nucleotide other than guanine (G), and can therefore represent adenine (A), thymine (T), or cytosine (C). This code helps exclude guanine (G) at specific positions, as its larger size may affect structural or functional aspects of the sequence. "Y" represents a pyrimidine nucleotide, namely thymine (T) or cytosine (C). Pyrimidines are commonly found in specific regions of DNA and RNA, and this code helps include them without specifying the exact nucleotide. "W" represents a weak base, which can be either adenine (A) or thymine (T). This distinction is biochemically relevant because adenine and thymine have similar properties in certain situations, such as hydrogen bonding. "V" reflects any nucleotide other than thymine (T), and therefore includes adenine (A), cytosine (C), or guanine (G). This code is helpful when thymine is not needed or preferred due to its unique chemical properties within pyrimidines. “M” represents adenine (A) or cytosine (C).
[0117] In some embodiments, the Cas protein as described in this invention is capable of recognizing at least one protospacer adjacent motif (PAM) having or containing the following sequences: NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, or NGGHA. In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 12, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the NRRANH sequence; in some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, or 84% amino acid sequence identity with SEQ ID NO: 12. 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and able to recognize the instruction manual, page 20 / 316, document number 26, CN 121729488 A.The Cas protein possesses a protospacer adjacent motif (PAM) with the sequence NRHACT; in some embodiments, the amino acid sequence identity of the Cas protein with SEQ ID NO: 19 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NRAAR; in some embodiments, the amino acid sequence identity of the Cas protein with SEQ ID NO: 21 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, or 80%. 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NNNCCY; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NNRYY ... The amino acid sequence identity of SEQ ID NO: 24 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing protospacer adjacent motifs (PAMs) with the NGG sequence; in some embodiments, the amino acid sequence identity of the Cas protein with SEQ ID NO: 25 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%.94%, 95%, 96%, 97%, 98%, 99%, or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NNNCAA; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRNACN ...98%, 99%, or 100%, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRNACN; in some embodiments, the Cas protein has an amino acid sequence identity of at The amino acid sequence identity of 29 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NNGR; in some embodiments, the amino acid sequence identity of the Cas protein with SEQ ID NO: 30 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, or 86%. 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NGGNR; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NNNCCH; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89% or 100%, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NNNCCH; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 99% or 100%, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NNNCCH; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, The amino acid sequence identity of 36 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, and 82%.83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRRAAG; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRHRAC ...99% or 100 The amino acid sequence identity of 38 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRYART; in some embodiments, the amino acid sequence identity of the Cas protein with SEQ ID NO: 43 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, ... (Specification 21 / 316 pages 27 CN 121729488 A) 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the NRHACC sequence; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% with the NRAAR ...94%, 95%, 96%, 97%, 98%, 99% or 100% with the NRAAR sequence The amino acid sequence identity of 51 is at least 70%, 71%, and 72%.73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRNVHH; in some embodiments, the amino acid sequence identity of the Cas protein to SEQ ID NO: 56 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, or 89%. 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and capable of recognizing a protospacer adjacent motif (PAM) with the sequence YMACAW; in some embodiments, the Cas protein has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 59, and is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NAHAA; in some embodiments, the Cas protein has 97%, 98%, 99%, or 100% amino acid identity with SEQ ID NO: 60, and is capable of recognizing a protospacer adjacent motif (PAM) with the sequence NNNCA ... The amino acid sequence identity of SEQ ID NO: 27 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing a proto-spacer adjacent motif (PAM) with the sequence NRNACN; in some embodiments, the amino acid sequence identity of the Cas protein with SEQ ID NO: 29 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%.93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NNGR; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NGGNR ...99%, or 100%, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NGGNR; in some embodiments, the Cas protein has an amino acid sequence identity The amino acid sequence identity of 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NNNCCH; in some embodiments, the amino acid sequence identity of the Cas protein with SEQ ID NO: 36 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, or 86%. 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRRAAG; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRHRAC ...99%, or 100%, and is capable of recognizing proto-spacer adjacent mot The amino acid sequence identity of 38 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, and 82%.83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NRYART; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and is ...RHACC; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 89%, The amino acid sequence identity of 44 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and it is capable of recognizing proto-spacer adjacent motifs (PAMs) with NRAAR sequences; in some embodiments, the amino acid sequence identity of the Cas protein with SEQ ID NO: 51 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%. The amino acid sequence identity of the Cas protein with SEQ ID NO: 56 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and the Cas protein is capable of recognizing the proto-spacer adjacent motif (PAM) with the sequence YMACAW. In some embodiments, the amino acid sequence identity of the Cas protein with SEQ ID NO: 59 is at least 70%, 71%.72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NHAA; in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and is capable of recognizing proto-spacer adjacent motifs (PAMs) with the sequence NHAA; The original spacer adjacent motif (PAM); in some embodiments, the Cas protein has an amino acid sequence identity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and is capable of recognizing the original spacer adjacent motif (PAM) having the sequence NGGHA.
[0118] As described herein, the terms “recognize,” “recognize,” or “recognize” refer to the ability of the Cas protein to form a functional complex with sgRNA at a DNA target site where the sgRNA hybridizes (i.e., the location where the spacer sequence of the sgRNA hybridizes), and this location is surrounded by a PAM sequence, wherein the Cas protein is capable of performing its natural function, i.e., DNA cleavage or DNA binding. In this context, it is important to note that such DNA cleavage excludes the possibility that the type II Cas protein is a type II Cas nuclease with lost catalytic activity. For example, in the case of an inactivated type II Cas nuclease (e.g., an inactivated type II Cas nuclease), a complex can still be formed between the type II Cas nuclease, sgRNA, and the corresponding target if the desired PAM sequence is present, but such a complex will not lead to DNA cleavage.
[0119] In some embodiments, the Cas protein is a cleavage enzyme or an inactivated Cas protein. The DNA cleavage domain of the active Cas protein of the present invention comprises two subdomains, namely the HNH nuclease subdomain and the RuvC subdomain. Mutations within these subdomains can inhibit the nuclease activity of the Cas protein. In some embodiments, the Cas protein is identical to the amino acid sequence of SEQ ID NO: 9.At least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and including mutations at residue D11 or H859; or the amino acid sequence identity with SEQ ID NO: 12 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, and including mutations at residue D10 or H862 ... The amino acid sequence identity of NO: 31 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes mutations at residues D12 or H903. In some embodiments, except for the amino acid "M" at position 1 in the sequence, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 12, and includes mutations at residues D11 or H859; or except for the amino acid "M" at position 1 in the sequence, the amino acid sequence shares at least 70%, 71%, 72%, 73%, or 74% identity with the amino acid sequence of SEQ ID NO: 12. 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and including mutations at residue D10 or H862; or, except for amino acid "M" at position 1 in the sequence, the amino acid sequence identity with the amino acid sequence of SEQ ID NO: 31 is at least 70%, 71%, 72%, or 73%.74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and including mutations at residues D12 or H903. In some embodiments, the mutation at residue D11 or H859 in SEQ ID NO: 9 is D11A or H859A; the mutation at residue D10 or H862 in SEQ ID NO: 12 is D10A or H862A; and the mutation at residue D12 or H903 in SEQ ID NO: 31 is D12A or H903A. In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any amino acid sequence in SEQ ID NO: 869-877.
[0120] The present invention also provides an engineered, non-naturally occurring polynucleotide encoding the type II CRISPR-associated (Cas) protein disclosed herein.
[0121] According to the description of the invention, the term "polynucleotide" refers to a polymeric form of nucleotides of any length, whether deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. Therefore, this term includes, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers composed of purine and pyrimidine bases or other naturally, chemically, or biochemically modified non-natural or derived nucleotide bases. In some embodiments, the polynucleotide encoding the type II Cas protein described in this invention can be operatively linked to each other and to associated regulatory sequences (such as promoters, enhancers, and termination regions). For example, a functional link between a regulatory sequence and a foreign nucleic acid sequence can result in the expression of the latter. In other embodiments, the first nucleic acid sequence is said to be operatively linked to the second nucleic acid sequence when the first nucleic acid sequence is functionally related to the second nucleic acid sequence. For example, if a promoter affects the transcription or expression of a coding sequence, then the promoter is operatively linked to the coding sequence. Typically, the operatively linked DNA sequences are contiguous, and the coding regions are linked to the same reading frame when necessary or helpful. In some embodiments, the promoter is a constitutive promoter, a tissue-specific promoter.A promoter may be induced. In some embodiments, the term may also include all introns and other DNA sequences spliced from mRNA transcripts, as well as variants resulting from alternative splicing sites. These nucleic acid sequences may be DNA strand sequences transcribed into RNA or RNA sequences translated into proteins. Nucleic acid sequences include full-length nucleic acid sequences and non-full-length sequences derived from full-length proteins. Sequences may also include degenerate codons of the original sequence or sequences introduced to provide codon preference in a particular cell type.
[0122] In some embodiments, a polynucleotide encoding a Cas protein is operatively linked to a promoter and presented in a vector; optionally, the vector is selected from the group consisting of retroviral vectors, lentiviral vectors, phage vectors, adenovirus vectors, adeno-associated virus vectors, herpes simplex virus vectors, and plasmid vectors.
[0123] In some embodiments, the polynucleotide is a ribonucleotide sequence or a deoxyribonucleotide sequence, or an analogue thereof; optionally, the polynucleotide is codon-optimized for expression in cells of interest; in some embodiments, the polynucleotide is codon-optimized for expression in eukaryotic cells. In some embodiments, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cells are mammalian cells, preferably human cells.
[0124] In some embodiments, the polynucleotide is mRNA and further comprises a 5' cap sequence and / or a poly-A tail sequence. In some embodiments of the invention, the mRNA used may be modified to enhance its functional properties and stability. Specifically, in some embodiments, the modification process involves replacing uridine (represented by the letter "U") with N1-methylpseudouridine or pseudouridine. This substitution aims to enhance the resistance of mRNA to degradation by ribonucleases, potentially increasing its half-life and translation efficiency within cells. Incorporating N1-methylpseudouridine or pseudouridine into the mRNA structure can also positively influence the immunogenicity profile, as these modifications have been shown to reduce the immunogenicity of mRNA molecules compared to unmodified mRNA molecules. This is particularly critical for the development of mRNA therapeutics and vaccines, where minimizing adverse immune responses is essential.
[0125] In some embodiments, the polynucleotides of this invention are codon-optimized for expression in eukaryotic cells;Optionally, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, and non-human primate cells.
[0126] In some embodiments, the polynucleotide has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the nucleotide sequences in SEQ ID NO: 161-231, 241-311, 851-852, 861-863.
[0127] The present invention also provides an engineered, non-naturally occurring CRISPR-Cas system comprising: a) a type II Cas protein or a polynucleotide encoding a Cas protein as described herein; b) at least one engineered guide RNA or at least one engineered nucleic acid encoding a guide RNA, wherein the guide RNA comprises a spacer sequence complementary to the target nucleic acid and a Cas protein-binding segment interacting with the Cas protein, wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer.
[0128] According to the description of the invention, the term “complementary” describes the ability of two nucleic acid strands to pair with each other through their bases. This complementarity can occur through perfect base pairing, where each base on one strand forms a specific hydrogen bond pairing with its complementary base on the opposite strand, following the standard Watson-Crick base pairing rules: adenine (A) pairs with thymine (T) or uracil (U), and cytosine (C) pairs with guanine (G). This perfect base pairing allows for precise recognition and binding between the two strands. Furthermore, complementarity can also involve imperfect base pairing, including mismatches, insertions, or deletions that may result in non-standard base pairing or reduced interstrand affinity. Despite these imperfections, sufficient complementarity is maintained between strands to interact, although reduced specificity or stability may occur in duplexes.
[0129] According to the description of the invention, direct repeat (DR) sequences are derived from short repeat DNA sequence elements in a CRISPR array. This sequence is typically scattered among spacer sequences derived from exogenous genetic material such as bacteriophage or plasmid DNA. Posttranscriptionally, the DR sequence is part of the precrRNA transcript and is processed into mature CRISPR RNA (crRNA). The DR sequence in crRNAThe DR sequence is a key component in the formation of a double-stranded RNA (dsRNA) dimer by pairing with a complementary sequence in transcriptionally activated CRISPR RNA (tracrRNA). This is crucial for the function of the guide RNA (gRNA) complex in the CRISPR / Cas system. In the context of this invention, the DR sequence comprises the portion of the gRNA that pairs with the tracrRNA sequence to form a dsRNA dimer, thereby enabling the formation of the Cas protein binding segment. The expression "hybridization to form a double-stranded RNA (dsRNA) dimer," or similar expressions, refers in this invention to the process by which a direct repeat (DR) sequence in CRISPR RNA (crRNA) pairs with a complementary sequence in transcriptionally activated CRISPR RNA (tracrRNA) to form a stable double-stranded RNA structure. This dsRNA dimer is a key component of the guide RNA (gRNA) complex, facilitating Cas protein binding. Specifically, the DR sequence in crRNA and the complementary sequence in tracrRNA undergo base pairing, forming hydrogen bonds between their complementary bases, leading to the formation of the dsRNA dimer. This dimer is crucial for the function of the Cas protein binding segment and comprises a tracrRNA sequence and a DR sequence, which hybridize to form a dsRNA dimer.
[0130] According to the description of the invention, the term "target nucleic acid" refers to a specific nucleic acid substrate containing a nucleic acid sequence that is wholly or partially complementary to the RNA guide sequence. In some embodiments, the target nucleic acid comprises a gene or a sequence within a gene. In some embodiments, the target nucleic acid comprises a non-coding region (e.g., a promoter). In some embodiments, the target nucleic acid is single-stranded. In some embodiments, the target nucleic acid is double-stranded. The terms "target nucleic acid" or "target sequence" should be understood according to the context of the disclosure.
[0131] In some embodiments, the guide RNA is a dual guide RNA. In some embodiments, the guide RNA is a single guide RNA. In these embodiments, the guide RNA also includes a linker sequence connecting the tracrRNA sequence and the DR sequence. In some typical embodiments, the linker sequence comprises a short GAAA sequence. In some embodiments, the linker sequence is an artificial loop. In some embodiments, the sgRNA comprises the following sequence: a) a spacer sequence capable of hybridizing with the target nucleic acid sequence to be manipulated; b) a DR sequence; c) a linker sequence; and d) a tracrRNA sequence. The spacer sequence, DR sequence, linker sequence, and tracrRNA sequence are arranged in tandem in a 5' to 3' orientation, or a 3' to 5' orientation; in some embodiments, the sgRNA backbone comprises [the sequence name is missing from the original text].The identity of any sequence in SEQ ID NO: 571-835, 885, 901 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0132] In some embodiments, the sgRNA comprises a spacer sequence (such as any one of SEQ ID NO: 571-835, 885, 901) and a backbone sequence, wherein the spacer sequence is located at the 5' end of the backbone sequence (such as the sequence in SEQ ID NO: 903). In some embodiments, the sgRNA comprises a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the sequence of SEQ ID NO: 903.
[0133] In some embodiments, the spacer sequence hybridizes with one or more nucleic acids in a prokaryotic or eukaryotic cell. In some embodiments, the eukaryotic cell is selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the eukaryotic cell includes mammalian cells. In some embodiments, mammalian cells include human cells. In some embodiments, eukaryotic cells include plant cells.
[0134] In some embodiments, the system further includes a donor template nucleic acid.
[0135] As described in this disclosure, the term "donor template nucleic acid" refers to a nucleic acid molecule that one or more cellular proteins can use to alter the structure of a target nucleic acid after a Cas protein has altered the structure of that target nucleic acid. In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid. In some embodiments, the donor template nucleic acid is a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear. In some embodiments, the donor template nucleic acid is circular (e.g., a plasmid). In some embodiments, the donor template nucleic acid is a foreign nucleic acid molecule. In some embodiments, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome). In some embodiments, the donor template nucleic acid is DNA or RNA or a DNA-RNA hybrid.
[0136] The present invention also provides an engineered vector comprising the polynucleotides described in this disclosure.
[0137] As described in this disclosure, a “vector” is a tool that allows or facilitates the transfer of an entity from one environment to another. It is a replicon, such as a plasmid, bacteriophage, or cosmosome, that can insert another DNA fragment to achieve replication, as described in the insert fragment specification 26 / 316 pages 32 CN 121729488 A. Typically, a vector is capable of replication when associated with appropriate control elements. Generally, a “vector” refers to a nucleic acid molecule capable of transporting another nucleic acid molecule linked to it. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends or without free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other polynucleotide variants known in the art. One type of vector is a “plasmid,” which refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted, as through standard molecular cloning techniques. Another type of vector is a viral vector, in which a virus-derived DNA or RNA sequence is present in the vector for packaging into a virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses (AAVs)). Viral vectors also include polynucleotides carried by the virus for transfection into host cells. Some vectors are capable of autonomous replication in the host cells in which they are introduced (e.g., bacterial vectors with bacterial origins of replication and circular mammalian vectors). Other vectors (e.g., non-circular mammalian vectors) integrate into the host genome upon introduction into the host cell and replicate along with the host genome. Furthermore, some vectors are capable of directing the expression of genes operatively linked to them. These vectors are referred to herein as “expression vectors.” Expression vectors commonly used in recombinant DNA technologies are typically in plasmid form. Recombinant expression vectors may contain the nucleic acids of the present invention in a form suitable for expression in host cells, meaning that recombinant expression vectors include one or more regulatory elements selected based on the host cell used, operatively linked to the nucleic acid sequence to be expressed. In recombinant expression vectors, “operatively linked” means that the nucleotide sequence of interest is linked to the regulatory element in a manner that allows the nucleotide sequence to be expressed (e.g., in an in vitro transcription / translation system or when the vector is introduced into a host cell). In some embodiments, the vector is an expression vector. In some embodiments, the vector is an inducible, conditional, or constitutive expression vector. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA are located on the same vector or on different vectors.
[0138] The present invention also provides a vector system comprising one or more polynucleotides described herein and one or more polynucleotides encoding guide RNA; wherein the guide RNA comprises a spacer sequence complementary to the target nucleic acid and a spacer sequence complementary to the target nucleic acid.The Cas protein interacting Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence, which hybridize to form a double-stranded RNA (dsRNA) dimer. In some embodiments, the polynucleotide encoding the Cas protein and the polynucleotide encoding the guide RNA are located on the same vector or on different vectors.
[0139] The present invention also provides an engineered, non-naturally occurring cell comprising: the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, or the vector system described herein.
[0140] The present invention also provides a cell modified by utilizing the Cas protein described herein, the polynucleotide described herein, the CRISPR-Cas system described herein, the vector described herein, or the vector system described herein.
[0141] In some embodiments, the cell is a eukaryotic cell or a prokaryotic cell. In some embodiments, the eukaryotic cell is selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cell is a mammalian cell, a human cell, or a plant cell.
[0142] In some embodiments, the cell is a vertebrate, mammalian, rodent, goat, pig, bird, chicken, turkey, cow, horse, sheep, fish, primate, or human cell. In some embodiments, the cell is a mammalian cell. In one embodiment, the cell is a human cell. In some embodiments, the cell is a somatic cell, germ cell, or fetal cell. In some embodiments, the cell is a zygote cell, blastocyst cell, embryonic cell, stem cell, cell capable of mitosis, or cell capable of meiosis. (See page 27 / 316 of CN 121729488 A for details.) In some embodiments, the cell is not part of a human embryo. In some embodiments, the cell is a somatic cell. In one embodiment, the cell is a T cell, CD8+ T cell, CD8+ naive T cell, central memory T cell, effector memory T cell, CD4+ T cell, stem cell memory T cell, helper T cell, regulatory T cell, cytotoxic T cell, natural killer T cell, hematopoietic stem cell, long-term hematopoietic stem cell, short-term hematopoietic stem cell, pluripotent progenitor cell, lineage-restricted progenitor cell, lymphoid progenitor cell, myeloid progenitor cell, common myeloid progenitor cell, erythrocyte progenitor cell, megakaryocytic erythrocyte progenitor cell, retinal cell, photoreceptor cell, rod cell, cone cell, retinal pigment epithelial cell, aqueous reticulum cell, cochlear hair cell, outer hair cell, inner capillary cell.Cells, lung epithelial cells, bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, neurons, neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells, such as progenitor B cells, pre-B cells, original B cells, memory B cells, plasma cells, gastrointestinal epithelial cells, biliary epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, preadipocytes, pancreatic islet cells (such as β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0143] In some embodiments, this disclosure provides an isolated eukaryotic cell comprising a modified target site, wherein the target site has been modified using the methods described in this invention, or using the systems described in this invention, or using the Cas protein described in this invention, or using the polynucleotide described in this invention, or using the CRISPR-Cas system described in this invention, or using the vector described in this invention, or using the vector system described in this invention, or using the kit described in this invention, or using the pharmaceutical composition described in this invention.
[0144] In some embodiments, the cells are eukaryotic cells or prokaryotic cells. In some embodiments, the eukaryotic cell is selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and human cells. In some embodiments, the cell is a mammalian cell, a human cell, or a plant cell.
[0145] In some embodiments, the cell is a vertebrate, mammalian, rodent, goat, pig, bird, chicken, turkey, cow, horse, sheep, fish, primate, or human cell. In some embodiments, the cell is a mammalian cell. In one embodiment, the cell is a human cell. In some embodiments, the cell is a somatic cell, germ cell, or fetal cell. In some embodiments, the cell is a zygote cell, blastocyst cell, embryonic cell, stem cell, mitotic cell, or meiotic cell. In some embodiments, the cell is not part of a human embryo. In some embodiments, the cell is a somatic cell. In one embodiment, the cells are T cells, CD8+T cells, CD8+ naive T cells, central memory T cells, effector memory T cells, CD4+ T cells, stem cell memory T cells, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, hematopoietic stem cells, long-term hematopoietic stem cells, short-term hematopoietic stem cells, pluripotent progenitor cells, lineage-restricted progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythrocyte progenitor cells, megakaryocytic erythrocyte progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, aqueous reticulum cells, cochlear hair cells, outer hair cells, inner hair cells, lung epithelial cells, bronchial epithelial cells, alveolar epithelial cells, lung epithelial progenitor cells, striated muscle cells, cardiomyocytes, muscle satellite cells, neurons, neural stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells. Examples include progenitor B cells, pre-B cells, original B cells, memory B cells, plasma cells, gastrointestinal epithelial cells, biliary epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatocytes, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, preadipocytes, pancreatic islet cells (e.g., β cells, α cells, δ cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In some embodiments, the cells are T cells, hematopoietic stem cells, retinal cells, cochlear hair cells, lung epithelial cells, muscle cells, neurons, mesenchymal stem cells, induced pluripotent stem cells (iPS cells), or embryonic stem cells. In another embodiment, the cells are plant cells.
[0146] In some embodiments, a vector, such as a plasmid or viral vector, is delivered to the target tissue via intramuscular injection, intravenous administration, transdermal administration, intranasal administration, oral administration, or mucosal administration. Such administration can be a single dose or multiple doses. Skilled technicians understand that the actual dosage may vary considerably due to a variety of factors, such as vector selection, target cells, organism, tissue, general condition of the patient, required degree of transformation / modification, route of administration, method of administration, and type of transformation / modification required.
[0147] In some embodiments, delivery is made via adeno-associated virus (AAV), such as AAV2, AAV8, or AAV9, which may contain at least 1 × 10^5 adenovirus or adeno-associated virus particles (also referred to as particle units, pu) in a single dose. In some embodiments, the dose is at least about 1 × 10^6 particles, at least about 1 × 10^7 particles, at least about 1 × 10^8 particles, or at least about 1 × 10^9 particles of adeno-associated virus. Due to the limited genomic payload of recombinant AAV, the type II described in this inventionThe smaller size of Cas nucleases allows for greater flexibility in packaging effectors and RNA guidelines, as well as the appropriate control sequences (e.g., promoters) required for efficient and cell-type-specific expression.
[0148] In some embodiments, delivery is made via a recombinant adeno-associated virus (rAAV) vector. For example, in some embodiments, delivery may be made using a modified AAV vector. Modified AAV vectors may be based on one or more capsid types, including AAV1, AAV2, AAV5, AAV6, AAV8, AAV8.2, AAV9, AAV rh10, modified AAV vectors (e.g., modified AAV2, modified AAV3, modified AAV6), and pseudotyped AAVs (e.g., AAV2 / 8, AAV2 / 5, and AAV2 / 6).
[0149] In some embodiments, delivery is made via plasmids. The dosage may be sufficient to elicit a response. In some embodiments, a suitable amount of plasmid DNA in the plasmid composition may range from about 0.1 to about 2 mg. Plasmids typically include (i) a promoter; (ii) a sequence encoding a CRISPR enzyme targeting a nucleic acid, operatively linked to the promoter; (iii) a selectivity marker; (iv) an origin of replication; and (v) a transcription terminator, located downstream of (ii) and operatively linked to (ii). Plasmids may also encode RNA components of the CRISPR-Cas system, but one or more of these may be encoded on different vectors. Dosing frequency is determined by a medical or veterinary practitioner (e.g., a physician, veterinarian) or skilled technician.
[0150] The present invention also provides a kit comprising the Cas protein described in this disclosure, the polynucleotide described in this disclosure, the CRISPR-Cas system described in this disclosure, the vector described in this disclosure, the vector system described in this disclosure, or the cell described in this disclosure.
[0151] Kits described in this disclosure may comprise one or more containers containing components necessary to perform the methods of this disclosure and may include instructions for use. Any described kit may also include auxiliary components necessary to perform the editing methods. Each component of the kit may be provided in liquid form (e.g., dissolved in solution) or solid form (e.g., lyophilized powder), where applicable. In certain embodiments, some components may require reconstitution or other treatment (e.g., activation state) after the addition of suitable solvents or other substances (such as water or buffers), which may or may not be provided with the kit. In some embodiments, the kit may further include other suitable excipients, such as buffers or reagents, to facilitate the application of the kit. The kit can be used in a variety of applications, such as medical applications, including therapeutic and diagnostic, research, etc. Therefore, the type II Cas nuclease and kit of the present invention can be used to prepare for therapeutic and / or treatment purposes.Or the pharmaceutical or reagent being studied.
[0152] Cas proteins, the CRISPR-Cas system, and polynucleotides, as described herein, can be delivered by various delivery systems, such as vectors (e.g., plasmids), viral delivery vectors (e.g., adeno-associated virus (AAV), lentivirus, adenovirus, or other viral vectors), or methods (e.g., ribo-electroporation or electroporation of ribonucleoprotein complexes composed of type V-VI effectors and their corresponding RNA guides or guides). The protein and one or more RNA guides can be packaged into one or more vectors, such as plasmids or viral vectors. For bacterial applications, nucleic acids encoding any of the CRISPR system components described in this invention can be delivered to bacteria using bacteriophages. Exemplary bacteriophages include, but are not limited to, T4 phage, Mu, λ phage, T5 phage, T7 phage, T3 phage, Φ29, M13, MS2, Qβ, and ΦX174.
[0153] The present invention also provides a pharmaceutical composition comprising the Cas protein described in this disclosure, the polynucleotide described in this disclosure, the CRISPR-Cas system described in this disclosure, the vector described in this disclosure, the vector system described in this disclosure, or the cell described in this disclosure.
[0154] As stated in this disclosure, a “pharmaceutical composition” refers to a formulation intended for pharmaceutical use. In some embodiments, the pharmaceutical composition further includes an acceptable pharmaceutical excipient. In some embodiments, the pharmaceutical composition may contain other therapeutic agents. In some embodiments, the pharmaceutical composition is prepared according to standard procedures and can be administered to a subject, such as a human patient, via intravenous, intramuscular, intradermal, intra-articular, intralesional, intraperitoneal, intracardiac, intracerebrospinal fluid, intravenous, epidural, topical, subconjunctival, periocular, intraocular, vitreous, posterior sclera, penetrating sclera, suprascleral, subretinal, retroretinal, fundus, intranasal inhalation, pressurized inhalation, oral, subcutaneous, or topical routes. For example, a composition for injection may be provided as a sterile isotonic aqueous solution. If necessary, the pharmaceutical composition may also contain a solvent and a local anesthetic, such as lidocaine, to minimize discomfort at the injection site. Typically, the components may be provided individually or as a mixture of unit doses, for example as a lyophilized powder or an anhydrous concentrated solution, in a sealed container indicating the amount of the active ingredient. If the pharmaceutical composition is intended for infusion, it may be combined with an infusion bottle containing sterile pharmaceutical-grade water or physiological saline. If the pharmaceutical composition is intended for injection, sterile water for injection or physiological saline may be included to mix the components prior to administration. Furthermore, wetting agents, colorants, release agents, coating agents, sweeteners, flavoring agents, preservatives, and antioxidants may also be incorporated into the formulation as needed.
[0155] In some embodiments, the pharmaceutical composition also includes a delivery system selected from: AAV (adenocarcinoma-associated disease...)Viruses, adenoviruses, retroviruses, HSV (herpes simplex virus), γ-retroviruses, lentiviruses, eCIS (extracellular contractile injection system), eVLPs (engineered virus-like particles), VLPs (virus-like particles), liposomes, plasmids, LNPs (lipid nanoparticles), exosomes, microvesicles, nucleic acid nanoassemblies, gene guns, and / or implantable devices.
[0156] The present invention also provides methods for treating, preventing, diagnosing, or detecting diseases using the Cas protein, polynucleotide, CRISPR-Cas system, vector, vector system, cell, kit, or pharmaceutical composition described in this disclosure.
[0157] The present invention also provides a method for modifying or targeting a target DNA site, the method comprising delivering the Cas protein, polynucleotide, CRISPR-Cas system, vector, vector system, kit, or pharmaceutical composition described in this disclosure to the site.
[0158] In some embodiments, the disclosure also provides a method for targeting and cleaving target DNA, the method comprising: contacting the target DNA with a Cas protein described herein, a polynucleotide described herein, a CRISPR-Cas system described herein, a vector described herein, a vector system described herein, a kit described herein, or a pharmaceutical composition described herein.
[0159] In some embodiments, modifying or targeting a target site includes inducing DNA strand breaks. In some embodiments, modifying or targeting a target site includes inducing DNA double-strand breaks or DNA single-strand breaks. In some embodiments, modifying or targeting a target site includes altering the gene expression of one or more genes. In some embodiments, modifying or targeting a target site includes epigenetic modification of a target DNA site. In some embodiments, the method is a method of modifying a cell, cell line, or organism by manipulating one or more target sequences at a site of interest in the genome.
[0160] In some embodiments, cleaving the target DNA or target sequence results in the formation of an indel or the insertion of a nucleotide sequence. In some embodiments, cleaving the target DNA or target nucleotide includes cleaving the target DNA or target sequence at two sites, resulting in deletion or inversion of the sequence between the two sites. In some embodiments, the target DNA is double-stranded DNA or single-stranded DNA, or a DNA-RNA hybrid.
[0161] In some embodiments, modifying or targeting the target site includes inducing DNA strand breaks, altering the gene expression of one or more genes, or epigenetic modification of the target DNA site; optionally, DNA strand breaks include DNA double-strand breaks or DNA single-strand breaks.
[0162] In some embodiments, the method is performed in vitro or in vivo.
[0163] The present invention also provides isolated eukaryotic cells comprising a modified target site, wherein the target site has been modified by the method described in the present invention, or using the system described in the present invention, or using the Cas protein described in the present invention, or using the polynucleotide described in the present invention, or using the CRISPR-Cas system described in the present invention, or using the vector described in the present invention, or using the vector system described in the present invention, or using the kit described in the present invention, or using the pharmaceutical composition described in the present invention.
[0164] The present invention also provides a system for detecting the presence of a nucleic acid target sequence in an in vitro sample, comprising: a) the Cas protein described in this disclosure; b) at least one guide polynucleotide comprising a guide sequence capable of binding to the target sequence and designed to form a complex with the Cas protein; and c) a nucleic acid-based masking structure comprising a non-target sequence, wherein the Cas protein exhibits incidental cleavage activity against RNA and / or ssDNA and cleaves the non-target sequence in the nucleic acid-based masking structure activated by the target sequence.
[0165] The present invention also provides a method for detecting target nucleic acids in a sample, comprising: contacting one or more samples with a) the Cas protein described in this disclosure; b) at least one guide polynucleotide comprising a guide sequence designed to be complementary to a target sequence and designed to form a complex with the Cas protein; and c) a nucleic acid-based masking structure comprising a non-target sequence, wherein the Cas protein exhibits incidental cleavage activity against RNA and / or ssDNA and cleaves the non-target sequence in the nucleic acid-based masking structure activated by the target sequence; and detecting a signal of non-target sequence cleavage, thereby detecting one or more target sequences in the sample.
[0166] As described in this disclosure, a “sample” may comprise whole cells and / or live cells and / or cell debris. A sample may comprise (or be derived from) “body fluids”. This disclosure includes examples of bodily fluids selected from amniotic fluid, aqueous humor, vitreous humor, bile, serum, breast milk, cerebrospinal fluid, earwax, lymph, gastric juice, mucus (including nasal secretions and sputum), ascites, pericardial fluid, peritoneal fluid, pleural fluid, pus, tears, saliva, sebum (skin oil), semen, sputum, synovial fluid, sweat, urine, vaginal secretions, vomitus, and one or more mixtures thereof. Samples include cell cultures, bodily fluids, and cell cultures derived from bodily fluids. Bodily fluids can be obtained from mammals, for example by puncture or other collection or sampling procedures.
[0167] The present invention also provides a guide RNA (gRNA) comprising: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence of at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence with SEQ ID NO:The 901 sequence has at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity spacer sequences. In some embodiments, the gRNA further includes a Cas protein-binding segment; wherein the Cas protein-binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence, which hybridize to form a double-stranded RNA (dsRNA) dimer. In some embodiments, the gRNA is a dual-guide RNA. In some embodiments, the gRNA is a single-guide RNA. In some embodiments, the gRNA is modified. In some embodiments, at least three nucleotides of the gRNA are modified. In one embodiment, the gRNA contains a 5' end modification that includes at least two phosphothioester (PS) bonds in the first seven nucleotides of the 5' end. In some embodiments, as described on pages 31 / 316 of the specification (CN 121729488 A), the gRNA contains a 3' end modification that includes at least two phosphothioester (PS) bonds in the first seven nucleotides of the 3' end. In some embodiments, the gRNA is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any nucleotide sequence of SEQ ID NO: 903.
[0168] The present invention also provides a polynucleotide encoding the gRNA described in this document; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any of the sequences in SEQ ID NO: 903.
[0169] An engineered, non-naturally occurring CRISPR-Cas system comprising: a) a Cas protein or encoding CasThe Cas protein is a polynucleotide; b) at least one gRNA or at least one engineered nucleic acid encoding a gRNA as described in this disclosure, wherein the gRNA further comprises a Cas protein binding segment that interacts with the Cas protein; wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 903.
[0170] In some embodiments, the Cas protein comprises a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to any one of the sequences in SEQ ID NO: 12; or, except for the amino acid "M" at position 1 in the sequence, the Cas protein comprises a sequence with at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, or 91% identity to any one of the sequences in SEQ ID NO: 12. 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the sequence.
[0171] In some embodiments, the Cas protein further comprises an effector domain (or functional domain). Such an effector domain may have one or more types of enzymatic activity, including polymerase activity, ligase activity, reverse transcriptase activity, deaminase activity, replication activity, or proofreading activity; in some embodiments, the effector domain includes nuclease, nickase, deaminase, reverse transcriptase, recombinant nuclease, ligase, cleavage enzyme, ligase ...Hisizases, methyltransferases, methyltransferases, acetyltransferases, acetyltransferases, transcription activators, transcription repressors, cryptochromes, photoinducible / controllable domains, or chemically inducible / controllable domains.
[0172] In some embodiments, the Cas protein further comprises one or more nuclear localization signal sequences, nuclear output signal sequences, cell-penetrating peptide sequences, and affinity tags. Type II Cas proteins comprise one or more nuclear localization signals (NLS). The NLS may be located at the end of the peptide chain or at other parts. The NLS located at both ends or at other parts of the Cas9 amino acid sequence may be the same or different. In some embodiments, the N-terminal NLS and the C-terminal NLS are the same. In some embodiments, the N-terminal NLS and the C-terminal NLS are different. In some embodiments, the N-terminus of the Cas9 amino acid sequence contains an NLS and the C-terminus contains an NLS. The amino acid sequence of the NLS is fused to the N-terminus and / or C-terminus of the Cas9 amino acid sequence, respectively. The NLS may be an SV40 (monkey virus 40) NLS, a c-Myc NLS, or other suitable monomeric NLS. The NLS may be fused to the N-terminus and / or C-terminus of the Cas protein. In some embodiments, the Cas protein is purified by affinity chromatography using an affinity tag (such as GST, FLAG, or a six-histidine sequence). In some embodiments, the amino acid sequence of the A-terminal NLS is given in SEQ ID NO: 881 or 882. In some embodiments, the amino acid sequence of the C-terminal FLAG sequence is given in SEQ ID NO: 883. Other available sequences and different combinations may also be selected for the NLS and FLAG sequences.
[0173] In some embodiments, the Cas protein comprises an amino acid sequence that is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 12.
[0174] In some embodiments, the Cas protein shares at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO: 12, and is capable of recognizing protospacer adjacent motifs (PAMs) with the sequence NRHACT.
[0175] In some embodiments, the Cas protein is a cleavage enzyme or an inactivated Cas protein. The DNA cleavage domain of the active Cas protein in this invention comprises two subdomains: the HNH nuclease subdomain and the RuvC subdomain. Mutations within these subdomains can inhibit the nuclease activity of the Cas protein. In some embodiments, the amino acid sequence identity of the Cas protein with SEQ ID NO: 12 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and includes mutations at residues D10 or H862; in some embodiments, except for amino acid "M" at position 1 in the sequence, the amino acid sequence identity of the Cas protein with SEQ ID NO: 12 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, or 81%. 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, and including mutations at residues D10 or H862.
[0176] In some embodiments, the mutation at residues D10 or H862 in SEQ ID NO: 12 is D10A or H862A; in some embodiments, the identity of the Cas protein with any amino acid sequence in SEQ ID NO: 872-874 is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0177] The present invention also provides an engineered vector comprising a polynucleotide encoding the gRNA disclosed herein; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence.The 903 sequence identity is at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0178] In some embodiments, the vector may be an inducible, conditional, or constitutive expression vector.
[0179] The present invention also provides a vector system comprising one or more polynucleotides encoding the gRNA disclosed herein; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 871 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence having at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the sequence of SEQ ID NO: 903.
[0180] The present invention also provides a pharmaceutical composition comprising the gRNA disclosed herein; a polynucleotide encoding the gRNA; a CRISPR-Cas system comprising the gRNA or polynucleotide; a vector comprising the gRNA coding sequence; and a vector system comprising the gRNA coding sequence; wherein the guide RNA (gRNA) comprises: a) a spacer sequence of SEQ ID NO: 901; b) a spacer sequence comprising at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; c) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence. In some embodiments, the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a spacer sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 903 sequence.The sequences are 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0181] The present invention also provides a method for treating, preventing, or diagnosing diseases associated with RHO gene loci, comprising administering a composition to a subject in need, wherein the composition comprises: a) a guide RNA comprising a spacer sequence of SEQ ID NO: 901; b) a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c) a guide RNA comprising a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence.
[0182] The present invention also provides a method for treating, preventing or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a) a guide RNA comprising a spacer sequence of SEQ ID NO: 901; b) a guide RNA comprising at least 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c) a guide RNA comprising a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence.
[0183] The present invention also provides a method for modifying RHO gene sites, comprising delivering a composition to cells, wherein the composition comprises: a) an sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b) an sgRNA comprising an sgRNA sequence having at least 90% identity with the sequence of SEQ ID NO: 903; or c) an sgRNA comprising an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 903.
[0184] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a desired subject, wherein the composition comprises: a) an sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b) an sgRNA comprising an sgRNA sequence having at least 90% identity with the sequence of SEQ ID NO: 903; or c) an sgRNA comprising an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the sequence of SEQ ID NO: 903.The sgRNA having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the 903 sequence.
[0185] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a) a guide RNA comprising a spacer sequence of SEQ ID NO: 901; b) a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c) a guide RNA comprising a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence.
[0186] The present invention also provides a method for treating, preventing or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a. a guide RNA comprising a guide sequence of SEQ ID NO: 901; b. a guide RNA comprising at least 17, 18, 19 or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c. a guide RNA comprising a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the SEQ ID NO: 901 sequence.
[0187] The present invention also provides a method for modifying RHO gene sites, comprising delivering a composition to cells, wherein the composition comprises: a. sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b. sgRNA comprising an sgRNA sequence that is at least 90% identical to the SEQ ID NO: 903 sequence; or c. sgRNA comprising an sgRNA sequence that is at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identical to the SEQ ID NO: 903 sequence. Specification page 34 / 316 40 CN 121729488 A
[0188] The present invention also provides a method for treating, preventing or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a. an sgRNA comprising the sgRNA sequence of SEQ ID NO: 903; b. an sgRNA comprising an sgRNA sequence having at least 90% identity with the sequence of SEQ ID NO: 903; or c. an sgRNA comprising an sgRNA sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91% or 90% identity with the sequence of SEQ ID NO: 903.
[0189] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: a. a guide RNA comprising a spacer sequence of SEQ ID NO: 901; b. a guide RNA comprising at least 17, 18, 19, or 20 consecutive nucleotides of the SEQ ID NO: 901 sequence; or c. a guide RNA comprising a guide sequence having at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identity with the SEQ ID NO: 901 sequence.
[0190] The present invention also provides a method for modifying RHO gene sites, comprising delivering a composition to a cell, wherein the composition comprises: (i) a Cas protein, wherein: a. the Cas protein contains a sequence that is at least 90% identical to SEQ ID NO: 12 or 92; and / or b. an RNA-guided DNA binder that contains a sequence that is at least 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 12 or 92; and / or (ii) a guide RNA or a vector encoding the guide RNA, wherein the guide RNA contains a spacer sequence of SEQ ID NO: 901.
[0191] The present invention also provides a method for treating, preventing, or diagnosing RHO-related diseases, comprising administering a composition to a subject in need, wherein the composition comprises: (i) an RNA-guided DNA binder, wherein: a. the RNA-guided DNA binder comprises a sequence that is at least 90% identical to SEQ ID NO: 12 or 92; and / or b. the RNA-guided DNA binder comprises a sequence that is at least 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 12 or 92; and / or (ii) an sgRNA or a vector encoding sgRNA, wherein the sgRNA comprises the sequence of SEQ ID NO: 903.
[0192] Table 1: Exemplary amino acid sequences involved in this disclosure. Instruction manual, pages 35 / 316, 41, CN 121729488 A
[0193] ; Instruction manual, pages 36 / 316, 42, CN 121729488 A; Instruction manual, pages 37 / 316, 43, CN 121729488 A; Instruction manual, pages 38 / 316, 44, CN 121729488 A; Instruction manual, pages 39 / 316, 45, CN 121729488 A; Instruction manual, pages 40 / 316, 46, CN 121729488 A; Instruction manual, pages 41 / 316, 47, CN 121729488 A; Instruction manual, pages 42 / 316, 48, CN 121729488 APages 43 / 316, 49 CN 121729488 A, Instruction Manual; Pages 44 / 316, 50 CN 121729488 A, Instruction Manual; Pages 45 / 316, 51 CN 121729488 A, Instruction Manual; Pages 46 / 316, 52 CN 121729488 A, Instruction Manual; Pages 47 / 316, 53 CN 121729488 A, Instruction Manual; Pages 48 / 316, 54 CN 121729488 A, Instruction Manual; Pages 49 / 316, 55 CN 121729488 A, Instruction Manual; Pages 50 / 316, 56 CN 121729488 A, Instruction Manual; Pages 51 / 316, 57 CN 121729488 A, Instruction Manual; Pages 52 / 316, 58 CN 121729488 A, Instruction Manual; Pages 53 / 316, 59 CN 121729488 A Instruction Manual 54 / 316 pages 60 CN 121729488 A Instruction Manual 55 / 316 pages 61 CN 121729488 A Instruction Manual 56 / 316 pages 62 CN 121729488 A Instruction Manual 57 / 316 pages 63 CN 121729488 A Instruction Manual 58 / 316 pages 64 CN 121729488 A Instruction Manual 59 / 316 pages 65 CN 121729488 A Instruction Manual 60 / 316 pages 66 CN 121729488 A Instruction Manual 61 / 316 pages 67 CN 121729488 A Instruction Manual 62 / 316 pages 68 CN 121729488 A Instruction Manual 63 / 316 pages 69 CN 121729488 A Instruction manual pages 64 / 316, 70 CN 121729488 A; Instruction manual pages 65 / 316, 71 CN 121729488 A; Instruction manual pages 66 / 316, 72 CN 121729488 A; Instruction manual pages 67 / 316, 73 CN 121729488 A; Instruction manual pages 68 / 316, 74 CN 121729488 A; Instruction manual pages 69 / 316, 75 CN 121729488 A; Instruction manual pages 70 / 316, 76 CN 121729488 A; Instruction manual pages 71 / 316, 77 CN 121729488 A; Instruction manual pages 72 / 316, 78 CN 121729488 A; Instruction manual pages 73 / 316.79 CN 121729488 A Instruction Manual 74 / 316 pages 80 CN 121729488 A Instruction Manual 75 / 316 pages 81 CN 121729488 A Instruction Manual 76 / 316 pages 82 CN 121729488 A Instruction Manual 77 / 316 pages 83 CN 121729488 A Instruction Manual 78 / 316 pages 84 CN 121729488 A Instruction Manual 79 / 316 pages 85 CN 121729488 A Instruction Manual 80 / 316 pages 86 CN 121729488 A Instruction Manual 81 / 316 pages 87 CN 121729488 A Instruction Manual 82 / 316 pages 88 CN 121729488 A Instruction Manual 83 / 316 pages 89 CN 121729488 A Specification 84 / 316 pages 90 CN 121729488 A Specification 85 / 316 pages 91 CN 121729488 A Specification 86 / 316 pages 92 CN 121729488 A Specification 87 / 316 pages 93 CN 121729488 A Specification 88 / 316 pages 94 CN 121729488 A Specification 89 / 316 pages 95 CN 121729488 A Specification 90 / 316 pages 96 CN 121729488 A Specification 91 / 316 pages 97 CN 121729488 A Specification 92 / 316 pages 98 CN 121729488 A
[0194] Table 2: Exemplary polynucleotides encoding Cas proteins described in this disclosure. Instruction manual 93 / 316 pages 99 CN 121729488 A
[0195] Instruction manual 94 / 316 pages 100 CN 121729488 A Instruction manual 95 / 316 pages 101 CN 121729488 A Instruction manual 96 / 316 pages 102 CN 121729488 A Instruction manual 97 / 316 pages 103 CN 121729488 A Instruction manual 98 / 316 pages 104 CN 121729488 A Instruction manual 99 / 316 pages 105 CN 121729488 A Instruction manual 100 / 316 pages 106 CN 121729488 A Instruction manual 101 / 316 pages 107 CN 121729488 A Instruction manualPages 102 / 316, 108 CN 121729488 A, Instruction Manual; Pages 103 / 316, 109 CN 121729488 A, Instruction Manual; Pages 104 / 316, 110 CN 121729488 A, Instruction Manual; Pages 105 / 316, 111 CN 121729488 A, Instruction Manual; Pages 106 / 316, 112 CN 121729488 A, Instruction Manual; Pages 107 / 316, 113 CN 121729488 A, Instruction Manual; Pages 108 / 316, 114 CN 121729488 A, Instruction Manual; Pages 109 / 316, 115 CN 121729488 A, Instruction Manual; Pages 110 / 316, 116 CN 121729488 A, Instruction Manual; Pages 111 / 316, 117 CN 121729488 A Instruction Manual 112 / 316 pages 118 CN 121729488 A Instruction Manual 113 / 316 pages 119 CN 121729488 A Instruction Manual 114 / 316 pages 120 CN 121729488 A Instruction Manual 115 / 316 pages 121 CN 121729488 A Instruction Manual 116 / 316 pages 122 CN 121729488 A Instruction Manual 117 / 316 pages 123 CN 121729488 A Instruction Manual 118 / 316 pages 124 CN 121729488 A Instruction Manual 119 / 316 pages 125 CN 121729488 A Instruction Manual 120 / 316 pages 126 CN 121729488 A Instruction Manual 121 / 316 Page 127 CN 121729488 A Instruction Manual 122 / 316 Page 128 CN 121729488 A Instruction Manual 123 / 316 Page 129 CN 121729488 A Instruction Manual 124 / 316 Page 130 CN 121729488 A Instruction Manual 125 / 316 Page 131 CN 121729488 A Instruction Manual 126 / 316 Page 132 CN 121729488 A Instruction Manual 127 / 316 Page 133 CN 121729488 A Instruction Manual 128 / 316 Page 134 CN 121729488 A Instruction Manual 129 / 316 Page 135 CN 121729488 A Instruction Manual 130 / 316 Page 136 CN121729488 A Instruction Manual 131 / 316 pages 137 CN 121729488 A Instruction Manual 132 / 316 pages 138 CN 121729488 A Instruction Manual 133 / 316 pages 139 CN 121729488 A Instruction Manual 134 / 316 pages 140 CN 121729488 A Instruction Manual 135 / 316 pages 141 CN 121729488 A Instruction Manual 136 / 316 pages 142 CN 121729488 A Instruction Manual 137 / 316 pages 143 CN 121729488 A Instruction Manual 138 / 316 pages 144 CN 121729488 A Instruction Manual 139 / 316 pages 145 CN 121729488 A Instruction Manual 140 / 316 Page 146 CN 121729488 A Instruction Manual 141 / 316 Page 147 CN 121729488 A Instruction Manual 142 / 316 Page 148 CN 121729488 A Instruction Manual 143 / 316 Page 149 CN 121729488 A Instruction Manual 144 / 316 Page 150 CN 121729488 A Instruction Manual 145 / 316 Page 151 CN 121729488 A Instruction Manual 146 / 316 Page 152 CN 121729488 A Instruction Manual 147 / 316 Page 153 CN 121729488 A Instruction Manual 148 / 316 Page 154 CN 121729488 A Instruction Manual 149 / 316 Page 155 CN 121729488 A Instruction Manual 150 / 316 pages 156 CN 121729488 A Instruction Manual 151 / 316 pages 157 CN 121729488 A Instruction Manual 152 / 316 pages 158 CN 121729488 A Instruction Manual 153 / 316 pages 159 CN 121729488 A Instruction Manual 154 / 316 pages 160 CN 121729488 A Instruction Manual 155 / 316 pages 161 CN 121729488 A Instruction Manual 156 / 316 pages 162 CN 121729488 A Instruction Manual 157 / 316 pages 163 CN 121729488 A Instruction Manual 158 / 316 pages 164 CN 121729488 APages 159 / 316, 165 CN 121729488 A, Instruction Manual; Pages 160 / 316, 166 CN 121729488 A, Instruction Manual; Pages 161 / 316, 167 CN 121729488 A, Instruction Manual; Pages 162 / 316, 168 CN 121729488 A, Instruction Manual; Pages 163 / 316, 169 CN 121729488 A, Instruction Manual; Pages 164 / 316, 170 CN 121729488 A, Instruction Manual; Pages 165 / 316, 171 CN 121729488 A, Instruction Manual; Pages 166 / 316, 172 CN 121729488 A, Instruction Manual; Pages 167 / 316, 173 CN 121729488 A, Instruction Manual; Pages 168 / 316, 174 CN 121729488 A Instruction Manual 169 / 316 pages 175 CN 121729488 A Instruction Manual 170 / 316 pages 176 CN 121729488 A Instruction Manual 171 / 316 pages 177 CN 121729488 A Instruction Manual 172 / 316 pages 178 CN 121729488 A Instruction Manual 173 / 316 pages 179 CN 121729488 A Instruction Manual 174 / 316 pages 180 CN 121729488 A Instruction Manual 175 / 316 pages 181 CN 121729488 A Instruction Manual 176 / 316 pages 182 CN 121729488 A Instruction Manual 177 / 316 pages 183 CN 121729488 A Instruction Manual 178 / 316 Page 184 CN 121729488 A Instruction Manual 179 / 316 Page 185 CN 121729488 A Instruction Manual 180 / 316 Page 186 CN 121729488 A Instruction Manual 181 / 316 Page 187 CN 121729488 A Instruction Manual 182 / 316 Page 188 CN 121729488 A Instruction Manual 183 / 316 Page 189 CN 121729488 A Instruction Manual 184 / 316 Page 190 CN 121729488 A Instruction Manual 185 / 316 Page 191 CN 121729488 A Instruction Manual 186 / 316 Page 192 CN 121729488 A Instruction Manual 187 / 316 Page 193 CN121729488 A Instruction Manual 188 / 316 pages 194 CN 121729488 A Instruction Manual 189 / 316 pages 195 CN 121729488 A Instruction Manual 190 / 316 pages 196 CN 121729488 A Instruction Manual 191 / 316 pages 197 CN 121729488 A Instruction Manual 192 / 316 pages 198 CN 121729488 A Instruction Manual 193 / 316 pages 199 CN 121729488 A Instruction Manual 194 / 316 pages 200 CN 121729488 A Instruction Manual 195 / 316 pages 201 CN 121729488 A Instruction Manual 196 / 316 pages 202 CN 121729488 A Instruction Manual 197 / 316 Page 203 CN 121729488 A Instruction Manual 198 / 316 Page 204 CN 121729488 A Instruction Manual 199 / 316 Page 205 CN 121729488 A Instruction Manual 200 / 316 Page 206 CN 121729488 A Instruction Manual 201 / 316 Page 207 CN 121729488 A Instruction Manual 202 / 316 Page 208 CN 121729488 A Instruction Manual 203 / 316 Page 209 CN 121729488 A Instruction Manual 204 / 316 Page 210 CN 121729488 A Instruction Manual 205 / 316 Page 211 CN 121729488 A Instruction Manual 206 / 316 Page 212 CN 121729488 A Instruction Manual 207 / 316 pages 213 CN 121729488 A Instruction Manual 208 / 316 pages 214 CN 121729488 A Instruction Manual 209 / 316 pages 215 CN 121729488 A Instruction Manual 210 / 316 pages 216 CN 121729488 A Instruction Manual 211 / 316 pages 217 CN 121729488 A Instruction Manual 212 / 316 pages 218 CN 121729488 A Instruction Manual 213 / 316 pages 219 CN 121729488 A Instruction Manual 214 / 316 pages 220 CN 121729488 A Instruction Manual 215 / 316 pages 221 CN 121729488 APages 216 / 316, 222 CN 121729488 A, Instruction Manual; Pages 217 / 316, 223 CN 121729488 A, Instruction Manual; Pages 218 / 316, 224 CN 121729488 A, Instruction Manual; Pages 219 / 316, 225 CN 121729488 A, Instruction Manual; Pages 220 / 316, 226 CN 121729488 A, Instruction Manual; Pages 221 / 316, 227 CN 121729488 A, Instruction Manual; Pages 222 / 316, 228 CN 121729488 A, Instruction Manual; Pages 223 / 316, 229 CN 121729488 A, Instruction Manual; Pages 224 / 316, 230 CN 121729488 A, Instruction Manual; Pages 225 / 316, 231 CN 121729488 A Instruction Manual 226 / 316 pages 232 CN 121729488 A Instruction Manual 227 / 316 pages 233 CN 121729488 A Instruction Manual 228 / 316 pages 234 CN 121729488 A Instruction Manual 229 / 316 pages 235 CN 121729488 A Instruction Manual 230 / 316 pages 236 CN 121729488 A Instruction Manual 231 / 316 pages 237 CN 121729488 A Instruction Manual 232 / 316 pages 238 CN 121729488 A Instruction Manual 233 / 316 pages 239 CN 121729488 A Instruction Manual 234 / 316 pages 240 CN 121729488 A Instruction Manual 235 / 316 Page 241 CN 121729488 A Instruction Manual 236 / 316 Page 242 CN 121729488 A Instruction Manual 237 / 316 Page 243 CN 121729488 A Instruction Manual 238 / 316 Page 244 CN 121729488 A Instruction Manual 239 / 316 Page 245 CN 121729488 A Instruction Manual 240 / 316 Page 246 CN 121729488 A Instruction Manual 241 / 316 Page 247 CN 121729488 A Instruction Manual 242 / 316 Page 248 CN 121729488 A Instruction Manual 243 / 316 Page 249 CN 121729488 A Instruction Manual 244 / 316 Page 250 CN121729488 A Instruction Manual 245 / 316 pages 251 CN 121729488 A Instruction Manual 246 / 316 pages 252 CN 121729488 A Instruction Manual 247 / 316 pages 253 CN 121729488 A Instruction Manual 248 / 316 pages 254 CN 121729488 A Instruction Manual 249 / 316 pages 255 CN 121729488 A Instruction Manual 250 / 316 pages 256 CN 121729488 A Instruction Manual 251 / 316 pages 257 CN 121729488 A Instruction Manual 252 / 316 pages 258 CN 121729488 A Instruction Manual 253 / 316 pages 259 CN 121729488 A Instruction Manual 254 / 316 Page 260 CN 121729488 A Instruction Manual 255 / 316 Page 261 CN 121729488 A Instruction Manual 256 / 316 Page 262 CN 121729488 A Instruction Manual 257 / 316 Page 263 CN 121729488 A Instruction Manual 258 / 316 Page 264 CN 121729488 A Instruction Manual 259 / 316 Page 265 CN 121729488 A Instruction Manual 260 / 316 Page 266 CN 121729488 A Instruction Manual 261 / 316 Page 267 CN 121729488 A Instruction Manual 262 / 316 Page 268 CN 121729488 A Instruction Manual 263 / 316 Page 269 CN 121729488 A Instruction Manual 264 / 316 pages 270 CN 121729488 A Instruction Manual 265 / 316 pages 271 CN 121729488 A Instruction Manual 266 / 316 pages 272 CN 121729488 A Instruction Manual 267 / 316 pages 273 CN 121729488 A Instruction Manual 268 / 316 pages 274 CN 121729488 A Instruction Manual 269 / 316 pages 275 CN 121729488 A Instruction Manual 270 / 316 pages 276 CN 121729488 A Instruction Manual 271 / 316 pages 277 CN 121729488 A Instruction Manual 272 / 316 pages 278 CN 121729488 APages 273 / 316, 279 CN 121729488 A Specification Pages 274 / 316, 280 CN 121729488 A Specification Pages 275 / 316, 281 CN 121729488 A Specification Pages 276 / 316, 282 CN 121729488 A
[0196] Table 3: Exemplary direct repeat sequences described in this disclosure.
[0197] Specification Pages 277 / 316, 283 CN 121729488 A
[0198] In some embodiments, this disclosure provides an engineered, non-naturally occurring crRNA or a variant thereof, wherein the crRNA comprises a nucleotide sequence having at least 90% identity with any of the sequences in SEQ ID NO: 341-417 (Table 3). In some embodiments, the crRNA comprises a nucleotide sequence having at least 95% or 98% identity with any of the sequences in SEQ ID NO: 341-417. In some embodiments, the crRNA comprises any of the sequences in SEQ ID NO: 341-417.
[0199] The following non-limiting examples provide further illustration of embodiments of this disclosure. Those skilled in the art should understand that the techniques disclosed in the following examples represent effective methods found in practicing this disclosure and can therefore be considered as examples of their modes of practice. Those skilled in the art should understand, in consideration of this disclosure, that many changes can be made in the specific embodiments disclosed and similar or identical results can still be obtained without departing from the spirit and scope of this disclosure.
[0200] Example 1: Metagenomic Analysis Method for Proteins
[0201] Metagenomic sequence data from public databases were retrieved using a hidden Markov model based on known Cas protein sequences (including type II Cas effector proteins). Potential active sites were determined by comparing the identified CRISPR-Cas proteins with known proteins. After screening hundreds of potential sequences, this metagenomic workflow ultimately identified type II Cas proteins, as detailed in Table 1.
[0202] A phylogenetic tree was generated using MUSCLE 3 (Veen et al., 2020) to explore the relationships between orthologs at the primary amino acid level. This study utilized hundreds of Class 2 Type II-A / B / C sequences from the National Center for Biotechnology Information (NCBI) and various publications and patents. Notably, the phylogenetic tree indicates that the Cas protein described in this invention is distinct from previously known Cas proteins. The Type II Cas protein detailed in this disclosure shows low similarity to other known Cas proteins.
[0203] The structure of type II Cas protein was modeled using AlphaFold2. By annotating its domain arrangement, it was found that the Cas protein disclosed in this invention contains a RuvC domain, a BH (bridged helix) domain, a REC domain, an HNH domain, and a CTD (C-terminal domain). Notably, the RuvC domain contains three distinct RuvC subdomains and a BH domain. Figure 1 shows a visualization of some example protein structures.
[0204] Example 2: Experimental Protocol for Predicting RNA Folding
[0205] The predicted RNA folding of the putative guide RNA sequence matching the Cas protein was calculated using the RNAfold web server developed by Lorenz et al. in 2011.
[0206] Example 3: Identification of PAM in Mammalian Cell Lines
[0207] In one set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). Mix 450 µL of cells at a density of 120,000 cells / well with 50 μL of a mixture containing Lipofectamine™ 3000 (ThermoFisher Scientific, catalog number L3000008), Opti-Mem (volume adjusted to 50 μL), 1 μL dsODN (2.5 pM), 100 ng (approximately 1 μL) of psgRNA carrying an sgRNA scaffold (Table 6) and the coding sequence for the Humanspacer3 spacer sequence (SEQ ID NO: 885, Table 4), and 400 ng (approximately 1 μL) of pCas protein particles carrying the coding sequences for Cas9 CDS, NLS, and FLAG (Table 2) according to the manufacturer's protocol. The cell mixture was then seeded into 24-well plates and cultured at 37°C and 5% CO2. 10 pM dsODN was prepared by annealing dsODN-Top and dsODN-BoT oligonucleotides before transfection.
[0208] 72 hours after transfection, the supernatant was removed and the cell layer was washed with PBS. Genomic DNA was then extracted from each well of a 24-well plate using DNA extraction solution (Denogen (Beijing) Biotechnology Co., Ltd., catalog number DNS033-48) according to the manufacturer's protocol. All DNA samples (500 ng, 260 / 280 value: 1.8–2.0) were analyzed by Guide-Seq NGS.
[0209] The basic method for Guide-Seq library preparation was described by Nikolay et al. (Nat. Protoc. 2021).Okay. The extracted DNA sample was first cleaved using the KAPA Frag Kit (Catalog No. KK8602, Roche). The cleaved DNA was purified and then phosphorylated using T4 polynucleotide kinase (Catalog No. M0201S, NEB). The SS5 adapter (generated by annealing 10 μM SS5TOP oligonucleotide and 10 μM SS5BTM oligonucleotide) was ligated to the cleaved DNA using the Quick Ligation™ Kit (Catalog No. M2200S, NEB), followed by two-step off-target PCR to add the chemical modifications required for sequencing.
[0210] Off-target PCR 1 was performed using Platinum™ Taq DNA polymerase (Catalog No. 15966005, Invitrogen) with GSP1 (a mixture of GSP1-Top and GSP1-BoT) and Y_XX oligonucleotides. Off-target PCR2 was performed using Platinum™ Taq DNA polymerase with GSP2 (a mixture of GSP2-TopA / B / C and GSP1-BoTA / B / C), Y_XX (same as PCR1), and i753_XX oligonucleotides. DNA products from each of the above steps were purified using SPRI Select (catalog number B23318, Beckman Coulter). The final library was quantified by qPCR and sequenced on an Illumina NextSeq 1000. Reads were aligned to a reference genome after removing low-quality fractions. The Q30 rate was greater than 0.9. Read lengths were between 130 bp and 140 bp. The resulting file containing the read data was mapped to a reference genome (BAM file), where reads overlapping with the target region were selected.
[0211] Table 4: Nucleotide sequences mentioned in the example. Instruction manual 279 / 316 pages 285 CN 121729488 A
[0212]
[0213] Note: “p” indicates phosphorylation modification; “ ” indicates phosphate thioester (PS) bond; “N” indicates any natural or non-natural nucleotide.
[0214] Figure 2 and Table 5 show the PAM preferences of some example Cas proteins when using the corresponding sgRNA scaffold in the HEK293 cell line.
[0215] Table 5: PAM preferences of example Cas proteins Instruction manual 280 / 316 pages 286 CN 121729488 A
[0216] Example 4: Screening for in vitro editing efficiency of CRISPR-Cas system in mammalian cell lines
[0217] In one set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). existTwenty-four hours prior to transfection, 250 μL of cells at a density of 50,000 cells / well were seeded into 48-well plates. Following the manufacturer's protocol, cells were transfected using a lipid complex containing Lipofectamine™ 3000 (0.4 μL / well), P3000 (2 μL / well), pgRNA / pCas protein particles (125 ng / well and 375 ng / well, respectively), and Opti-Mem (to a final volume of 25 μL / well). The seeded cells were then incubated statically and adherently in a tissue culture incubator at 37°C and 5% CO2 for 72 hours. The nucleotide sequences of the pgRNAs used in this embodiment contain sequences encoding the corresponding Cas sgRNA scaffolds (Table 6; SEQ ID NO: 465 (GEBx0305), SEQ ID NO: 451 (GEBx0308), SEQ ID NO: 454 (GEBx0308-Rq-V3), SEQ ID NO: 478 (spCas9), SEQ ID NO: 491, and SEQ ID NO: 548) and the corresponding spacer sequences (Table 7).
[0218] 72 hours after transfection, the supernatant was removed and the cell layer was washed with PBS. Genomic DNA was then extracted from each well of a 24-well plate using DNA extraction solution (Denogen (Beijing) Biotechnology Co., Ltd., catalog number DNS033-48) according to the manufacturer's protocol. All DNA samples (500 ng, 260 / 280 value: 1.8–2.0) were subjected to amplicon NGS analysis.
[0219] To quantitatively determine the editing efficiency at target locations in the genome, NGS was used to identify insertions and deletions introduced by gene editing. Primers for NGS were designed to be positioned around the target region of the endogenous gene. Additional PCR was performed according to the manufacturer's protocol (Illumina) to add the chemical modifications required for sequencing. Amplicon sequencing was performed on an Illumina iSeq 100. Reads were aligned to a reference genome after removing low-quality fractions. The Q30 rate was greater than 0.9. Read lengths were between 130 bp and 140 bp. The generated file containing read data was mapped to a reference genome (BAM file), where reads overlapping with the target region were selected, and the number of wild-type reads and reads containing insertions, substitutions, or deletions were calculated. The number of reads mapped to the reference genome exceeded 1000.
[0220] Table 6: sgRNA scaffold sequences of the corresponding Cas proteins (Instruction manual 282 / 316 pages 288 CN 121729488 A; Instruction manual 283 / 316 pages 289 CN 121729488 A; Instruction manual 284 / 316 pages 290 CN)121729488 A Instruction Manual 285 / 316 pages 291 CN 121729488 A Instruction Manual 286 / 316 pages 292 CN 121729488 A Instruction Manual 287 / 316 pages 293 CN 121729488 A
[0221] The interval sequence is located at the 5' end of the stent. The 3' end of each interval sequence is directly connected to the 5' end of the subsequent stent sequence, forming a typical repeat-interval pattern.
[0222] Table 7: Example spacer sequence Specification 288 / 316 pages 294 CN 121729488 A Specification 289 / 316 pages 295 CN 121729488 A Specification 290 / 316 pages 296 CN 121729488 A Specification 291 / 316 pages 297 CN 121729488 A Specification 292 / 316 pages 298 CN 121729488 A Specification 293 / 316 pages 299 CN 121729488 A
[0223] Figure 3 shows the insertion and deletion (indel) levels of GEBx0305 against 16 targets with GGAAAA-PAM in the HEK293T cell line. In this experiment, the sgRNA sequence used for GEBx0305 included the GEBx0305-HPT-V1 scaffold (WT, SEQ ID NO: 465) and 20nt spacer sequences (SEQ ID NO: 571, 573, 575, 577, 579, 581, 583, 585, 587, 589, 591, 593, 595, 597, 599, and 601). SpCas9 targeting corresponding sites (SEQ ID NO: 572, 574, 576, 578, 580, 582, 584, 586, 588, 590, 592, 594, 596, 598, 600, and 602) served as a positive control. Figure 4 shows the insertion and deletion levels of GEBx0308 targeting 19 sites with GGTACT-PAM in the HEK293T cell line. The sgRNA sequences used for GEBx0308 in this experiment included the GEBx0308-PT-V1 scaffold (WT, SEQ ID NO: 451) and 20nt spacer sequences (SEQ ID NO: 611, 613, 615, 617, 619, 621, 623, 625, 627, 629, 631, 633, 635, 637, 639, 641, 643, 645, 647, and 649). SpCas9 targets the corresponding sites (SEQ ID NO: 612, 614, 616, 618, 620, 622, 624).626, 628, 630, 632, 634, 636, 638, 640, 642, 644, 646, 648 and 650) were used as positive controls.
[0224] Figure 5 shows the insertion and deletion levels of GEBx0308 targeting 15 sites with CATACT-PAM in the HEK293T cell line. The psgRNA sequence used for GEBx0308 in this experiment contains the GEBx0308-Rq-V3 scaffold (MO, SEQ ID NO: 454) and a 20nt spacer sequence (SEQ ID NO: 659-674).
[0225] Figure 6 shows the insertion and deletion levels of GEBx0328 targeting 20 sites with NNGCCT-PAM in the HEK293T cell line. The sgRNA sequence used for GEBx0328 in this experiment contained the GEBx0328-PT-V1 scaffold (WT, SEQ ID NO: 491) and a 20nt spacer sequence (SEQ ID NO: 765-784). GEBx0328 showed moderate insertion-deletion activity at 20 target sites.
[0226] Figures 7 and 8 show the insertion-deletion levels of GEBx0361 at 19 target sites with GGTACC or TGTACC PAM in the HEK293T cell line. The sgRNA sequence used for GEBx0361 in this experiment contained the GEBx0361-HPT-V2 scaffold (SEQ ID NO: 548) and a 20nt spacer sequence (SEQ ID NO: 817-835). GEBx0361 showed moderate insertion-deletion activity at 19 target sites.
[0227] Example 5: Optimization of an Example CRISPR / Cas System
[0228] To further improve the editing efficiency of Cas proteins, guide sequences and sgRNA scaffolds of different lengths were tested. In one set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 24 hours before transfection, 250 μL of cells at a density of 50,000 cells / well were seeded into 48-well plates. Cells were transfected using a lipid complex containing Lipofectamine™ 3000 (0.4 μL / well), P3000 (2 μL / well), pgRNA / pCas protein particles (125 ng / well and 375 ng / well, respectively), and Opti-Mem (volume supplemented to 25 μL / well) according to the manufacturer's protocol. The inoculated cells were incubated at 37°C and 5%The tissue culture was statically incubated and adhered to the walls of a CO2 incubator for 72 hours. The nucleotide sequence of the pgRNA used in this embodiment includes the sequence encoding the corresponding sgRNA scaffold (SEQ ID NO: 465 (GEBx0305), SEQ ID NO: 451 (GEBx0308) or SEQ ID NO: 491 (GEBx0328)) and the corresponding spacer sequence (SEQ ID NO: 575, 599 or 695-707 for GEBx0305; SEQ ID NO: 625, 676 or 708-721 for GEBx0308; SEQ ID NO: 769, 782, 785-798 for GEBx0328).
[0229] As shown in Figure 9, the insertion / deletion level of GEBx0305 varies depending on the length of the guide sequence. The sgRNA sequences used in this experiment contained a GEBx0305-HPT-V1 scaffold (WT, SEQ ID NO: 465) and a CFTR-NGGAAAA-T3 or POLQ-NGGAAAA-T5 spacer sequence, ranging in length from 18 nt to 25 nt (SEQ ID NO: 575, 599, or 695-707). The CFTR-NGGAAAA-T3-23 nt and POLQ-NGGAAAA-T5-25 nt spacer sequences showed the highest insertion / deletion levels in each group.
[0230] Figure 10 (A and B) shows that the insertion / deletion levels of GEBx0308 varied with the length of the guide sequence. The sgRNA sequences used in this experiment contained a GEBx0308-PT-V1 scaffold (WT, SEQ ID NO: 451) and a CFTR-NGGTACT-T4 or CD34-TATACT-T2 spacer sequence, ranging in length from 18 nt to 25 nt (SEQ ID NO: 625, 676, or 708-721). The CFTR-NGGTACT-T4-21 nt spacer sequence showed a higher level of insertion / deletion than other spacer sequences of different lengths.
[0231] Figure 11 (A and B) shows that the insertion / deletion level of GEBx0328 varies with the length of the guide sequence. The sgRNA sequences used in this experiment contained a GEBx0328-PT-V1 scaffold (WT, SEQ ID NO: 491) and a TTR-NGGCCT-T1 or CFTR-NGGCCT-T5 spacer sequence, ranging in length from 18 nt to 25 nt (SEQ ID NO: 769, 782, 785-798). The TTR-NGGCCT-T1-24nt and CFTR-NGGCCT-T5-24nt interval sequences showed the highest levels of insertion / deletion in each group.
[0232] In another set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 24 hours prior to transfection, 100 μL of cells at a density of 25,000 cells / well were seeded into 96-well plates. Cells were transfected using a lipid complex containing Lipofectamine™ 3000 (0.4 μL / well), P3000 (2 μL / well), pCas protein-gRNA plasmid (300 ng / well), and Opti-Mem (to a final volume of 25 μL / well) according to the manufacturer's protocol. The seeded cells were then statically and adherently cultured in a tissue culture incubator at 37°C and 5% CO2 for 72 hours. The nucleotide sequence of the Cas protein-gRNA used in this embodiment consists of Cas CDS, Cas sgRNA scaffold (Table 6), and corresponding spacer sequences (Table 7).
[0233] In another set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 24 hours before transfection, 100 μL of cells at a density of 25,000 cells / well were seeded into 96-well plates. Cells were transfected using a lipid complex containing Lipofectamine™ 3000 (0.4 μL / well), P3000 (2 μL / well), pCas protein-gRNA plasmid (300 ng / well), and Opti-Mem (volume supplemented to 25 μL / well) according to the manufacturer's protocol. The seeded cells were statically incubated and adhered in a tissue culture incubator at 37°C and 5% CO2 for 72 hours. The nucleotide sequence of the Cas protein-gRNA used in this example consists of Cas CDS, Cas sgRNA scaffold, and corresponding spacer sequences.
[0234] Figure 12 (A and B) shows the insertion and deletion levels of GEBx0305 using the modified RNA scaffold when targeting endogenous genes. The sgRNA sequence used in this experiment contains a scaffold of GEBx0305-HPT-V1 (WT, SEQ ID NO: 465), GEBx0305-Rq-V2 (M0, SEQ ID NO: 466), or GEBx0305-M1 to M6 (SEQ ID NO: 467-472), and a 20nt spacer sequence, targeting 8 endogenous gene sites (SEQ ID NO: 467-472).575, 579, 585, 587, 589, 591, 599, and 601). The GEBx0305-M5 and M6 scaffolds showed the highest mean insertion / deletion values.
[0235] Figure 13 (A and B) shows the insertion / deletion levels of GEBx0308 when using modified RNA scaffolds to target endogenous genes. The sgRNA sequences used in this experiment contained GEBx0308-PT-V1 (WT, SEQ ID NO: 451), GEBx0308-Rq-V3 (M0, SEQ ID NO: 454), or GEBx0308-M1 to M6 (SEQ ID NO: 455-460) scaffolds, along with a 20nt spacer sequence, targeting 7 endogenous gene sites (SEQ ID NO: 621, 625, 627, 631, 637, 639, 641). The GEBx0308-M3 and M4 scaffolds showed the highest mean insertion / deletion values.
[0236] Figure 14 (A and B) shows the insertion and deletion levels of GEBx0305 targeting endogenous genes under optimized conditions. The optimized sgRNA sequence used in this experiment contains a GEBx0305-M5 (SEQ ID NO: 471) scaffold and a 21nt spacer sequence, targeting 27 endogenous gene sites (SEQ ID NO: 652, 654, 655, 658, 696, 703, 722-725, 729-745). Compared with WT sgRNA (GEBx0305-HPT-V1 scaffold (SEQ ID NO: 465) and 20nt spacer sequences (SEQ ID NO: 571, 573, 575, 577, 579, 581, 583, 585, 587, 589, 591, 593, 595, 597, 599, 651-658, 746-749)), optimized sgRNA significantly improved the insertion and deletion levels of GEBx0305 (P<0.0001).
[0237] Figure 15 (A and B) shows the insertion and deletion levels of GEBx0308 targeting endogenous genes under optimized conditions. The optimized sgRNA sequence used in this experiment contained a GEBx0308-M4 (SEQ ID NO: 458) scaffold and a 21nt spacer sequence, targeting 25 endogenous gene sites (SEQ ID NO: 710, 659, 665, 667, 668, 672, 674, 726, 727, 728, 750-764). This was compared with the WT sgRNA (GEBx0308-PT-V1 scaffold (SEQ ID NO: 451) and 20nt spacer sequence (SEQ ID NO: 611, 613, ...).Compared to 615, 617, 619, 621, 623, 625, 627, 629, 631, 633, 635, 637, 639, 641, 645, 647, 659, 662, 663, 665, 667, 668, 672, 674), optimized sgRNA slightly increased the insertion and deletion levels of GEBx0308 (P=0.03).
[0238] Figure 16 (A and B) shows the insertion and deletion levels of GEBx0328 when using modified RNA scaffolds to target endogenous genes. The sgRNA sequences used in this experiment contained scaffolds of GEBx0328-PT-V1 (WT, SEQ ID NO: 491), GEBx0328-Rq-V1 (M0, SEQ ID NO: 492), or GEBx0328-M1 to M8 (SEQ ID NO: 493-500), and 20nt spacer sequences, targeting 8 endogenous gene sites (SEQ ID NO: 765, 767, 769, 770, 771, 776, 782, and 783). The GEBx0328-M6 scaffold showed the highest mean insertion / deletion.
[0239] Figure 17 (A and B) shows the insertion / deletion levels of GEBx0328 targeting endogenous genes under optimized conditions. The optimized sgRNA sequence used in this experiment contained a GEBx0328-M6 (SEQ ID NO: 498) scaffold and a 24nt spacer sequence, targeting 19 endogenous gene loci (SEQ ID NO: 790, 797, 799-807, 809-816). Compared with WT sgRNA (GEBx0328-PT-V1 scaffold and 20nt spacer sequence), the optimized sgRNA significantly improved the insertion and deletion levels of GEBx0328 (P=0.0002).
[0240] Example 6: In vitro gene editing using RNA In one set of experiments, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 24 hours before transfection, 100 μL of cells at a density of 25,000 cells / well were seeded into 96-well plates. Following the manufacturer's protocol, cells were transfected using a lipid complex containing Lipofectamine™ RNAiMAX (Invitrogen™) and RNA (20 ng sgRNA, sgRNA:mRNA = 1:1–1:8 w / w), and Opti-Mem (volume adjusted to 25 μL / well). The seeded cells were then statically incubated in a 37°C, 5% CO2 tissue culture incubator for 72 hours. The sgRNA and mRNA sequences are shown in Table 8. The mRNA used in this example is N1-methyl-pseudomonas aeruginosa.Uracil modified with uridine.
[0241] Primary human hepatocytes (PHH) were thawed and suspended in hepatocyte thawing medium (Lonza, catalog number MCHT50) supplemented with additives, and then centrifuged at 100 g for 10 minutes. The supernatant was discarded, and the cell pellet was resuspended in hepatocyte plating medium (Lonza, catalog number MP100) with 10% fetal bovine serum added. After cell counting, the cells were seeded at a density of 40,000 cells / well in 96-well ultra-low adsorption cell culture plates (Liver Biotech, catalog number LV-ULA002-96W). The seeded cells were incubated statically and adherently in a tissue culture incubator at 37°C and 5% CO2 for 24 hours. After incubation, the cell monolayer formation was checked, and the medium was replaced with hepatocyte medium (Lonza, catalog number CC-3198) with 10% fetal bovine serum added. Following the manufacturer's protocol, cells were transfected using a lipid complex containing Lipofectamine™ RNAiMAX (Invitrogen™) and RNA (30 ng gRNA, gRNA:mRNA = 1:1 - 1:16 w / w) and Opti-Mem (volume supplemented to 25 μL / well). Seeded cells were statically and adherently cultured in a tissue culture incubator at 37°C and 5% CO2 for 72 hours.
[0242] Genomic DNA extraction and NGS methods were the same as described in Example 4.
[0243] Figure 18 shows the insertion / deletion levels of GEBx0305 targeting endogenous genes after transfection of HEK293T cells with a lipid complex containing a fixed amount (20 ng) of sgRNA targeting the POLQ-NGGAAAA-T5 site (SEQ ID NO: 855) and different proportions of GEBx0305 mRNA (SEQ ID NO: 851). Equal doses of SpCas9 mRNA and sgRNA (SEQ ID NO: 853 and 854) were used as positive controls. GEBx0305 showed insertion / deletion efficiency comparable to SpCas9.
[0244] Figure 19 shows the insertion / deletion levels of GEBx0308 targeting endogenous genes after transfection of a lipid complex containing a fixed amount (20 ng) of sgRNA targeting the CFTR-NGGTACT-T4 or POLQ-NGGTACT-T1 site (SEQ ID NO: 856 / 857) and different proportions of GEBx0308 mRNA (SEQ ID NO: 852) in HEK293T cells.
[0245] Figure 20 shows the insertion / deletion levels of GEBx0308 targeting endogenous genes after transfection of primary human hepatocytes (PHH) containing a fixed amount (20 ng) of sgRNA targeting the POLQ-NGGAAAA-T5 site (SEQ ID NO: 855) and different proportions of GEBx0305 mRNA (SEQ ID NO: 852).Following transfection of a lipid complex containing a fixed amount (20 ng) of sgRNA targeting the POLQ-NGGTACT-T1 site (SEQ ID NO: 857) and different proportions of GEBx0308 mRNA (SEQ ID NO: 852, Table 8) into primary human hepatocytes (PHH), the levels of insertion and deletion of endogenous genes targeted by GEBx0305 were observed.
[0246] Figure 21 shows the levels of insertion and deletion of endogenous genes targeted by GEBx0308 after transfection of a lipid complex containing a fixed amount (20 ng) of sgRNA targeting the POLQ-NGGTACT-T1 site (SEQ ID NO: 857) and different proportions of GEBx0308 mRNA (SEQ ID NO: 852, Table 8) into primary human hepatocytes (PHH).
[0247] Table 8: Example mRNA and gRNA sequences for gene editing 297 / 316 pages 303 CN 121729488 A 298 / 316 pages 304 CN 121729488 A 299 / 316 pages 305 CN 121729488 A 300 / 316 pages 306 CN 121729488 A 301 / 316 pages 307 CN 121729488 A
[0248] Example 7: Off-target analysis in cell lines using GUIDE-Seq
[0249] GUIDE-Seq utilizes dsODN to insert double-strand break sites generated by CRISPR / Cas. HEK293T cells were cultured in advanced DMEM medium supplemented with 5% fetal bovine serum (Gibco™). 24 hours prior to transfection, cells were seeded into 24-well plates at a density of 100,000 cells / well. Following the manufacturer's protocol, 400 ng pCas protein plasmid, 150 ng pgRNA plasmid, and 2.5 pmol dsODN were transfected using Lipofectamine 3000 (Invitrogen™), cultured at 37°C and 5% CO2, and harvested on day 3 post-transfection.
[0250] For GUIDE-Seq library construction, 500 ng genomic DNA was used. Briefly, the DNA was fragmented using the KAPA Frag Kit (KAPA Biosystems), followed by aptamer ligation and two rounds of semi-nested PCR enrichment of the dsODN integration fragment. The final sequencing library was quantified using KAPA Library Quantification Kits and sequenced on an Illumina NextSeq 1000 system. Data demultiplexing for Index 1 is performed using bcl2fq (version 2.19), followed by index demultiplexing via a custom script.2. Reuse was performed, aptamer trimming was done using the BBduk tool, and analysis was conducted using GUIDE-seq software. In short, FASTQ files with unique molecular indexes (UMIs) were integrated to generate UMI-consistent sequences and aligned with the human reference genome (hg19) using BWA MEM. High-quality alignment (MAPQ ≥ 50) was used to identify genomic sites containing dsODNs as potential off-target sites. Candidate sites with up to 6 mismatches with the corresponding target protospacer sequence were identified as true off-target sites.
[0251] Figure 22 shows the GUIDE-seq insertion sites for GEBx0305. No detectable off-target sites were detected at site 1 (CFTR-NGGAAAA-T5) and site 2 (EMX1-NGGAAAA-T5).
[0252] Figure 23 shows the GUIDE-seq insertion sites for GEBx0308. No off-target sites were detected at site 1 (CD34-NGGTACT-T4) and site 2 (POLQ-NGGTACT-T1).
[0253] Figure 24 shows the GUIDE-seq insertion sites of GEBx0328. No detectable off-target sites were detected at site 2 (CFTR-NGGCCT-T5), while only one off-target site was detected at site 1 (CFTR-NGGCCT-T3).
[0254] Example 8: Detection of base editing efficiency of nCas protein
[0255] In order to generate Cas base editor, E. coli tRNA adenosine deaminase (TadA-8e) was fused to the N-terminus of GEBx0305, GEBx0308, and GEB328 nickases (GEBx0305-D11A, GEBx0308-D10A, GEB328-D12A) and linked using an ABE-connector (SEQ ID NO: 868). HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). For reverse transfection, HEK293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (Gibco™). 24 hours prior to transfection, 100 μL of cells at a density of 25,000 cells / well were seeded into 96-well plates. Cells were transfected using a lipid complex containing Lipofectamine™ 3000 (0.4 μL / well), P3000 (2 μL / well), pABE-gRNA plasmid (300 ng / well), and Opti-Mem (to a final volume of 25 μL / well) according to the manufacturer's protocol. The seeded cells were then statically incubated and adherent in a tissue culture incubator at 37°C and 5% CO2 for 72 hours. The methods used in this example...The nucleotide sequence of the pABE / gRNA plasmid contains Cas-ABE CDS (SEQ ID NO: 861, 862, 863, Table 9), Cas sgRNA scaffold, and corresponding spacer sequences. After transfection for 72 hours, the supernatant was removed and the cell layer was washed with PBS. Genomic DNA was then extracted from each well of a 24-well plate using DNA extraction solution (Denogen (Beijing) Biotechnology Co., Ltd., catalog number DNS033-48) according to the manufacturer's protocol. All DNA samples were subjected to amplicon NGS analysis.
[0256] Table 9: CDS sequences of example ABE sequences. Instruction manual 303 / 316 pages 309 CN 121729488 A
[0257] Instruction manual 304 / 316 pages 310 CN 121729488 A Instruction manual 305 / 316 pages 311 CN 121729488 A Instruction manual 306 / 316 pages 312 CN 121729488 A Instruction manual 307 / 316 pages 313 CN 121729488 A Instruction manual 308 / 316 pages 314 CN 121729488 A Instruction manual 309 / 316 pages 315 CN 121729488 A Instruction manual 310 / 316 pages 316 CN 121729488 A Instruction manual 311 / 316 pages 317 CN 121729488 A Instruction manual 312 / 316 pages 318 CN 121729488 A
[0258] Figure 25 shows the base editing efficiency of GEBx0305-ABE in A-to-G conversion of adenine at five endogenous gene sites in HEK293T cells. The sgRNA sequences used in this experiment contained the GEBx0305-M5 (SEQ ID NO: 471) scaffold and 21nt spacer sequences (SEQ ID NO: 703, 722, 723, 725). GEBx0305-ABE exhibited efficient A-to-G conversion at these sites.
[0259] Figure 26 shows the base editing efficiency of GEBx0308-ABE in A-to-G conversion of adenine at five endogenous gene sites in HEK293T cells. The sgRNA sequences used in this experiment contained the GEBx0308-M4 (SEQ ID NO: 458) scaffold and 21nt spacer sequences (SEQ ID NO: 710, 726, 727, 728, 667). GEBx0308-ABE exhibited efficient A-to-G conversion at these sites.
[0260] Figure 27 shows the A conversion of adenine at five endogenous gene sites by GEBx0328-ABE in HEK293T cells.Base editing efficiency for A-to-G conversion. The sgRNA sequences used in this experiment contained the GEBx0328-M6 (SEQ ID NO: 498) scaffold and 24nt spacer sequences (SEQ ID NO: 790, 797, 801, 809, 815). GEBx0328-ABE exhibited efficient A-to-G conversion at these sites.
[0261] Example 8: Evaluation of the specific knockout effect of GEBx0308 on the RHO-P23H site in HEK293T cells. Design of editor plasmid and target plasmid: An editor plasmid containing SaCas9 (SEQ ID NO: 906) or GEBx0308 encoding gene (Section 313 / 316, CN 121729488 A (SEQ ID NO: 252) and corresponding guide sequences (RHO-P23H sgRNA (SaCas9), i.e., the single guide RNA of SaCas9, sequence SEQ ID NO: 904; and RHO-P23H sgRNA, i.e., the single guide RNA of GEBx0308, sequence SEQ ID NO: 902) was designed and generated. Design and generate a target plasmid containing RHO-WT (approximately 170 bp), RHO-P23H (approximately 170 bp, c.68C>A) and a 500 bp irrelevant sequence (used to separate RHO-WT and RHO-P23H).
[0262] The specific knockout effect of GEBx0308 on the RHO-P23H site in HEK293T cells was evaluated. HEK293T cells were cultured in Dulbecco modified Eagle medium (DMEM, CORNING) containing 2 mM L-glutamine (GlutaMAX™-1, Gibco) and 10% fetal bovine serum (FETAL BOVINE SERUM, GEMINI) at 37°C in a 5% CO2 buffer incubator.
[0263] To evaluate the specific knockout efficiency of SaCas9 (SEQ ID NO: 906, Table 10) and GEBx0308 (SEQ ID NO: 252) at the RHO-P23H site, editor plasmids and target plasmids were co-transfected into HEK293T cells. 1.0 × 10^5 cells were seeded into each well of a 24-well plate. Editor plasmids and target plasmids were co-transfected into HEK293T cells at 2:1 and 3:1 (w / w) ratios using Lipofectamine 3000 reagent, following the manufacturer's instructions (Life Technologies). After 72 hours of incubation, cells were harvested and lysed to release the target plasmid. Target plasmid-specific primers were used.Yes, using potentially editable target plasmids as templates, the editing efficiency of RHO-WT and RHO-P23H sites was analyzed by NGS. The results are shown in Figure 28.
[0264] Table 10: Example sequences associated with RHO-P23H site-specific knockout. Specification 314 / 316 pages 320 CN 121729488 A
[0265] Specification 315 / 316 pages 321 CN 121729488 A
[0266] Although preferred embodiments of the invention have been shown and described herein, these embodiments are provided by way of example only to those skilled in the art. The invention is not intended to be limited to the specific embodiments provided in the specification. Although the invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein should not be construed as limiting. Various changes, modifications and substitutions can now be made by those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the invention are not limited to the specific descriptions, configurations or relative proportions presented herein, which depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in carrying out this invention. Therefore, this invention should also cover any such alternatives, modifications, variations, or equivalents. The scope of this invention is intended to be defined by the following claims and to cover all methods and structures within the scope of these claims and their equivalents. Instruction manual 316 / 316 pages 322 CN 121729488 A Figure 1 Instruction manual Figure 1 / 31 pages 323 CN 121729488 A Figure 2A Instruction manual Figure 2 / 31 pages 324 CN 121729488 A Figure 2B Instruction manual Figure 3 / 31 pages 325 CN 121729488 A Figure 2C Instruction manual Figure 4 / 31 pages 326 CN 121729488 A Figure 2D Instruction manual Figure 5 / 31 pages 327 CN 121729488 A Figure 2E Instruction manual Figure 6 / 31 pages 328 CN 121729488 A Figure 3 Figure 4 Instruction manual Figure 7 / 31 pages 329 CN 121729488 A Figure 5 Instruction manual Figure 8 / 31 pages 330 CN 121729488 A Figure 6 Instruction manual Figure 9 / 31 pages 331 CN Figure 7 of the instruction manual, page 332, 10 / 31; Figure 8 of the instruction manual, page 333, 11 / 31; Figure 9 of the instruction manual, page 333, 12 / 31.Page 334 CN 121729488 A Figure 10 Instruction Manual Drawing 13 / 31 Page 335 CN 121729488 A Figure 11 Instruction Manual Drawing 14 / 31 Page 336 CN 121729488 A Figure 12 Instruction Manual Drawing 15 / 31 Page 337 CN 121729488 A Figure 13 Instruction Manual Drawing 16 / 31 Page 338 CN 121729488 A Figure 14 Instruction Manual Drawing 17 / 31 Page 339 CN 121729488 A Figure 15 Instruction Manual Drawing 18 / 31 Page 340 CN 121729488 A Figure 16 Instruction Manual Drawing 19 / 31 Page 341 CN 121729488 A Figure 17 Instruction Manual Drawing 20 / 31 Page 342 CN 121729488 A Figure 18 Instruction Manual Drawing 21 / 31 Page 343 CN 121729488 A Figure 19 Instruction Manual Illustration 22 / 31 Page 344 CN 121729488 A Figure 20 Instruction Manual Illustration 23 / 31 Page 345 CN 121729488 A Figure 21 Instruction Manual Illustration 24 / 31 Page 346 CN 121729488 A Figure 22 Instruction Manual Illustration 25 / 31 Page 347 CN 121729488 A Figure 23 Instruction Manual Illustration 26 / 31 Page 348 CN 121729488 A Figure 24 Instruction Manual Illustration 27 / 31 Page 349 CN 121729488 A Figure 25 Instruction Manual Illustration 28 / 31 Page 350 CN 121729488 A Figure 26 Instruction Manual Illustration 29 / 31 Page 351 CN 121729488 A Figure 27 Instruction Manual Illustration 30 / 31 Page 352 CN 121729488 A Figure 28 Instruction Manual Drawings 31 / 31 Page 353 CN 121729488 A
Claims
1. An engineered, non-naturally occurring type II CRISPR-associated (Cas) protein or a variant thereof, having at least 70% sequence identity with any amino acid sequence in SEQ ID NO:1-71.
2. The Cas protein according to claim 1, wherein the sequence identity of the Cas protein with any amino acid sequence in SEQ ID NO: 1-71 is at least 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%.
3. An engineered, non-naturally occurring type II CRISPR-associated (Cas) protein, wherein, except for amino acid "M" at position 1 of the sequence, the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with any amino acid sequence in SEQ ID NO: 1-71.
4. The Cas protein according to any one of claims 1-3, wherein the sequence identity of the Cas protein with any amino acid sequence of SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60 and 68 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%.
5. The Cas protein according to any one of claims 1-3, wherein, Except for amino acid "M" at position 1 in the sequence, the Cas protein has sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with any of the amino acid sequences in SEQ ID NO: 9, 12, 19, 21, 22, 24, 25, 27, 29, 30, 31, 36, 37, 38, 43, 44, 51, 56, 59, 60, and 68.
6. The Cas protein according to any one of claims 1-5, wherein the Cas protein is capable of recognizing a group consisting of the following sequence adjacent motifs (PAMs): NRRANH, NRHACT, NRAAR, NNNCCY, NNRYYYY, NGG, NNNCAA, NRNACN, NNGR, NGGNR, NNNCCH, NRRAAG, NRHRAC, NRYART, NRHACC, NRAAR, NRNVHH, YMACAW, NAHAA, NRHAYY, and NGGHA.
7. The Cas protein according to any one of claims 1-6, wherein: (1) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 9, and is able to recognize PAM with the sequence NRRANH; (2) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 12, and is able to recognize PAM with the sequence NRHACT; (3) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 19, and is able to recognize PAM with the sequence NRAAR; (4) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 9, and is able to recognize PAM with the sequence NRAAR; The amino acid sequence identity of 21 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NNNCCY; (5) The amino acid sequence identity of Cas protein with SEQ ID NO: 22 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NNRYYYY; (6) The amino acid sequence identity of Cas protein with SEQ ID NO: 24 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NGG; (7) The amino acid sequence identity of Cas protein with SEQ ID NO: The amino acid sequence identity of 25 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NNNCAA; (8) The amino acid sequence identity of Cas protein with SEQ ID NO: 27 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NRNACN; (9) The amino acid sequence identity of Cas protein with SEQ ID NO: 29 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NNGR; (10) The amino acid sequence identity of Cas protein with SEQ ID NO: The sequence identity of the amino acid sequence of 30 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%, and it is able to recognize PAM with the sequence NGGNR.(11) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 31, and is capable of recognizing PAM with the sequence NNNCCH; (12) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 36, and is capable of recognizing PAM with the sequence NRRAAG; (13) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 37, and is capable of recognizing PAM with the sequence NRHRAC; (14) The Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 37, and is capable of recognizing PAM with the sequence NRHRAC; The amino acid sequence identity of 38 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the NRYART sequence; (15) The amino acid sequence identity of the Cas protein with SEQ ID NO: 43 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the NRHACC sequence; (16) The amino acid sequence identity of the Cas protein with SEQ ID NO: 44 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the NRAAR sequence; (17) The amino acid sequence identity of the Cas protein with SEQ ID NO: The amino acid sequence identity of 51 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NRNVHH; (18) The amino acid sequence identity of Cas protein with SEQ ID NO: 56 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence YMACAW; (19) The amino acid sequence identity of Cas protein with SEQ ID NO: 59 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and it is able to recognize PAM with the sequence NAHAA; (20) The amino acid sequence identity of Cas protein with SEQ ID NO: The amino acid sequence of 60 has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and is able to recognize PAM with the sequence NRHAYY.Or (21) the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100% with the amino acid sequence of SEQ ID NO: 68, and is able to recognize PAM with the sequence NGGHA.
8. The Cas protein according to any one of claims 1-7, wherein the Cas protein is a cleavage enzyme or an inactivated Cas protein.
9. The Cas protein according to claim 8, wherein the Cas protein has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 9, and has a mutation at residue D11 or H859; or has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 12, and has a mutation at residue D10 or H862; or has a sequence identity of at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100% with the amino acid sequence of SEQ ID NO: 31, and has a mutation at residue D12 or H903.
10. The Cas protein according to claim 8, wherein, except for amino acid "M" at position 1 in the sequence, the sequence identity of the Cas protein with the amino acid sequence of SEQ ID NO: 9 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and there is a mutation at residue D11 or H859; or except for amino acid "M" at position 1 in the sequence, the sequence identity with the amino acid sequence of SEQ ID NO: 12 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and there is a mutation at residue D10 or H862; or except for amino acid "M" at position 1 in the sequence, the sequence identity with the amino acid sequence of SEQ ID NO: 31 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99%, or 100%, and there is a mutation at residue D12 or H903.
11. The Cas protein according to any one of claims 8-10, wherein the mutation at residue D11 or H859 of SEQ ID NO: 9 is D11A or H859A; the mutation at residue D10 or H862 of SEQ ID NO: 12 is D10A or H862A; or the mutation at residue D12 or H903 of SEQ ID NO: 31 is D12A or H903A.
12. The Cas protein according to any one of claims 8-11, wherein the sequence identity of the Cas protein with any amino acid sequence in SEQ ID NO: 869-877 is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 98%, 99% or 100%.
13. The Cas protein according to any one of claims 8-12, wherein the Cas protein further comprises a sequence selected from the following sequence group: nuclear localization signal sequence, nuclear exit signal sequence, cell penetration peptide sequence, affinity tag sequence, deaminase sequence, reverse transcriptase sequence, recombinase sequence, methyltransferase sequence, methyltransferase sequence, acetyltransferase sequence, acetyltransferase sequence, transcription activator sequence, transcription repressor domain sequence, cryptochrome sequence, photoinducible / controllable domain sequence, and chemically inducible / controllable domain sequence.
14. The Cas protein according to any one of claims 1-13, wherein the Cas protein comprises an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 92%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the amino acid sequences in SEQ ID NO: 81-151 and 864-866.
15. An engineered, non-naturally occurring polynucleotide encoding a type II CRISPR-associated (Cas) protein as described in any one of claims 1-14.
16. The polynucleotide of claim 15, wherein the polynucleotide is a ribonucleotide sequence or a deoxyribonucleotide sequence, or an analogue thereof; optionally, the polynucleotide is codon-optimized for expression in cells of interest; preferably, the polynucleotide is mRNA and further comprises a 5' cap sequence and / or a poly-A tail sequence.
17. The polynucleotide of claim 16, wherein the polynucleotide is codon-optimized for expression in eukaryotic cells; optionally, the eukaryotic cells are selected from the group consisting of: plant cells, fungal cells, unicellular eukaryotes, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasitic cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells, and / or human cells.
18. The polynucleotide according to any one of claims 15-17, wherein the sequence identity of the polynucleotide with any of the nucleotide sequences in SEQ ID NO: 161-231, 241-311, 851-852 and 861-863 is at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99% or 100%.
19. According to an engineered, non-naturally occurring CRISPR-Cas system, comprising: a) The type II Cas protein or the polynucleotide encoding the Cas protein as described in any one of claims 1-14; b) at least one engineered guide RNA or at least one engineered nucleic acid encoding the guide RNA, wherein the guide RNA comprises a spacer sequence complementary to the target nucleic acid and a Cas protein binding segment that interacts with the Cas protein, wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer.
20. The system of claim 19, wherein the guide RNA further comprises a linker sequence connecting the tracrRNA sequence and the DR sequence to form the sgRNA backbone.
21. The system of claim 20, wherein the sgRNA backbone comprises a sequence that is at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identical to any one of the sequences in SEQ ID NO: 431-562.
22. The system according to any one of claims 19-21, wherein the polynucleotide encoding the Cas protein is operatively linked to a promoter; optionally, the promoter is a constitutive promoter, a tissue-specific promoter, or an inducible promoter.
23. The system of claim 22, wherein the polynucleotide encoding the Cas protein is operatively linked to a promoter and is present in a vector; optionally, the vector is selected from the group consisting of: retroviral vectors, lentiviral vectors, phage vectors, adenovirus vectors, adeno-associated virus vectors, herpes simplex virus vectors, and plasmid vectors.
24. An engineered vector comprising the polynucleotide of any one of claims 15-18, wherein the vector may be an inducible, conditional, or constitutive expression vector.
25. A vector system comprising one or more polynucleotides according to any one of claims 15-18 and one or more polynucleotides encoding a guide RNA; wherein the guide RNA comprises a spacer sequence complementary to a target nucleic acid and a Cas protein binding segment that interacts with the Cas protein, wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer.
26. An engineered, non-naturally occurring cell, comprising: The Cas protein as described in any one of claims 1-14, the polynucleotide as described in any one of claims 15-18, the CRISPR-Cas system as described in any one of claims 19-23, the vector as described in claim 24, and the vector system as described in claim 25.
27. A cell modified with any one of the Cas protein as claimed in claims 1-14, any one of the polynucleotides as claimed in claims 15-18, any one of the CRISPR-Cas systems as claimed in claims 19-23, the vector as claimed in claim 24, or the vector system as claimed in claim 25.
28. A kit comprising: the Cas protein as described in any one of claims 1-14, the polynucleotide as described in any one of claims 15-18, the CRISPR-Cas system as described in any one of claims 19-23, the vector as described in claim 24, the vector system as described in claim 25, or the cell as described in claim 26 or 27.
29. A pharmaceutical composition comprising: the Cas protein as described in any one of claims 1-14, the polynucleotide as described in any one of claims 15-18, the CRISPR-Cas system as described in any one of claims 19-23, the vector as described in claim 24, the vector system as described in claim 25, or the cell as described in claim 26 or 27.
30. The pharmaceutical composition of claim 29, wherein the pharmaceutical composition further comprises a delivery system selected from the following options: AAV (adeno-associated virus), adenovirus, retrovirus, HSV (herpes simplex virus), Gammaretrovirus, LV (lentivirus), eCIS (extracellular contractile injection system), eVLPs (engineered virus-like particles), VLPs (virus-like particles), liposomes, plasmids, LNPs (lipid nanoparticles), exosomes, microvesicles, nucleic acid nanoassemblies, gene guns, and / or implantable devices.
31. A method for treating, preventing, diagnosing, or detecting a disease, comprising a Cas protein as claimed in any one of claims 1-14, a polynucleotide as claimed in any one of claims 15-18, a CRISPR-Cas system as claimed in any one of claims 19-23, a vector as claimed in claim 24, a vector system as claimed in claim 25, a cell as claimed in claim 26 or 27, a kit as claimed in claim 28, or a pharmaceutical composition as claimed in claim 29 or 30.
32. A method for modifying or targeting a target DNA site, comprising providing the site with the Cas protein as described in any one of claims 1-14, the polynucleotide as described in any one of claims 15-18, the CRISPR-Cas system as described in any one of claims 19-23, the vector as described in claim 24, or the vector system as described in claim 25.
33. The method of claim 32, wherein modifying or targeting the target DNA site comprises inducing DNA strand breaks, altering the gene expression of one or more genes, or epigenetically modifying the target DNA site; optionally, the DNA strand breaks comprise DNA double-strand breaks or DNA single-strand breaks.
34. The method according to any one of claims 32 or 33, wherein the method is performed in vitro or in vivo.
35. A method for targeting and cleaving double-stranded target DNA, comprising: The double-stranded target DNA is contacted with the Cas protein as described in any one of claims 1-14, the polynucleotide as described in any one of claims 15-18, the CRISPR-Cas system as described in any one of claims 19-23, or the pharmaceutical composition as described in claim 29 or 30.
36. An isolated eukaryotic cell comprising a modified target site, wherein the target site has been modified by the method of any one of claims 32-35, or using a pharmaceutical composition as described in claims 29 or 30, or using a CRISPR-Cas system as described in any one of claims 19-23.
37. A system for detecting the presence of a target nucleic acid sequence in an in vitro sample, comprising: a) The Cas protein as described in any one of claims 1-14; b) At least one guide polynucleotide containing a guide sequence capable of binding to the target sequence, designed to form a complex with the Cas protein; c) A nucleic acid-based masking structure containing a non-target sequence; The Cas protein exhibits incidental cleavage activity on RNA and / or ssDNA, and cleaves non-target sequences in nucleic acid-based masking structures activated by the target sequence.
38. A method for detecting a target nucleic acid in a sample, comprising: The sample is subjected to one or more samples with a) the Cas protein as described in any one of claims 1-14; b) At least one guide polynucleotide containing a guide sequence designed to be complementary to the target sequence to form a complex with the Cas protein; c) A nucleic acid-based masking structure contact containing a non-target sequence; The Cas protein exhibits incidental cleavage activity of RNA and / or ssDNA, and cleaves non-target sequences in nucleic acid-based masking structures activated by the target sequence; and detects signals from non-target sequence cleavage, thereby detecting one or more target sequences in the sample.
39. A guide RNA (gRNA) comprising: a) a spacer sequence of SEQ ID NO: 901; a spacer sequence having at least 15, 16, 17, 18, 19 or 20 consecutive nucleotides of the sequence SEQ ID NO: 901; or a spacer sequence having at least 99%, at least 98%, at least 97%, at least 96%, at least 95%, at least 94%, at least 93%, at least 92%, at least 91% or at least 90% sequence identity with the sequence SEQ ID NO: 901; and b) a Cas protein binding segment; wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence, the DR sequence hybridizing to form a double-stranded RNA (dsRNA) duplex.
40. The gRNA according to claim 39, wherein the gRNA is a double-stranded gRNA or a single-stranded gRNA.
41. The gRNA of claim 40, wherein the gRNA has at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 98%, at least 99%, or 100% sequence identity with the nucleotide sequence of SEQ ID NO:
903.
42. The gRNA of claim 41, wherein at least three nucleotides of the gRNA are modified.
43. A polynucleotide encoding the gRNA of any one of claims 39-42.
44. An engineered, non-naturally occurring CRISPR-Cas system, comprising: a) a Cas protein or a polynucleotide encoding the Cas protein; b) at least one guide RNA (gRNA) or at least one engineered nucleic acid encoding the guide RNA as described in any one of claims 39-42, wherein the gRNA further comprises a Cas protein binding segment that interacts with the Cas protein; wherein the Cas protein binding segment comprises a tracrRNA sequence and a direct repeat (DR) sequence that hybridize to form a double-stranded RNA (dsRNA) dimer.
45. The CRISPR-Cas system of claim 44, wherein the gRNA comprises: a) the sequence of SEQ ID NO: 903; or b) a sequence having at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identity with the sequence of SEQ ID NO:
903.
46. The CRISPR-Cas system according to claim 44 or 45, wherein the Cas protein comprises a sequence having at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identity with SEQ ID NO: 12; or except for the amino acid "M" at position 1 in the sequence, the Cas protein comprises a sequence having at least 70%, 75%, 80%, 85%, 88%, 90%, 92%, 94%, 95%, 96%, 98%, 99%, or 100% identity with SEQ ID NO:
12.
47. An engineered vector comprising the polynucleotide of claim 43, wherein the vector is optionally an inducible, conditional, or constitutive expression vector.
48. A carrier system comprising one or more polynucleotides as described in claim 43.
49. A pharmaceutical composition comprising gRNA as claimed in any one of claims 39-42, a polynucleotide as claimed in claim 43, a CRISPR-Cas system as claimed in any one of claims 44-46, a vector as claimed in claim 47, or a vector system as claimed in claim 48.
50. A method for treating, preventing, or diagnosing diseases associated with the RHO locus, comprising contacting target cells of a subject requiring treatment with a gRNA as described in any one of claims 39-42, a polynucleotide as described in claim 43, a CRISPR-Cas system as described in any one of claims 44-46, or a pharmaceutical composition as described in claim 49.