Cas proteins, crisper-cas systems, and applications thereof

CN122270552APending Publication Date: 2026-06-23GUANGZHOU REFORGENE MEDICINE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU REFORGENE MEDICINE CO LTD
Filing Date
2024-11-14
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

The existing CRISPR-Cas13 system is inefficient in editing and has high cytotoxicity in mammalian cells, making it difficult to achieve RNA targeting/cleaving activity and cell safety.

Method used

A novel Cas protein with high amino acid sequence identity and known Cas13 protein sequence is developed, which can form a CRISPR complex with guide polynucleotides, and achieve sequence-specific binding and cleavage of targeted RNA.

Benefits of technology

Improves the efficiency and safety of RNA editing in mammalian cells, reduces cytotoxicity, and achieves a compact Cas13 system suitable for AAV delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000119_0000
    Figure 00000119_0000
  • Figure 00000119_0001
    Figure 00000119_0001
  • Figure 00000119_0002
    Figure 00000119_0002
Patent Text Reader

Abstract

A CRISPR-Cas system and applications thereof are provided, also related to a Cas protein, a fusion protein and a guide polynucleotide. The Cas protein has at least 50% sequence identity compared to the sequence set forth in any one of SEQ ID NOs: 1-184, 426-428. The fusion protein comprises the Cas protein fused to a protein domain and / or a polypeptide tag. The guide polynucleotide comprises a direct repeat sequence having at least 70% sequence identity to any one of SEQ ID NOs: 185-199, 245-247, 287-303, 369 and 370-373, and a guide sequence engineered to hybridize to a target nucleic acid. The CRISPR-Cas system comprises a Cas protein having at least 90% sequence identity compared to the sequence set forth in any one of SEQ ID NOs: 1-184, 426-428, or a nucleic acid encoding thereof, and a guide polynucleotide or a nucleic acid encoding thereof.
Need to check novelty before this filing date? Find Prior Art

Description

Cas proteins, CRISPR-Cas systems and their applications

[0001] This application claims priority to Chinese Patent Application No. 2023115084846 filed on November 14, 2023, and Chinese Patent Application No. 2024106414288 filed on May 22, 2024. This application incorporates the entirety of the aforementioned Chinese patent applications. Technical Field

[0002] The present disclosure relates to the field of CRISPR gene editing, and specifically to a Cas protein, a CRISPR-Cas system, and applications thereof. Background Art

[0003] The CRISPR-Cas system is a nucleic acid targeting and editing system based on the bacterial immune system that protects bacteria from viruses. CRISPR-Cas systems include CRISPR-Cas9, CRISPR-Cas12, and CRISPR-Cas13. The CRISPR-Cas13 system is similar to the CRISPR-Cas9 system, but unlike the Cas9 protein, which targets DNA, the Cas13 protein targets RNA.

[0004] CRISPR-Cas13 belongs to the Type VI CRISPR-Cas13 system, which contains a single Cas effector protein. Currently, CRISPR-Cas13 can be divided into multiple subtypes (such as Cas13a, Cas13b, Cas13c and Cas13d) based on phylogeny. However, there is still an urgent need to discover new Cas13 systems with compact size (for example, suitable for AAV delivery), high editing efficiency in mammalian cells (for example, RNA targeting / cleavage activity) and / or low cytotoxicity (for example, cell dormancy and apoptosis caused by bystander RNA degradation).

[0005] Summary of the Invention

[0006] The present disclosure provides a Cas protein, a CRISPR-Cas system, and applications thereof.

[0007] One aspect of the present disclosure relates to a Cas protein having an amino acid sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 1-184, 426-428.

[0008] In some embodiments, the amino acid sequence of the Cas protein has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to the sequence shown in any one of SEQ ID NOs: 2, 103-105, 61-62, 106-107.

[0009] In some embodiments, the amino acid sequence of the Cas protein has 100% sequence identity to the sequence shown in any one of SEQ ID NOs: 1-184, 426-428.

[0010] In some embodiments, the amino acid sequence of the Cas protein comprises the sequence shown in any one of SEQ ID NOs: 1-184, 426-428. In some embodiments, the amino acid sequence of the Cas protein is as shown in any one of SEQ ID NOs: 1-184, 426-428.

[0011] The Cas proteins disclosed herein with sequences of SEQ ID NOs: 1-184, 426-428 were identified based on bioinformatics analysis of genomes in databases such as NCBI Genbank and subsequent activity verification.

[0012] In some embodiments, the Cas protein is capable of forming a CRISPR complex with a guide polynucleotide comprising a direct repeat sequence linked to a guide sequence engineered to direct sequence-specific binding of the CRISPR complex to a target nucleic acid.

[0013] Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the sequence shown in any one of SEQ ID NOs: 185-368, 429-431, or the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the reverse complement of any one of SEQ ID NOs: 185-368, 429-431.

[0014] In some embodiments, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 185-368, 429-431.

[0015] In some embodiments, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to any one of SEQ ID NOs: 186, 287-291, 245-246.

[0016] In some embodiments, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the reverse complement of any one of SEQ ID NOs: 185-368, 429-431.

[0017] In some embodiments, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the reverse complement of any one of SEQ ID NOs: 186, 287-291, 245-246.

[0018] In some embodiments, the direct repeat sequence has 100% sequence identity to any one of SEQ ID NOs: 185-368, 429-431. In some embodiments, the direct repeat sequence has 100% sequence identity to the reverse complement of any one of SEQ ID NOs: 185-368, 429-431.

[0019] In some embodiments, the direct repeat sequence has 100% sequence identity to any one of SEQ ID NOs: 186, 287-291, 245-246. In some embodiments, the direct repeat sequence has 100% sequence identity to the reverse complement of any one of SEQ ID NOs: 186, 287-291, 245-246.

[0020] In some embodiments, the direct repeat sequence comprises or is a sequence shown in any one of SEQ ID NOs: 185-368, 429-431. In some embodiments, the direct repeat sequence comprises or is a reverse complementary sequence of a sequence shown in any one of SEQ ID NOs: 185-368, 429-431.

[0021] In some embodiments, the guide sequence is located 3' to the direct repeat sequence.

[0022] In some embodiments, the guide sequence is located 5' to the direct repeat sequence.

[0023] In some embodiments, the Cas protein can be guided to the target nucleic acid by a guide polynucleotide; or

[0024] The Cas protein can be guided to a target nucleic acid by a guide polynucleotide and target or modify the target nucleic acid.

[0025] In some embodiments, the target nucleic acid is RNA or DNA.In some embodiments, the target nucleic acid is PTBP1 mRNA, AQpl mRNA, VEGFA mRNA, VEGFR1 mRNA, VEGFR2 mRNA, or AR mRNA.

[0026] In some embodiments, the Cas protein is derived from: the same kingdom, phylum, class, order, family, genus or species as the Cas protein having an amino acid sequence comprising the sequence shown in any one of SEQ ID NOs: 1-184, 426-428; and / or

[0027] The Cas protein is a CRISPR-Cas effector protein and / or a CRISPR-Cas auxiliary protein.

[0028] Another aspect of the present disclosure relates to a Cas protein, wherein the amino acid sequence of the Cas protein comprises the amino acid sequence shown in any one of the following motifs 1-9:

[0029] Motif 1: R-[NH]-x(2)-[FNI]-H,

[0030] Motif 2: Px(4)-[VFL]-x(2)-[RK],

[0031] Motif 3: [LIM]-[KR]-x(2)-Yx(3)-[FI],

[0032] Motif 4: [RK]-x-[LVI]-x(6)-[GN]-x(4)-[LIV],

[0033] Motif 5: [RL]-[ARK]-[NH]-x(2)-[RTE]-[LY],

[0034] Motif 6: [LIK]-[MLI]-x(2)-[[VLI]-x-[GSA]-[RKQ]-[LW],

[0035] Motif 7: [LIR]-[WA]-ERD-[LRI]-x-[YF]-x(2)-[LK],

[0036] Motif 8: [LI]-[RV]-[NGR]-x-[FIL]-x-[HS]-[FMC],

[0037] Motif 9: [YHA]-[DR]-x-[KLN]-x(4)-V; or,

[0038] The amino acid sequence of the Cas protein comprises the amino acid sequence shown in any one of the following motifs 10-13:

[0039] Motif 10: [KSQ]-[LMI]-x(3)-RN-[YTF]-X-[SAT]-H,

[0040] Motif 11: G-[LRA]-x-[YF]-[FL]-x-[SC]-[FL]-[FA]-L,

[0041] Motif 12: [KQCG]-[NTQ]-xD-[IRS]-[VLI]-[LT],

[0042] Motif 13: RNx(3)-H-[NMR]; or,

[0043] The amino acid sequence of the Cas protein comprises the amino acid sequence shown in any one of the following motifs 14-25:

[0044] Motif 14: Rx(3)-[VA]-H,

[0045] Motif 15: [KLN]-[NSMY]-x-[GP]-[FMYLV]-[STDN]-x-[KTER]-x-[LYIV]-[RF]-E,

[0046] Motif 16: [DSN]-[STPKF]-x-[REK]-x-[KIR]-[LIMAVF],

[0047] Motif 17: Kx(3)-Yx(2)-[EA]-Ax(2)-L,

[0048] Motif 18:

[0049] [FL]-[LMIQ]-[DYEN]-[GIDAS]-[KQ]- -[IVT]-[NS]-x-[LKFM]-x(4)-[ISTAM],

[0050] Motif 19: [LIF]-x(4)-[SQGN]-[FVI]-x-[RK]-[MISF],

[0051] Motif 20: [AS]-x(3)-[LAIF]-[GDN],

[0052] Motif 21: [NS]-N-[VI]-[IVL],

[0053] Motif 22: [NCS]-x(3)-[VIR]-x-[FSLY]-[VFAP]-[LGITF],

[0054] Motif 23:

[0055] K-[NSTQ]-[LMG]-[VMI]-x-[VWIT]-[NL]-[SAT],

[0056] Motif 24: R-[NR]-x(3)-H,

[0057] Motif 25: Yx(3)-[RS]-[YFK]-x-[NLS]-[LT]-[SAT];

[0058] Among them, A, F, C, U, D, N, E, Q, G, H, L, I, K, O, M, P, R, S, T, V, W, and Y are standard amino acid codes, "x" is any amino acid, the number in the brackets after x represents multiple consecutive xs, "[]" is an optional amino acid code, and "-" is a separator.

[0059] In some embodiments, the amino acid sequence of the Cas protein comprises the amino acid sequence shown in the following motifs 26-34 from N-terminus to C-terminus:

[0060] Motif 26: RH-[ARY]-[STL]-[FNI]-H,

[0061] Motif 27: P-[RKS]-[LF]-[NMS]-[RSK]-[VF]-[IL]-[NRT]-[RK]-[AL]-[RK],

[0062] Motif 28: [LI]-[KR]-[MHN]-[LI]-Y-[EAK]-[QTI]-[GD]-[FI],

[0063] Motif 29: R-[SG]-[LI]-[RL]-[EQ]-x(2)-[RH]-x-[GN]-x(2)-[DS]-x-[LI],

[0064] Motif 30: [RL-[AR]-[NH]-x(2)-[RT]-[LY],

[0065] Motif 31: [LK]-[ML]-[MI]-xVx-[GSA]-[RK]-[LW],

[0066] Motif 32: [LIR]-[WA]-ERD-[LR]-[YNH]-[YF]-[VL]-[TIL]-[LK],

[0067] Motif 33: [LI]-RNx-[FI]-[SA]-H-[FM]-[NY],

[0068] Motif 34: Y-[DR]-[RS]-[KL]-[LY]-[KN]-[NS]-[SA]-V; or,

[0069] The amino acid sequence of the Cas protein comprises the amino acid sequence shown in the following motifs 35-46 from N-terminus to C-terminus:

[0070] Motif 35: [LA]-R-[NQ]-x(2)-[VA]-Hx(2)-E,

[0071] Motif 36: KNxGF-[ST]-x-[KT]-xLRE,

[0072] Motif 37: D-[ST]-[IV]-RxKL,

[0073] Motif 38: Kx-[KL]-[LI]-Yx-[DH]-[EA]-Ax(2)-L-[WY],

[0074] Motif 39:

[0075] GKEIN-[ED]-LLTT-[LC]-INKF-[DE]-NI,

[0076] Motif 40: [EQ]-Lx-[LE]-x-[KN]-SF-[VA]-[KR]-[SM],

[0077] Motif 41: Ax(2)-[IL]-LG,

[0078] Motif 42: [NS]-NV-[VL],

[0079] Motif 43: N-[RE]-A-[VI]-[VI]-[RD]-FVL,

[0080] Motif 44:

[0081] KN-[LM]-V-[NY]-[VI]-[N]-[SA],

[0082] Motif 45: R-[N]-x(3)-H,

[0083] Motif 46: Y-[NC]-x(2)-R-[YF]-KNL-[ST]-Ix(2)-LFD;

[0084] Among them, A, F, C, U, D, N, E, Q, G, H, L, I, K, O, M, P, R, S, T, V, W, and Y are standard amino acid codes, "x" is any amino acid, the number in the brackets after x represents multiple consecutive xs, "[]" is an optional amino acid code, and "-" is a separator.

[0085] In some embodiments, the Cas protein comprises a sequence as shown in any one of SEQ ID NOs: 1-15, SEQ ID NOs: 61-63, SEQ ID NOs: 103-119, or a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a sequence shown in any one of SEQ ID NOs: 1-15, SEQ ID NOs: 61-63, SEQ ID NOs: 103-119.

[0086] In some embodiments, any amino acid residue in the amino acid sequence of the Cas protein, except for the amino acids determined by motifs 1-25, undergoes conservative amino acid substitutions based on the wild-type sequence, and the wild-type sequence includes the sequences shown in SEQ ID NO: 1-15, SEQ ID NO: 61-63, and SEQ ID NO: 103-119.

[0087] In some embodiments, the Cas protein is capable of forming a CRISPR complex with a guide polynucleotide comprising a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to direct sequence-specific binding of the CRISPR complex to a target nucleic acid;

[0088] Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the sequence shown in any one of SEQ ID NOs: 185-368, 429-431, or the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the reverse complement of any one of SEQ ID NOs: 185-368, 429-431.

[0089] In some embodiments, the Cas protein can be guided to the target nucleic acid by a guide polynucleotide; or

[0090] The Cas protein can be guided to a target nucleic acid by a guide polynucleotide and target or modify the target nucleic acid.

[0091] In some embodiments, the target nucleic acid is PTBP1 mRNA, AQp1 mRNA, VEGFA mRNA, VEGFR1 mRNA, VEGFR2 mRNA, or AR mRNA.

[0092] In some embodiments, the Cas protein is derived from: the same kingdom, phylum, class, order, family, genus or species as the Cas protein comprising the sequence shown in any one of SEQ ID NOs: 1-15, SEQ ID NOs: 61-63, and SEQ ID NOs: 103-119; and / or

[0093] The Cas protein is a CRISPR-Cas effector protein and / or a CRISPR-Cas auxiliary protein.

[0094] Another aspect of the present disclosure relates to a fusion protein comprising: the Cas protein or a functional fragment thereof described in the present disclosure.

[0095] In some embodiments of the present disclosure, the Cas protein or its functional fragment is fused to a protein domain and / or a polypeptide tag, and the fusion does not change the original function of the Cas protein or its functional fragment.

[0096] In some embodiments of the present disclosure, the Cas protein or a functional fragment thereof is fused to a localization tag that provides subcellular localization;

[0097] Optionally, the localization tag is selected from a nuclear localization signal (NLS) or a nuclear export signal (NES).

[0098] In some embodiments of the present disclosure, the Cas protein or a functional fragment thereof is fused to any one or more protein domains and / or polypeptide tags selected from the group consisting of a cytosine deaminase domain, an adenosine deaminase domain, a translation activation domain, a translation repression domain, an RNA methylation domain, an RNA demethylation domain, a nuclease domain, a splicing factor domain, a reporter domain, an affinity domain, a subcellular localization signal, a reporter tag, and an affinity tag.

[0099] In some embodiments of the present disclosure, the Cas protein or a functional fragment thereof is covalently linked to the Cas protein domain; and / or

[0100] The Cas protein or its functional fragment is covalently or non-covalently linked to a conjugate molecule, the conjugate molecule does not change the original function of the Cas protein or its functional fragment, and the conjugate molecule is selected from at least one of a polymer, a small molecule compound, an antibody, an oligopeptide, a detectable label, and a polypeptide other than the Cas protein.

[0101] In some embodiments of the present disclosure, the structure of the fusion protein is NLS-Cas protein-SV40NLS-nucleoplasminNLS.

[0102] Another aspect of the present disclosure relates to a conjugate, which comprises: the Cas protein or a functional fragment thereof described in the present disclosure.

[0103] In some embodiments, the conjugate comprises: a Cas protein or a functional fragment thereof as described herein, and a conjugated moiety.

[0104] In some embodiments, the conjugate comprises: a Cas protein or a functional fragment thereof as described herein, a linker, and a conjugated moiety.

[0105] In some embodiments, the conjugated moiety is not an amino acid sequence. In some embodiments, the conjugated moiety is an amino acid sequence.

[0106] In some embodiments, the linker is not an amino acid sequence. In some embodiments, the linker is a small molecule structure. In some embodiments, the linker is a small molecule chemical group.

[0107] Another aspect of the present disclosure relates to a guide polynucleotide comprising: (i) a direct repeat sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to any one of SEQ ID NOs: 185-368, 429-431, or the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to any one of SEQ ID NOs: 185-368, 429-431; The reverse complement sequence of any one of NO:185-368, 429-431 has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity, and the direct repeat sequence is linked to (ii) a guide sequence engineered to hybridize to a target nucleic acid; the guide polynucleotide is capable of forming a CRISPR complex with a Cas protein and guiding sequence-specific binding of the CRISPR complex to a target nucleic acid, the Cas protein being a CRISPR-Cas effector protein and / or a CRISPR-Cas accessory protein.

[0108] In some embodiments, the Cas protein has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity compared to the amino acid sequence shown in SEQ ID NOs: 1-184, 426-428; or

[0109] The amino acid sequence of the Cas protein comprises, from N-terminus to C-terminus, the amino acid sequence shown in motifs 1-9, the amino acid sequence shown in motifs 10-13, or the amino acid sequence shown in motifs 14-25.

[0110] In some embodiments, the direct repeat has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity compared to any one of SEQ ID NOs: 185-199, 245-247, 287-303; or

[0111] The direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity compared to the reverse complement of any one of SEQ ID NOs: 185-199, 245-247, 287-303.

[0112] In some embodiments, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to any one of SEQ ID NOs: 186, 287-291, 245-246.

[0113] In some embodiments, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the reverse complement of any one of SEQ ID NOs: 186, 287-291, 245-246.

[0114] In some embodiments, the guide sequence is located 3' to the direct repeat sequence; or

[0115] The guide sequence is located at the 5' end of the direct repeat sequence.

[0116] In some embodiments, the guide sequence comprises 15-35 nucleotides; and / or

[0117] The guide sequence hybridizes to the target nucleic acid with no more than a one nucleotide mismatch.

[0118] In some embodiments, the direct repeat sequence comprises 25 to 40 nucleotides.

[0119] In some embodiments, the guide polynucleotide further comprises an aptamer sequence.

[0120] In some embodiments, the aptamer sequence is inserted into a loop of the guide polynucleotide; and / or

[0121] The aptamer sequence includes an MS2 aptamer sequence, a PP7 aptamer sequence or a Qβ aptamer sequence.

[0122] In some embodiments, the guide polynucleotide comprises modified nucleotides.

[0123] In some embodiments, the modification comprises a 2'-O-methyl, 2'-O-methyl-3'-phosphorothioate, or 2'-O-methyl-3'-thioPACE modification.

[0124] In some embodiments, the target nucleic acid is located in the nucleus of a eukaryotic cell.

[0125] In some embodiments, the target nucleic acid is selected from TTR RNA, SOD1 RNA, PCSK9 RNA, VEGFA RNA, VEGFR1 RNA, PTBP1 RNA, AQp1 RNA, ANGPTL3 RNA, or AR RNA.

[0126] In some embodiments, the guide sequence is selected from the sequence shown in SEQ ID NO: 380-381, SEQ ID NO: 398, or SEQ ID NO: 406.

[0127] Another aspect of the present disclosure relates to a CRISPR-Cas system comprising:

[0128] The Cas protein or fusion protein, recombinant protein or conjugate described herein, or a nucleic acid encoding the Cas protein or fusion protein, recombinant protein or conjugate; and

[0129] Any guide polynucleotide, or a nucleic acid encoding the same, comprising a direct repeat sequence linked to a guide sequence engineered to hybridize to a target nucleic acid;

[0130] The guide polynucleotide is capable of forming a CRISPR complex with the Cas protein or fusion protein and guiding the sequence-specific binding of the CRISPR complex to the target nucleic acid.

[0131] In some embodiments, the direct repeat sequence has at least 70% sequence identity compared to any one of SEQ ID NOs: 185-368, 429-431.

[0132] In some embodiments, the CRISPR-Cas system is a CRISPR-Cas13 system.

[0133] In some embodiments, the target nucleic acid is selected from TTR RNA, SOD1 RNA, PCSK9 RNA, VEGFA RNA, VEGFR1 RNA, PTBP1 RNA, AQp1 RNA, ANGPTL3 RNA, or AR RNA;

[0134] Optionally, the guide sequence is selected from the sequences shown in SEQ ID NO: 380-381, SEQ ID NO: 398 or SEQ ID NO: 406.

[0135] Another aspect of the present disclosure relates to a CRISPR-Cas system comprising:

[0136] The guidance polynucleotides described herein, or nucleic acids encoding the same, and

[0137] Any CRISPR-Cas effector protein and / or CRISPR-Cas accessory protein, fusion protein thereof, recombinant protein thereof, conjugate thereof, or nucleic acid encoding the same.

[0138] Another aspect of the present disclosure relates to a vector system comprising the CRISPR-Cas system described herein, wherein the vector system comprises one or more vectors comprising a polynucleotide sequence encoding the Cas protein, the fusion protein, a recombinant protein thereof, a conjugate thereof, and a polynucleotide sequence encoding a guide polynucleotide.

[0139] Another aspect of the present disclosure relates to an adeno-associated virus (AAV) vector comprising the CRISPR-Cas system described herein, wherein the AAV vector comprises a DNA sequence encoding the Cas13 protein, the fusion protein, the recombinant protein or the conjugate described herein and a DNA sequence of the guide polynucleotide.

[0140] Another aspect of the disclosure relates to a lentiviral vector comprising a CRISPR-Cas system as described herein, the lentiviral vector comprising a guide polynucleotide as described herein and an mRNA encoding a Cas protein or fusion protein, recombinant protein or conjugate as described herein;

[0141] Optionally, the lentiviral vector is pseudotyped with an envelope protein;

[0142] Optionally, the mRNA encoding the Cas protein or fusion protein is linked to an aptamer sequence.

[0143] In some embodiments, the lentiviral vector is pseudotyped with a homologous or heterologous envelope protein, such as VSV-G. In some embodiments, the mRNA encoding the Cas protein or fusion protein is linked to an aptamer sequence.

[0144] Another aspect of the present disclosure relates to a ribonucleoprotein complex comprising a CRISPR-Cas system described herein, wherein the ribonucleoprotein complex is formed by a guide polynucleotide described herein and a Cas protein or fusion protein described herein.

[0145] Another aspect of the present disclosure relates to virus-like particles comprising a CRISPR-Cas system as described herein, comprising a ribonucleoprotein complex formed by a guide polynucleotide as described herein and a Cas protein or fusion protein as described herein. In some embodiments, the Cas protein or fusion protein is fused to a gag protein.

[0146] Another aspect of the present disclosure relates to an isolated nucleic acid encoding a Cas protein described herein or a fusion protein, recombinant protein or conjugate described herein.

[0147] Another aspect of the disclosure relates to an isolated nucleic acid encoding a guide polynucleotide described herein.

[0148] Another aspect of the present disclosure relates to a delivery composition comprising: a delivery system loaded with a Cas protein described herein, a fusion protein, a recombinant protein or conjugate described herein, a guide polynucleotide described herein, a CRISPR-Cas system described herein, a vector system described herein, an adeno-associated viral vector described herein, a lentiviral vector described herein, a ribonucleoprotein complex described herein, a virus-like particle described herein, or a nucleic acid described herein.

[0149] In some embodiments, the delivery system is selected from at least one of lipid nanoparticles and extracellular vesicles.

[0150] Another aspect of the disclosure relates to a eukaryotic cell comprising a Cas protein described herein, a fusion protein, a recombinant protein or a conjugate described herein, a guide polynucleotide described herein, or a CRISPR-Cas system described herein;

[0151] Optionally, the eukaryotic cell is a mammalian cell.

[0152] In some embodiments, the eukaryotic cell is a human cell.

[0153] Another aspect of the present disclosure relates to a pharmaceutical composition comprising a CRISPR-Cas system described herein, a Cas protein described herein, a fusion protein, a recombinant protein or conjugate described herein, a guide polynucleotide described herein, or a CRISPR-Cas system described herein, a vector system described herein, an adeno-associated viral vector described herein, a lentiviral vector described herein, a ribonucleoprotein complex described herein, a virus-like particle described herein, a nucleic acid described herein, a delivery composition described herein, or a eukaryotic cell described herein.

[0154] Another aspect of the present disclosure relates to a pharmaceutical composition comprising a CRISPR-Cas13 system.

[0155] Another aspect of the present disclosure relates to an in vitro composition comprising a CRISPR-Cas system described herein, and a labeled detector RNA that is incapable of hybridizing to or being targeted by a guide polynucleotide described herein.

[0156] Another aspect of the present disclosure relates to the use of a Cas protein described herein, a fusion protein, a recombinant protein or conjugate described herein, a guide polynucleotide described herein, a CRISPR-Cas system described herein, or an isolated nucleic acid described herein in detecting a target nucleic acid in a nucleic acid sample suspected of containing the target nucleic acid or in preparing a reagent for detecting a target nucleic acid in a nucleic acid sample suspected of containing the target nucleic acid.

[0157] Another aspect of the present disclosure relates to the use of the CRISPR-Cas system described herein for detecting a target nucleic acid in a nucleic acid sample suspected of containing the target nucleic acid or for preparing a reagent for detecting a target nucleic acid in a nucleic acid sample suspected of containing the target nucleic acid.

[0158] Another aspect of the present disclosure relates to the use of a composition comprising a CRISPR-Cas system described herein, a Cas protein described herein, a fusion protein, a recombinant protein or conjugate described herein, a guide polynucleotide described herein, an isolated nucleic acid described herein, a vector system described herein, a delivery composition described herein, a lentiviral vector described herein, a ribonucleoprotein complex described herein, a virus-like particle described herein, or a eukaryotic cell described herein in any of the following or in the preparation of an agent for achieving any of the following:

[0159] Cleavage or nicking of one or more target nucleic acid molecules, activation or upregulation of one or more target nucleic acid molecules, activation or inhibition of translation of one or more target nucleic acid molecules, inactivation of one or more target nucleic acid molecules, visualization, labeling or detection of one or more target nucleic acid molecules, binding of one or more target nucleic acid molecules, transport of one or more target nucleic acid molecules, and sequestration of one or more target nucleic acid molecules.

[0160] Another aspect of the present disclosure relates to the use of a composition comprising a CRISPR-Cas system described herein, a Cas protein described herein, a fusion protein, a recombinant protein or conjugate described herein, a guide polynucleotide described herein, a nucleic acid described herein, a vector system described herein, a delivery composition described herein, a lentiviral vector described herein, a ribonucleoprotein complex described herein, a virus-like particle described herein, or a eukaryotic cell described herein for cleaving one or more target nucleic acid molecules or preparing an agent for cleaving one or more target nucleic acid molecules.

[0161] Another aspect of the present disclosure relates to the use of a CRISPR-Cas system described herein, a Cas protein described herein, a fusion protein, a recombinant protein or conjugate described herein, a guide polynucleotide described herein, a nucleic acid described herein, a vector system described herein, a delivery composition described herein, a lentiviral vector described herein, a ribonucleoprotein complex described herein, a virus-like particle described herein, or a eukaryotic cell described herein in binding one or more target nucleic acid molecules.

[0162] Another aspect of the present disclosure relates to the use of the CRISPR-Cas system described herein in the preparation of a reagent that binds to or cleaves one or more target nucleic acid molecules.

[0163] Another aspect of the present disclosure relates to the use of a composition comprising a CRISPR-Cas system as described herein, a Cas protein as described herein, a fusion protein, a recombinant protein or conjugate as described herein, a guide polynucleotide as described herein, a nucleic acid as described herein, a vector system as described herein, a delivery composition as described herein, a lentiviral vector as described herein, a ribonucleoprotein complex as described herein, a virus-like particle as described herein, or a eukaryotic cell as described herein for cutting or editing a target nucleic acid in a mammalian cell, wherein the editing is base editing.

[0164] Another aspect of the present disclosure relates to the use of the CRISPR-Cas system described herein in preparing a reagent for cutting or editing a target nucleic acid in a mammalian cell, wherein the editing is base editing.

[0165] Another aspect of the present disclosure relates to the use of a CRISPR-Cas system described herein, a Cas protein described herein, a fusion protein described herein, a guide polynucleotide described herein, a nucleic acid described herein, a vector system described herein, a delivery composition described herein, a lentiviral vector described herein, a ribonucleoprotein complex described herein, a virus-like particle described herein, or a eukaryotic cell described herein in activating or upregulating one or more target nucleic acid molecules or in the preparation of an agent that activates or upregulates one or more target nucleic acid molecules.

[0166] Another aspect of the present disclosure relates to the use of a composition comprising a CRISPR-Cas system described herein, a Cas protein described herein, a fusion protein, a recombinant protein or conjugate described herein, a guide polynucleotide described herein, a nucleic acid described herein, a vector system described herein, a delivery composition described herein, a lentiviral vector described herein, a ribonucleoprotein complex described herein, a virus-like particle described herein, or a eukaryotic cell described herein in inhibiting the translation of one or more target nucleic acid molecules or in the preparation of an agent that inhibits the translation of one or more target nucleic acid molecules.

[0167] Another aspect of the present disclosure relates to the use of a CRISPR-Cas system described herein, a Cas protein described herein, a fusion protein described herein, a guide polynucleotide described herein, a nucleic acid described herein, a vector system described herein, a delivery composition described herein, a lentiviral vector described herein, a ribonucleoprotein complex described herein, a virus-like particle described herein, or a eukaryotic cell described herein in inactivating one or more target nucleic acid molecules or in preparing an agent that inactivates one or more target nucleic acid molecules.

[0168] Another aspect of the present disclosure relates to the use of a composition comprising a CRISPR-Cas system described herein, a Cas protein described herein, a fusion protein, a recombinant protein or conjugate described herein, a guide polynucleotide described herein, a nucleic acid described herein, a vector system described herein, a delivery composition described herein, a lentiviral vector described herein, a ribonucleoprotein complex described herein, a virus-like particle described herein, or a eukaryotic cell described herein for visualizing, labeling or detecting one or more target nucleic acid molecules or for preparing a reagent for visualizing, labeling or detecting one or more target nucleic acid molecules.

[0169] Another aspect of the present disclosure relates to the use of a composition comprising a CRISPR-Cas system as described herein, a Cas protein as described herein, a fusion protein, a recombinant protein or conjugate as described herein, a guide polynucleotide as described herein, a nucleic acid as described herein, a vector system as described herein, a lipid nanoparticle as described herein, a lentiviral vector as described herein, a ribonucleoprotein complex as described herein, a virus-like particle as described herein, or a eukaryotic cell as described herein for transporting one or more target nucleic acid molecules or for preparing an agent for transporting one or more target nucleic acid molecules.

[0170] Another aspect of the present disclosure relates to the use of a composition comprising a CRISPR-Cas system as described herein, a Cas protein as described herein, a fusion protein, a recombinant protein or conjugate as described herein, a guide polynucleotide as described herein, a nucleic acid as described herein, a vector system as described herein, a lipid nanoparticle as described herein, a lentiviral vector as described herein, a ribonucleoprotein complex as described herein, a virus-like particle as described herein, or a eukaryotic cell as described herein for masking one or more target nucleic acid molecules or for preparing an agent for masking one or more target nucleic acid molecules.

[0171] Another aspect of the present disclosure relates to the use of a CRISPR-Cas system described herein, a Cas protein described herein, a fusion protein, a recombinant protein or conjugate described herein, a guide polynucleotide described herein, a nucleic acid described herein, a vector system described herein, a delivery composition described herein, a lentiviral vector described herein, a ribonucleoprotein complex described herein, a virus-like particle described herein, or a eukaryotic cell described herein in the diagnosis, treatment, or prevention of a disease or disorder associated with a target nucleic acid.

[0172] Another aspect of the present disclosure relates to the use of the CRISPR-Cas system described herein in the diagnosis, treatment, or prevention of a disease or disorder associated with a target nucleic acid.

[0173] Another aspect of the present disclosure relates to a method of diagnosing, treating or preventing a disease or condition associated with a target nucleic acid, the method comprising administering a Cas protein as described herein, a fusion protein as described herein, a guide polynucleotide as described herein, a CRISPR-Cas system as described herein, or an isolated nucleic acid as described herein to a sample of a subject in need thereof or to a subject in need thereof.

[0174] Another aspect of the present disclosure relates to the use of a CRISPR-Cas system described herein, a Cas protein described herein, a fusion protein, a recombinant protein or conjugate described herein, a guide polynucleotide described herein, a nucleic acid described herein, a vector system described herein, a delivery composition described herein, a lentiviral vector described herein, a ribonucleoprotein complex described herein, a virus-like particle described herein, or a eukaryotic cell described herein in the preparation of a medicament for diagnosing, treating or preventing a disease or disorder associated with a target nucleic acid.

[0175] Another aspect of the present disclosure relates to the use of the CRISPR-Cas system described herein in the preparation of a medicament for diagnosing, treating, or preventing a disease or condition associated with a target nucleic acid. The above preferred conditions may be arbitrarily combined, consistent with common knowledge in the art, to yield preferred embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0176] Figures 1A-1D show partial alignment results of multiple sequence alignments of multiple Cas13a proteins and RcCas13a in Example 7.

[0177] Figures 2A-2C show partial alignment results of multiple sequence alignments of multiple Cas13b proteins and PbuCas13b in Example 7.

[0178] Figures 3A-3B show partial alignment results of multiple sequence alignments of multiple Cas13d proteins and EsiCas13d in Example 7.

[0179] Figures 4 to 6 respectively show the secondary structures of the direct repeat sequences corresponding to C13-52 / C13-55 / C13-88 predicted using RNA fold.

[0180] FIG7 shows the test results of targeted knockdown of Aqp1 RNA. DETAILED DESCRIPTION

[0181] The present disclosure is further illustrated by way of examples below, but the present disclosure is not limited to the scope of the examples. Experimental methods in the following examples without specifying specific conditions were performed according to conventional methods and conditions, or selected according to the product specifications.

[0182] As used herein, the term "sequence identity" (identity or percent identity) is used to refer to the matching of sequences between two polypeptides or between two nucleic acids. When a certain position in the two sequences being compared is occupied by the same base or amino acid monomer subunit (for example, a certain position in each of the two DNA molecules is occupied by adenine, or a certain position in each of the two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percent sequence identity" (percent identity) between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared × 100%. For example, if 6 out of 10 positions in two sequences match, then the two sequences have 60% sequence identity. Typically, comparisons are made when two sequences are aligned to produce maximum sequence identity. Such comparisons can be performed using published and commercially available alignment algorithms and programs, such as, but not limited to, ClustalΩ, MAFFT, Probcons, T-Coffee, Probalign, BLAST, which can be reasonably selected by one of ordinary skill in the art. Those skilled in the art can determine appropriate parameters for aligning sequences, including, for example, any algorithms needed to achieve better alignment or optimal comparison over the entire length of the sequences being compared, as well as any algorithms needed to achieve better alignment or optimal comparison over a portion of the sequences being compared.

[0183] As used herein, the term "guide polynucleotide" is used to refer to the molecule that forms a CRISPR complex with the Cas protein in the CRISPR-Cas system and guides the CRISPR complex to the target sequence. The English abbreviation can be expressed as gRNA, and "guide polynucleotide" can be used interchangeably with "gRNA" and "guide RNA". Typically, the guide polynucleotide comprises a backbone sequence connected to a guide sequence, and the guide sequence can hybridize with the target sequence. The backbone sequence usually comprises a direct repeat sequence and sometimes may also comprise a tracrRNA sequence. When the backbone sequence does not comprise a tracrRNA sequence, the guide polynucleotide comprises a guide sequence and a direct repeat sequence, and the guide polynucleotide may also be referred to as crRNA.

[0184] As used herein, the term "target nucleic acid" refers to a polynucleotide containing a target sequence, representing a specific sequence or its reverse complement that one wishes to bind, target, or modify using a gene editing system. Target nucleic acids include both target nucleic acids and target DNA. When using the CRISPR-Cas13 editing system, the "target nucleic acid" is a "target nucleic acid," such as a complete mature mRNA molecule or pre-mRNA molecule, or a fragment thereof; when using the CRISPR-Cas9 editing system, the "target nucleic acid" is a "target DNA."

[0185] As used herein, the term "extracellular vesicles" refers to endogenous nanoscale vesicles that transport certain substances (including but not limited to RNA and proteins). "Extracellular vesicles" include but are not limited to exosomes, microvesicles, etc.

[0186] Cas proteins

[0187] As used herein, Cas protein refers to the proteins involved in the CRISPR-Cas system, including but not limited to CRISPR-Cas effector proteins and CRISPR-Cas accessory proteins. In the present disclosure, Cas protein mainly refers to a class of proteins (often referred to as Cas13) applied to type VI CRISPR-Cas systems, which are mainly used for nucleic acid cutting, including Cas13 effector proteins and optional CRISPR-Cas accessory proteins encoding Cas13 effector proteins upstream or downstream. In certain embodiments, the accessory protein enhances the ability of Cas13 effector proteins to target RNA.

[0188] In the present disclosure, the amino acid sequence of the Cas protein has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity compared to the sequence shown in any one of SEQ ID NOs: 1-184, 426-428. In some embodiments, the amino acid sequence of the Cas protein has 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 66%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 77%, 77%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 88%, 88%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 99%, 99%, 97%, 98%, 99%, or 100% sequence identity compared to any one of SEQ ID NOs: 1-184, 426-428.

[0189] CRISPR-Cas13 system

[0190] Class 2 CRISPR-Cas systems confer diverse adaptive immunity mechanisms to microorganisms. This article presents an analysis of prokaryotic genomes and metagenomes to identify previously uncharacterized RNA-guided, RNA-targeting CRISPR-Cas13 systems, comprising hundreds of Cas proteins, including C13-82, which are classified as Type VI systems. Based on the robust activity of the engineered CRISPR-Cas13 system in human cells, as a compact, single-effector Cas13 enzyme, it can be flexibly packaged into AAV vectors. Our results demonstrate the use of Cas proteins as programmable RNA-binding modules for efficient targeting of cellular RNA, providing a versatile platform for transcriptome engineering, as well as therapeutic and diagnostic approaches.

[0191] One aspect of the present disclosure relates to a CRISPR-Cas system (particularly a CRISPR-Cas13 system), a composition or a kit comprising:

[0192] (a) a Cas protein or fusion protein having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to any one of SEQ ID NOs: 1-184, 426-428, or a nucleic acid encoding the Cas protein or fusion protein; and

[0193] (b) a guide polynucleotide or a nucleic acid encoding a guide polynucleotide;

[0194] The guide polynucleotide comprises a direct repeat sequence connected to a guide sequence, the guide sequence is engineered to hybridize with the target nucleic acid, and the guide polynucleotide is capable of forming a CRISPR complex with the Cas protein and guiding the sequence-specific binding of the CRISPR complex to the target nucleic acid.

[0195] Optionally, the guide sequence is located at the 3' end of the direct repeat sequence. Optionally, the guide sequence is located at the 5' end of the direct repeat sequence.

[0196] In some embodiments, the guide polynucleotide comprises a direct repeat sequence linked to a guide sequence, the guide sequence is engineered to hybridize to a target nucleic acid, and the guide polynucleotide is capable of forming a CRISPR complex with a Cas protein and directing the CRISPR complex sequence-specifically to bind to and cleave the target nucleic acid.

[0197] In some embodiments, the polynucleotide sequence encoding the Cas protein or fusion protein and / or the polynucleotide sequence encoding the guide polynucleotide is operably linked to a regulatory sequence. In some embodiments, the polynucleotide sequence encoding the Cas protein or fusion protein is operably linked to a regulatory sequence. In some embodiments, the polynucleotide sequence encoding the guide polynucleotide is operably linked to a regulatory sequence. In some embodiments, the regulatory sequence of the polynucleotide sequence encoding the Cas protein or fusion protein is the same as or different from the regulatory sequence of the polynucleotide sequence encoding the guide polynucleotide.

[0198] In some embodiments, the Cas protein herein has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% sequence identity compared to any one of SEQ ID NOs: 1-184, 426-428. When the CRISPR-Cas13 system comprises a fusion protein comprising a Cas13 protein and a protein domain and / or a polypeptide tag, the percentage of sequence identity between the Cas13 portion of the fusion protein and the reference sequence is calculated.

[0199] In some embodiments, a CRISPR-Cas system is provided, comprising:

[0200] (a) a Cas protein or fusion protein having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 2, 103-105, 61-62, 106-107, or a nucleic acid encoding the Cas protein or fusion protein; and

[0201] (b) a guide polynucleotide or a nucleic acid encoding a guide polynucleotide;

[0202] The guide polynucleotide comprises a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to hybridize to a target nucleic acid, the guide polynucleotide being capable of forming a CRISPR complex with the Cas protein and directing sequence-specific binding of the CRISPR complex to the target nucleic acid;

[0203] Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to any one of SEQ ID NOs: 186, 287-291, 245-246;

[0204] Optionally, the fusion protein comprises the Cas protein or a functional fragment thereof;

[0205] Optionally, the guide sequence is located at the 3' end of the direct repeat sequence.

[0206] In some embodiments, a CRISPR-Cas system is provided, comprising:

[0207] (a) a Cas protein or fusion protein having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequence shown in SEQ ID NO: 2, or a nucleic acid encoding the Cas protein or fusion protein; and

[0208] (b) a guide polynucleotide or a nucleic acid encoding a guide polynucleotide;

[0209] The guide polynucleotide comprises a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to hybridize to a target nucleic acid, the guide polynucleotide being capable of forming a CRISPR complex with the Cas protein and directing sequence-specific binding of the CRISPR complex to the target nucleic acid;

[0210] Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to the sequence shown in SEQ ID NO: 186;

[0211] Optionally, the fusion protein comprises the Cas protein or a functional fragment thereof;

[0212] Optionally, the guide sequence is located at the 3' end of the direct repeat sequence.

[0213] In some embodiments, a CRISPR-Cas system is provided, comprising:

[0214] (a) a Cas protein or fusion protein having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequence shown in SEQ ID NO: 103, or a nucleic acid encoding the Cas protein or fusion protein; and

[0215] (b) a guide polynucleotide or a nucleic acid encoding a guide polynucleotide;

[0216] The guide polynucleotide comprises a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to hybridize to a target nucleic acid, the guide polynucleotide being capable of forming a CRISPR complex with the Cas protein and directing sequence-specific binding of the CRISPR complex to the target nucleic acid;

[0217] Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to the sequence shown in SEQ ID NO: 287;

[0218] Optionally, the fusion protein comprises the Cas protein or a functional fragment thereof;

[0219] Optionally, the guide sequence is located at the 3' end of the direct repeat sequence.

[0220] In some embodiments, a CRISPR-Cas system is provided, comprising:

[0221] (a) a Cas protein or fusion protein having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequence shown in SEQ ID NO: 104, or a nucleic acid encoding the Cas protein or fusion protein; and

[0222] (b) a guide polynucleotide or a nucleic acid encoding a guide polynucleotide;

[0223] The guide polynucleotide comprises a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to hybridize to a target nucleic acid, the guide polynucleotide being capable of forming a CRISPR complex with the Cas protein and directing sequence-specific binding of the CRISPR complex to the target nucleic acid;

[0224] Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to the sequence shown in SEQ ID NO: 288;

[0225] Optionally, the fusion protein comprises the Cas protein or a functional fragment thereof;

[0226] Optionally, the guide sequence is located at the 3' end of the direct repeat sequence.

[0227] In some embodiments, a CRISPR-Cas system is provided, comprising:

[0228] (a) a Cas protein or fusion protein having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequence shown in SEQ ID NO: 105, or a nucleic acid encoding the Cas protein or fusion protein; and

[0229] (b) a guide polynucleotide or a nucleic acid encoding a guide polynucleotide;

[0230] The guide polynucleotide comprises a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to hybridize to a target nucleic acid, the guide polynucleotide being capable of forming a CRISPR complex with the Cas protein and directing sequence-specific binding of the CRISPR complex to the target nucleic acid;

[0231] Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to the sequence shown in SEQ ID NO: 289;

[0232] Optionally, the fusion protein comprises the Cas protein or a functional fragment thereof;

[0233] Optionally, the guide sequence is located at the 3' end of the direct repeat sequence.

[0234] In some embodiments, a CRISPR-Cas system is provided, comprising:

[0235] (a) a Cas protein or fusion protein having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequence shown in SEQ ID NO: 61, or a nucleic acid encoding the Cas protein or fusion protein; and

[0236] (b) a guide polynucleotide or a nucleic acid encoding a guide polynucleotide;

[0237] The guide polynucleotide comprises a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to hybridize to a target nucleic acid, the guide polynucleotide being capable of forming a CRISPR complex with the Cas protein and directing sequence-specific binding of the CRISPR complex to the target nucleic acid;

[0238] Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to the sequence shown in SEQ ID NO: 245;

[0239] Optionally, the fusion protein comprises the Cas protein or a functional fragment thereof;

[0240] Optionally, the guide sequence is located at the 3' end of the direct repeat sequence.

[0241] In some embodiments, a CRISPR-Cas system is provided, comprising:

[0242] (a) a Cas protein or fusion protein having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequence shown in SEQ ID NO: 62, or a nucleic acid encoding the Cas protein or fusion protein; and

[0243] (b) a guide polynucleotide or a nucleic acid encoding a guide polynucleotide;

[0244] The guide polynucleotide comprises a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to hybridize to a target nucleic acid, the guide polynucleotide being capable of forming a CRISPR complex with the Cas protein and directing sequence-specific binding of the CRISPR complex to the target nucleic acid;

[0245] Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to the sequence shown in SEQ ID NO: 246;

[0246] Optionally, the fusion protein comprises the Cas protein or a functional fragment thereof;

[0247] Optionally, the guide sequence is located at the 3' end of the direct repeat sequence.

[0248] In some embodiments, a CRISPR-Cas system is provided, comprising:

[0249] (a) a Cas protein or fusion protein having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequence shown in SEQ ID NO: 106, or a nucleic acid encoding the Cas protein or fusion protein; and

[0250] (b) a guide polynucleotide or a nucleic acid encoding a guide polynucleotide;

[0251] The guide polynucleotide comprises a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to hybridize to a target nucleic acid, the guide polynucleotide being capable of forming a CRISPR complex with the Cas protein and directing sequence-specific binding of the CRISPR complex to the target nucleic acid;

[0252] Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to the sequence shown in SEQ ID NO: 290;

[0253] Optionally, the fusion protein comprises the Cas protein or a functional fragment thereof;

[0254] Optionally, the guide sequence is located at the 3' end of the direct repeat sequence.

[0255] In some embodiments, a CRISPR-Cas system is provided, comprising:

[0256] (a) a Cas protein or fusion protein having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequence shown in SEQ ID NO: 107, or a nucleic acid encoding the Cas protein or fusion protein; and

[0257] (b) a guide polynucleotide or a nucleic acid encoding a guide polynucleotide;

[0258] The guide polynucleotide comprises a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to hybridize to a target nucleic acid, the guide polynucleotide being capable of forming a CRISPR complex with the Cas protein and directing sequence-specific binding of the CRISPR complex to the target nucleic acid;

[0259] Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity compared to the sequence shown in SEQ ID NO: 291;

[0260] Optionally, the fusion protein comprises the Cas protein or a functional fragment thereof;

[0261] Optionally, the guide sequence is located at the 3' end of the direct repeat sequence.

[0262] In some embodiments of the present disclosure, the Cas protein or its functional fragment is fused to a protein domain and / or a polypeptide tag, and the fusion does not change the original function of the Cas protein or its functional fragment.

[0263] In some embodiments of the present disclosure, the Cas protein or a functional fragment thereof is fused to a localization tag that provides subcellular localization;

[0264] Optionally, the localization tag is selected from a nuclear localization signal (NLS) or a nuclear export signal (NES).

[0265] In some embodiments of the present disclosure, the Cas protein or a functional fragment thereof is fused to any one or more protein domains and / or polypeptide tags selected from the group consisting of a cytosine deaminase domain, an adenosine deaminase domain, a translation activation domain, a translation repression domain, an RNA methylation domain, an RNA demethylation domain, a nuclease domain, a splicing factor domain, a reporter domain, an affinity domain, a subcellular localization signal, a reporter tag, and an affinity tag.

[0266] In some embodiments of the present disclosure, the Cas protein or a functional fragment thereof is covalently linked to the Cas protein domain; and / or

[0267] The Cas protein or its functional fragment is covalently or non-covalently linked to a conjugate molecule, the conjugate molecule does not change the original function of the Cas protein or its functional fragment, and the conjugate molecule is selected from at least one of a polymer, a small molecule compound, an antibody, an oligopeptide, a detectable label, and a polypeptide other than the Cas protein.

[0268] In some embodiments of the present disclosure, the structure of the fusion protein is NLS-Cas protein-SV40NLS-nucleoplasminNLS.

[0269] The CRISPR-Cas13 system herein can be introduced into a cell (or cell-free system) in a variety of non-limiting ways: (i) as Cas13 mRNA and a guide polynucleotide, (ii) as part of a single vector or plasmid, or divided into multiple vectors or plasmids, (iii) as a separate Cas13 protein and a guide polynucleotide, or (iv) as an RNP complex of a Cas13 protein and a guide polynucleotide.

[0270] In some embodiments, the CRISPR-Cas13 system, composition, or kit comprises a nucleic acid molecule encoding a Cas13 protein, wherein the coding sequence is codon optimized for expression in eukaryotic cells. In some embodiments, the CRISPR-Cas13 system, composition, or kit comprises a nucleic acid molecule encoding a Cas13 protein, wherein the coding sequence is codon optimized for expression in mammalian cells. In some embodiments, the CRISPR-Cas13 system, composition, or kit comprises a nucleic acid molecule encoding a Cas13 protein, wherein the coding sequence is codon optimized for expression in human cells.

[0271] In some embodiments, the nucleic acid molecule encoding the Cas13 protein is a plasmid. In some embodiments, the nucleic acid molecule encoding the Cas13 protein is part of a viral vector genome, such as a DNA genome of an AAV vector flanked by ITRs. In some embodiments, the nucleic acid molecule encoding the Cas13 protein is mRNA.

[0272] Guide polynucleotide

[0273] In some embodiments, the guidance polynucleotide of the CRISPR-Cas system is a guide RNA. In some embodiments, the guidance polynucleotide is a chemically modified guidance polynucleotide. In some embodiments, the guidance polynucleotide comprises at least one chemically modified nucleotide. In some embodiments, the guidance polynucleotide is a hybrid RNA-DNA guide. In some embodiments, the guidance polynucleotide is a hybrid RNA-LNA (locked nucleic acid) guide.

[0274] In some embodiments, the guide polynucleotide comprises at least one guide sequence (also known as a spacer sequence) linked to at least one direct repeat (DR). In some embodiments, the guide sequence is located at the 3' end of the direct repeat. In some embodiments, the guide sequence is located at the 5' end of the direct repeat.

[0275] In some embodiments, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 185-368, 429-431, or the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the reverse complement of any one of SEQ ID NOs: 185-368, 429-431. In some embodiments, the direct repeat sequence has 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 66%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 77%, 77%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 88%, 88%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 99%, 99%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 185-368, 429-431, or the direct repeat sequence has 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, The reverse complement of ID NOs: 185-368, 429-431 has a sequence identity of 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 66%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 77%, 77%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 88%, 88%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 99%, 99%, 97%, 98%, 99% or 100%.

[0276] In some embodiments, the guide sequence comprises at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, or at least 30 nucleotides. In some embodiments, the guide sequence comprises no more than 60 nucleotides, no more than 55 nucleotides, no more than 50 nucleotides, no more than 45 nucleotides, no more than 40 nucleotides, no more than 35 nucleotides, or no more than 30 nucleotides. In some embodiments, the guide sequence comprises 15-20 nucleotides, 20-25 nucleotides, 25-30 nucleotides, 30-35 nucleotides, or 35-40 nucleotides.

[0277] In some embodiments, the guide sequence has sufficient complementarity to the target nucleic acid sequence to hybridize to the target nucleic acid and guide sequence-specific binding of the CRISPR-Cas complex to the target nucleic acid. In some embodiments, the guide sequence has 100% complementarity to the target nucleic acid (or region of the RNA to be targeted), but the guide sequence can have less than 100% complementarity to the target nucleic acid, for example, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% complementarity.

[0278] In some embodiments, the guide sequence is engineered to hybridize to the target nucleic acid with no more than two nucleotide mismatches. In some embodiments, the guide sequence is engineered to hybridize to the target nucleic acid with no more than one nucleotide mismatches. In some embodiments, the guide sequence is engineered to hybridize to the target nucleic acid with or without mismatches.

[0279] In some embodiments, the same direction repeat sequence comprises at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides or at least 36 nucleotides. In some embodiments, the same direction repeat sequence comprises no more than 60 nucleotides, no more than 55 nucleotides, no more than 50 nucleotides, no more than 45 nucleotides, no more than 40 nucleotides or no more than 35 nucleotides. In some embodiments, the same direction repeat sequence comprises 20-25 nucleotides, 25-30 nucleotides, 30-35 nucleotides or 35-40 nucleotides.

[0280] In some embodiments, the CRISPR-Cas system, composition, or kit comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different guide polynucleotides. In some embodiments, the guide polynucleotides target at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different target nucleic acid molecules, or target at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different regions of one or more target nucleic acid molecules.

[0281] In some embodiments, the instructing polynucleotide includes a constant directional repeat sequence located upstream of the variable guide sequence. In some embodiments, a plurality of instructing polynucleotides is a part for an array (which can be a part for a vector, such as a viral vector or a plasmid). For example, the instructing array including the sequence DR-spacer-DR-spacer-DR-spacer can include three unique unprocessed instructing polynucleotides (one for each DR-spacer sequence). Once introduced into a cell or a cell-free system, the array is processed into three separate mature instructing polynucleotides by Cas13 protein. This allows multiplexing, such as by delivering a plurality of instructing polynucleotides to a cell or system to target a plurality of target nucleic acids or a plurality of regions within a single target nucleic acid.

[0282] The ability of a guide polynucleotide to direct sequence-specific binding of a CRISPR complex to a target nucleic acid can be assessed by any suitable assay. For example, components of a CRISPR system sufficient to form a CRISPR complex, including the guide polynucleotide to be tested, can be provided to a host cell with a corresponding target nucleic acid molecule, such as by transfection of a vector encoding the components of the CRISPR complex, and preferential cleavage within the target sequence can be assessed. Similarly, cleavage of a target nucleic acid sequence can be assessed in vitro by providing a target nucleic acid, components of a CRISPR complex, including the guide polynucleotide to be tested, and a control guide polynucleotide that is different from the test guide polynucleotide, and comparing the ability to bind to the target nucleic acid or the rate at which the target nucleic acid is cleaved between the test and control guide polynucleotides.

[0283] Cas13 mutants

[0284] In some embodiments, the Cas13 proteins provided herein comprise one or more mutations, such as single amino acid insertions, single amino acid deletions, single amino acid substitutions, or a combination thereof, compared to wild-type Cas13 proteins (SEQ ID NOs: 1-184, 426-428). In some instances, the Cas protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89 or 90 amino acid changes (e.g., insertions, deletions or substitutions), but retain the ability to bind to a target nucleic acid molecule that is complementary to the guide sequence of the guide polynucleotide and / or retain the ability to process guide array RNA transcripts into guide polynucleotide molecules. In some examples, the Cas13 protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 , 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89 or 90 amino acid changes (e.g., insertions, deletions or substitutions) but retains the ability to bind to a target nucleic acid molecule that is complementary to the guide sequence of the guide polynucleotide. In some examples, the Cas13 protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 amino acid changes (e.g., insertions, deletions, or substitutions) compared to a wild-type Cas13 protein, but retains the ability to bind to a target nucleic acid molecule that is complementary to the guide sequence of the guide polynucleotide, and / or retains the ability to process guide array RNA transcripts into guide polynucleotide molecules.

[0285] In some embodiments, the Cas13 protein comprises one or more mutations in the catalytic domain and has reduced RNA cleavage activity. In some embodiments, the Cas13 protein comprises a mutation in the catalytic domain and has reduced RNA cleavage activity. In some embodiments, the Cas13 protein comprises one or more mutations in one or two HEPN domains and substantially lacks RNA cleavage activity. In some embodiments, the Cas13 protein comprises a mutation in one or two HEPN domains and substantially lacks RNA cleavage activity. In some embodiments, "substantially lacking RNA cleavage activity" refers to retaining only ≤50%, ≤40%, ≤30%, ≤20%, ≤10%, ≤5% or ≤1% RNA cleavage activity compared to the wild-type Cas13 protein, or no detectable RNA cleavage activity.

[0286] In some embodiments, the Cas13 mutant protein has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity compared to any one of SEQ ID NOs: 1-184, 426-428. When the CRISPR-Cas13 system comprises a fusion protein of Cas13 with a protein domain and / or a polypeptide tag, the percentage of sequence identity between the Cas13 portion of the fusion protein and the reference sequence is calculated.

[0287] In some embodiments, the Cas13 mutant protein can form a CRISPR complex with the guide polynucleotide, and the CRISPR complex can specifically bind to the target nucleic acid sequence.

[0288] In some embodiments, the Cas13 mutant protein can form a CRISPR complex with a guide polynucleotide comprising a direct repeat sequence linked to a guide sequence engineered to direct sequence-specific binding of the CRISPR complex to a target nucleic acid.

[0289] A type of modification or mutation includes replacing an amino acid residue with a similar biochemical property with an amino acid, i.e., a conservative substitution (e.g., conservative substitution of 1-4, 1-8, 1-10, or 1-20 amino acids). Typically, conservative substitutions have little or no effect on the activity of the resulting protein or peptide. For example, conservative substitutions are amino acid substitutions in Cas13 proteins that substantially do not affect the binding of Cas13 proteins to target nucleic acid molecules complementary to the gRNA molecule guide sequence, and / or the process of processing guide array RNA transcripts into gRNA molecules. Alanine scanning can be used to identify which amino acid residues in Cas13 proteins can tolerate amino acid substitutions. In one example, when alanine or other conservative amino acids are replaced by 1-4, 1-8, 1-10, or 1-20 natural amino acids, the change in the ability of variant Cas13 proteins to modify gene expression in the CRISPR-Cas system is no more than 25%, for example, no more than 20%, for example, no more than 10%. Examples of amino acids that can be substituted and are considered conservative substitutions include: Ala for Ser; Arg for Lys; Asn for Gln or His; Asp for Glu; Cys for Ser; Gln for Asn; Glu for Pro for Gly; His for Asn or Gln; Ile for Leu or Val; Leu for Ile or Val; Lys for Arg or Gln; Met for Leu or Ile; Phe for Met, Leu or Tyr; Ser for Thr; Ser for Thr; Trp for Tyr; Tyr for Trp or Phe; Val for Ile or Leu.

[0290] More substantial changes can be made by using less conservative substitutions, for example, by selecting residues that differ more in maintaining: (a) the structure of the polypeptide backbone in the region where the substitution occurs, for example, as a helical or sheet conformation; (b) the charge or hydrophobicity of the region that interacts with the target site; or (c) the bulk of the side chain. Substitutions that would generally be expected to produce the greatest changes in polypeptide function are (a) substitutions between a hydrophilic residue (e.g., serine or threonine) and a hydrophobic residue (e.g., leucine, isoleucine, phenylalanine, valine, or alanine); (b) substitutions between cysteine ​​or proline and any other residue; (c) substitutions between a residue with a positively charged side chain (e.g., lysine, arginine, or histidine) and a negatively charged residue (e.g., glutamic acid or aspartic acid); or (d) substitutions between a residue with a bulky side chain (e.g., phenylalanine) and a residue without a side chain (e.g., glycine).

[0291] Subcellular localization signal (or localization signal)

[0292] In some embodiments, Cas13 protein or its functional fragment is fused with at least one homologous or heterologous subcellular localization signal.Exemplary subcellular localization signals include organelle localization signals, such as nuclear localization signal (NLS), nuclear export signal (NES) or mitochondrial localization signal.

[0293] In some embodiments, the Cas13 protein or its functional fragment is fused to at least one homologous or heterologous NLS. In some embodiments, the Cas13 protein or its functional fragment is fused to at least two NLSs. In some embodiments, the Cas13 protein or its functional fragment is fused to at least three NLSs. In some embodiments, the Cas13 protein or its functional fragment is fused to at least one N-terminal NLS and at least one C-terminal NLS. In some embodiments, the Cas13 protein or its functional fragment is fused to at least two C-terminal NLSs. In some embodiments, the Cas13 protein or its functional fragment is fused to at least two N-terminal NLSs.

[0294] In some embodiments, the Cas protein or its functional fragment is fused to a homologous or heterologous NES. In some embodiments, the Cas protein or its functional fragment is fused to at least two NESs. In some embodiments, the Cas protein or its functional fragment is fused to at least three NESs. In some embodiments, the Cas protein or its functional fragment is fused to at least one N-terminal NES and at least one C-terminal NES. In some embodiments, the Cas protein or its functional fragment is fused to at least two C-terminal NESs. In some embodiments, the Cas protein or its functional fragment is fused to at least two N-terminal NESs.

[0295] In some embodiments, the NES is independently selected from adenovirus type 5 E1B NES, HIV Rev NES, MAPK NES, and PTK2 NES.

[0296] In some embodiments, the Cas protein or its functional fragment is fused to a homologous or heterologous NLS and NES, with a cleavable linker between the NLS and the NES. In some embodiments, the NES promotes the production of delivery particles (e.g., virus-like particles) comprising the Cas protein or its functional fragment in the production cell line. In some embodiments, the cleavage of the linker in the target cell can expose the NLS and promote the nuclear localization of the Cas protein or its functional fragment in the target cell.

[0297] Protein domain and peptide tags

[0298] In some embodiments, the Cas protein or its functional fragment is covalently linked or fused to a homologous or heterologous protein domain and / or a polypeptide tag. In some embodiments, the Cas protein or its functional fragment is fused to a homologous or heterologous protein domain and / or a polypeptide tag.

[0299] In some embodiments, the protein domains and polypeptide tags are selected from the group consisting of: a cytosine deaminase domain, an adenosine deaminase domain, a translation activation domain, a translation repression domain, an RNA methylation domain, an RNA demethylation domain, a nuclease domain, a splicing factor domain, a reporter domain, an affinity domain, a subcellular localization signal, a reporter tag, and an affinity tag.

[0300] In some embodiments, the protein domain comprises a cytosine deaminase domain, an adenosine deaminase domain, a translation activation domain, a translation repression domain, an RNA methylation domain, an RNA demethylation domain, a ribonuclease domain, a splicing factor domain, a reporter domain, and an affinity domain. In some embodiments, the polypeptide tag comprises a reporter tag and an affinity tag.

[0301] In some embodiments, the amino acid sequence of the protein domain is ≥40 amino acids, ≥50 amino acids, ≥60 amino acids, ≥70 amino acids, ≥80 amino acids, ≥90 amino acids, ≥100 amino acids, ≥150 amino acids, ≥200 amino acids, ≥250 amino acids, ≥300 amino acids, ≥350 amino acids, or ≥400 amino acids in length.

[0302] Exemplary protein domains include domains that can cleave RNA (e.g., PIN endonuclease domains, NYN domains, SMR domains from SOT1, or RNase domains from staphylococcal nuclease), domains that can affect RNA stability (e.g., tristetraprolin (TTP) or domains from UPF1, EXOSC5, and STAU1), domains that can edit nucleotides or ribonucleotides (e.g., cytidine deaminases, PPR proteins, adenosine deaminases, ADAR family proteins, or APOBEC family proteins), domains that can activate translation (e.g., domains of eIF4E and other translation initiation factors, yeast poly(A) binding protein, or GLD2), domains that can inhibit translation (e.g., Pumilio or FBFPUF proteins, deadenylases, These proteins may be expressed in proteins that are expressed by CAF1, an Argonaute protein), a domain that can methylate RNA (e.g., a domain from an m6A methyltransferase factor such as METTL14, METTL3, or WTAP), a domain that can demethylate RNA (e.g., human alkylation repair homolog 5), a domain that can affect splicing (e.g., the RS-rich domain of SRSF1, the Gly-rich domain of hnRNPA1, the alanine-rich motif of RBM4, or the proline-rich motif of DAZAP1), a domain that can enable affinity purification or immunoprecipitation, and a domain that can enable proximity-dependent protein tagging and recognition (e.g., a biotin ligase such as BirA or a peroxidase such as APEX2 to biotinylate target DNA-interacting proteins).

[0303] In some embodiments, the protein domain comprises an adenosine deaminase domain. In some embodiments, a Cas protein with a mutant HEPN domain, a catalytically inactive Cas protein, or a functional fragment thereof is covalently linked or fused to an adenosine deaminase domain to guide the A-to-I deaminase activity of RNA transcripts in mammalian cells. Cox et al., Science 358 (6366): 1019-1027 (2017) describes an adenosine deaminase domain for targeting A-to-I RNA editing based on ADAR2 engineering, which is incorporated herein by reference in its entirety. In other embodiments, the adenosine deaminase domain is covalently linked or fused to a linker protein that is capable of binding to an aptamer sequence inserted into or attached to a guidance polynucleotide, thereby allowing the adenosine deaminase domain to be non-covalently linked to a Cas protein or a functional fragment thereof that is complexed with a guidance polynucleotide.

[0304] In some embodiments, the protein domain includes a cytosine deaminase domain. In some embodiments, the Cas protein with a mutant HEPN domain, a catalytically inactivated Cas protein, or its functional fragment is covalently linked to or fused with a cytosine deaminase domain to guide the C-to-U deaminase activity of RNA transcripts in mammalian cells. Abudayyehetal., Science 365 (6451): 382-386 (2019) describes a cytosine deaminase domain for targeting C-to-U RNA editing evolved from ADAR2, which is incorporated herein by reference in its entirety. In other embodiments, the cytosine deaminase domain is covalently linked to or fused with a linker protein, which is capable of binding to an aptamer sequence inserted into or attached to a guidance polynucleotide, thereby allowing the cytosine deaminase domain to be non-covalently linked to a Cas protein or its functional fragment compounded with a guidance polynucleotide.

[0305] In some embodiments, the protein domain comprises a splicing factor domain. In some embodiments, a Cas protein with a mutant HEPN domain, a catalytically inactive Cas protein, or a functional fragment thereof is covalently linked or fused to a splicing factor domain to guide the alternative splicing of a target nucleic acid in a mammalian cell. Konermann et al., Cell 173 (3): 665-676 (2018) describes a splicing factor domain for targeted alternative splicing, which is incorporated herein by reference in its entirety. Non-limiting examples of splicing factor domains include the RS-rich domain of SRSF1, the Gly-rich domain of hnRNPA1, the alanine-rich motif of RBM4, or the proline-rich motif of DAZAP1. In other embodiments, the splicing factor domain is covalently linked or fused to a linker protein that is capable of binding to an aptamer sequence inserted into or attached to a guide polynucleotide, thereby allowing the splicing factor domain to be non-covalently linked to a Cas protein or its functional fragment complexed with the guide polynucleotide.

[0306] In some embodiments, the protein domain comprises a translation activation domain. In some embodiments, the Cas protein with a mutation HEPN domain, a catalytically inactivated Cas protein, or its functional fragment is covalently linked or fused to the translation activation domain to activate or increase the expression of the target nucleic acid. Non-limiting examples of translation activation domains include eIF4E and other translation initiation factors, yeast poly (A) binding protein or the domain of GLD2. In other embodiments, the translation activation domain is covalently linked or fused to a linker protein that can bind to an aptamer sequence inserted or attached to a guide polynucleotide, thereby allowing the translation activation domain to be non-covalently linked to a Cas13 protein or its functional fragment compounded with a guide polynucleotide.

[0307] In some embodiments, the protein domain comprises a translation repression domain. In some embodiments, a Cas protein having a mutant HEPN domain, a catalytically inactive Cas protein, or a functional fragment thereof is covalently linked or fused to a translation repression domain to inhibit or reduce the expression of a target nucleic acid. Non-limiting examples of translation repression domains include Pumilio or FBFPUF proteins, deadenylase, CAF1, Argonaute proteins. In other embodiments, the translation repression domain is covalently linked or fused to a linker protein that is capable of binding to an aptamer sequence inserted into or attached to a guide polynucleotide, thereby allowing the translation repression domain to be non-covalently linked to a Cas protein or a functional fragment thereof complexed with a guide polynucleotide.

[0308] In some embodiments, the protein domain comprises an RNA methylation domain. In some embodiments, a Cas protein having a mutant HEPN domain, a catalytically inactive Cas protein, or a functional fragment thereof is covalently linked or fused to an RNA methylation domain for methylation of a target nucleic acid. Non-limiting examples of RNA methylation domains include m6A domains, such as METTL14, METTL3, or WTAP. In other embodiments, the RNA methylation domain is covalently linked or fused to an adapter protein that is capable of binding to an aptamer sequence inserted into or attached to a guide polynucleotide, thereby allowing the RNA methylation domain to be non-covalently linked to a Cas protein or a functional fragment thereof that is complexed with a guide polynucleotide.

[0309] In some embodiments, the protein domain comprises an RNA demethylation domain. In some embodiments, a Cas protein having a mutant HEPN domain, a catalytically inactive Cas13 protein, or a functional fragment thereof is covalently linked or fused to an RNA demethylation domain for demethylation of a target nucleic acid. Non-limiting examples of RNA demethylation domains include human alkylation repair homolog 5 or ALKBH5. In other embodiments, the RNA demethylation domain is covalently linked or fused to an adapter protein that is capable of binding to an aptamer sequence inserted into or attached to a guide polynucleotide, thereby allowing the RNA demethylation domain to be non-covalently linked to a Cas protein or a functional fragment thereof that is complexed with a guide polynucleotide.

[0310] In some embodiments, the protein domain comprises a ribonuclease domain. In some embodiments, a Cas protein having a mutant HEPN domain, a catalytically inactive Cas protein, or a functional fragment thereof is covalently linked or fused to a ribonuclease domain to cleave the target nucleic acid. Non-limiting examples of ribonuclease domains include a PIN endonuclease domain, a NYN domain, an SMR domain from SOT1, or an RNase domain from Staphylococcus nuclease.

[0311] In some embodiments, the protein domain comprises an affinity domain (affinity domain) and / or a reporter domain (reporter domain). In some embodiments, the Cas protein or its functional fragment is covalently linked or fused to a reporter domain, such as a fluorescent protein. Non-limiting examples of reporter domains include GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP.

[0312] In some embodiments, the Cas protein is covalently linked to or fused with a polypeptide tag. In some embodiments, the example of a polypeptide tag is a small polypeptide sequence. In some embodiments, the length of the amino acid sequence of the polypeptide tag is ≤50 amino acids, ≤40 amino acids, ≤30 amino acids, ≤25 amino acids, ≤20 amino acids, ≤15 amino acids, ≤10 amino acids or ≤5 amino acids. In some embodiments, the Cas13 protein is covalently linked to or fused with an affinity tag such as a purification tag. Non-limiting examples of affinity tags include HA-tags, His-tags (e.g., 6-His), Myc-tags, E-tags, S-tags, calmodulin tags, FLAG-tags, GST-tags, MBP-tags, Halo tags or biotin.

[0313] Aptamer Sequence

[0314] In some embodiments, the guidance polynucleotide further comprises an aptamer sequence. In some embodiments, the aptamer sequence is inserted into the loop of the guidance polynucleotide. In some embodiments, the aptamer sequence is inserted into the tetraloop of the guidance polynucleotide. In some embodiments, the aptamer sequence is attached to the end of the guidance polynucleotide.

[0315] Inserting an aptamer sequence into a guide polynucleotide of a CRISPR-Cas system is described in Konermann et al., Nature 517:583–588 (2015), which is incorporated herein by reference in its entirety. In some embodiments, the aptamer sequence comprises an MS2 aptamer sequence, a PP7 aptamer sequence, or a Qβ aptamer sequence.

[0316] Adaptor protein

[0317] In some embodiments, the CRISPR-Cas system further comprises a fusion protein comprising an adaptor protein and a homologous or heterologous protein domain and / or polypeptide tag, or a nucleic acid encoding the fusion protein, wherein the adaptor protein is capable of binding to an aptamer sequence.

[0318] Konermann et al., Nature 517: 583-588 (2015) describes a fusion protein of a connector protein and a protein domain, which is incorporated herein by reference in its entirety. In some embodiments, the connector protein includes MS2 phage coat protein (MCP), PP7 phage coat protein (PCP) or Qβ phage coat protein (QCP). In some embodiments, the protein domain includes a cytosine deaminase domain, an adenosine deaminase domain, a translation activation domain, a translation inhibition domain, an RNA methylation domain, an RNA demethylation domain, a nuclease domain, a splicing factor domain, an affinity domain or a reporter domain.

[0319] Modified guide polynucleotides

[0320] In some embodiments, the guidance polynucleotide comprises a modified nucleotide. In some embodiments, the modified nucleotide comprises 2'-O-methyl, 2'-O-methyl-3'-phosphorothioate or 2'-O-methyl-3'-thioPACE. In some embodiments, the guidance polynucleotide is a chemically modified guidance polynucleotide. Chemically modified guidance polynucleotides are described in Hendel et al., Nat. Biotechnol. 33(9): 985-989 (2015), which is incorporated herein by reference in its entirety.

[0321] In some embodiments, the guide polynucleotide is a hybrid RNA-DNA guide, a hybrid RNA-LNA (locked nucleic acid) guide, a hybrid DNA-LNA guide, or a hybrid DNA-RNA-LNA guide. In some embodiments, the direct repeat sequence comprises one or more ribonucleotides substituted with corresponding deoxyribonucleotides. In some embodiments, the guide sequence comprises one or more ribonucleotides substituted with corresponding deoxyribonucleotides. Hybrid RNA-DNA guide polynucleotides are described in WO2016 / 123230, which is incorporated herein by reference in its entirety.

[0322] Vector system

[0323] Another aspect of the present disclosure relates to a vector system comprising the CRISPR-Cas system herein, the vector system comprising one or more vectors comprising a polynucleotide sequence encoding a Cas protein or fusion protein and a polynucleotide sequence encoding a guide polynucleotide.

[0324] In some embodiments, the vector system comprises at least one plasmid or viral vector (e.g., retrovirus, lentivirus, adenovirus, adeno-associated virus, or herpes simplex virus). In some embodiments, the polynucleotide sequence encoding the Cas protein or fusion protein and the polynucleotide sequence encoding the guide polynucleotide are located on the same vector. In some embodiments, the polynucleotide sequence encoding the Cas protein or fusion protein and the polynucleotide sequence encoding the guide polynucleotide are located on multiple vectors.

[0325] In some embodiments, the polynucleotide sequence encoding Cas protein or fusion protein and / or the polynucleotide sequence encoding guidance polynucleotide are operably connected to a regulatory sequence. In some embodiments, the polynucleotide sequence encoding Cas protein or fusion protein is operably connected to a regulatory sequence. In some embodiments, the polynucleotide sequence encoding guidance polynucleotide is operably connected to a regulatory sequence. In some embodiments, the regulatory sequence of the polynucleotide sequence encoding Cas protein or fusion protein is identical or different from the regulatory sequence of the polynucleotide sequence encoding guidance polynucleotide. In some embodiments, regulatory sequence is optionally selected from promoter, enhancer, internal ribosome entry site (IRES) and other expression control elements (for example, transcription termination signal, such as polyadenylation signal and poly-U sequence). In some embodiments, regulatory sequence includes regulatory sequence that makes nucleotide sequence constitutively expressed in many types of host cells, and regulatory sequence (for example, tissue-specific regulatory sequence) that makes nucleotide sequence expressed only in certain host cells. Tissue-specific promoter can be directly expressed mainly in desired tissue of interest, such as muscle, neuron, bone, skin, blood, specific organs (such as liver, pancreas) or specific cell type (such as lymphocyte). The regulatory sequence can also direct expression in a time-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue or cell type specific. In some embodiments, the regulatory sequence is an enhancer element, such as the WPRE, the CMV enhancer, the R-U5 segment in the LTR of HTLV-1, the SV40 enhancer, or the intronic sequence between exons 2 and 3 of rabbit β-globin.

[0326] In some embodiments, the vector comprises a polIII promoter (e.g., U6 and H1 promoters), a polII promoter (e.g., the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), a cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), an SV40 promoter, a dihydrofolate reductase promoter, a β-actin promoter, a phosphoglycerol kinase (PGK) promoter, or an EF1α promoter), or a polIII promoter and a polII promoter.

[0327] In some embodiments, the promoter is a constitutive promoter, which is continuously active and not regulated by external signals or molecules. Suitable constitutive promoters include, but are not limited to, CMV, RSV, SV40, EF1α, CAG, and β-actin promoters. In some embodiments, the promoter is an inducible promoter regulated by external signals or molecules (e.g., transcription factors).

[0328] In some embodiments, the promoter is a tissue-specific promoter, which can be used to drive the tissue-specific expression of Cas proteins or fusion proteins. Suitable muscle-specific promoters include but are not limited to CK8, MHCK7, myoglobin promoter (Mb), desmin (Desmin) promoter, muscle creatine kinase promoter (MCK) and variants thereof, and SPc5-12 synthetic promoters. Suitable immune cell-specific promoters include but are not limited to B29 promoter (B cells), CD14 promoter (monocytes), CD43 promoter (leukocytes and platelets), CD68 (macrophages) and SV40 / CD43 promoter (leukocytes and platelets). Suitable blood cell-specific promoters include but are not limited to CD43 promoter (leukocytes and platelets), CD45 promoter (hematopoietic cells), INF-β (hematopoietic cells), WASP promoter (hematopoietic cells), SV40 / CD43 promoter (leukocytes and platelets), and SV40 / CD45 promoter (hematopoietic cells). Suitable pancreas-specific promoters include but are not limited to elastase-1 promoter. Suitable endothelial cell-specific promoters include, but are not limited to, the Fit-1 promoter and the ICAM-2 promoter. Suitable neuronal tissue / cell-specific promoters include, but are not limited to, the GFAP promoter (astroglial cells), the SYN1 promoter (neurons), and the NSE / RU5' (mature neurons). Suitable kidney-specific promoters include, but are not limited to, the NphsI promoter (podocytes). Suitable bone-specific promoters include, but are not limited to, the OG-2 promoter (osteoblasts, odontoblasts). Suitable lung-specific promoters include, but are not limited to, the SP-B promoter (lung). Suitable liver-specific promoters include, but are not limited to, the SV40 / Alb promoter. Suitable heart-specific promoters include, but are not limited to, α-MHC.

[0329] AAV vectors

[0330] Another aspect of the present disclosure relates to an adeno-associated viral (AAV) vector comprising the CRISPR-Cas system herein, wherein the adeno-associated viral (AAV) vector comprises DNA encoding the Cas protein or fusion protein herein and a guide polynucleotide.

[0331] Delivery of the CRISPR-Cas system by AAV vectors is described in Maeder et al., Nature Medicine 25: 229-233 (2019), which is incorporated herein by reference in its entirety. In some embodiments, the AAV vector comprises an ssDNA genome comprising a coding sequence for a Cas protein or fusion protein and a guide polynucleotide flanked by ITRs.

[0332] In some embodiments, the CRISPR-Cas system herein is packaged in an AAV vector, such as AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, and AAVrh74. In some embodiments, the CRISPR-Cas system herein is packaged in an AAV vector comprising an engineered capsid with tissue tropism, such as an engineered muscle tropism capsid. Engineering of AAV capsids with tissue tropism by directed evolution is described in Tabebordbare et al., Cell 184: 4919-4938 (2021), which is incorporated herein by reference in its entirety.

[0333] lipid nanoparticles

[0334] Another aspect of the present disclosure relates to lipid nanoparticles (LNPs) comprising the CRISPR-Cas system herein, wherein the LNP comprises a guide polynucleotide herein, and an mRNA encoding a Cas protein or fusion protein herein.

[0335] Gillmore et al., N. Engl. J. Med., 385:493-502 (2021) describes the LNP delivery of the CRISPR-Cas system, which is incorporated herein by reference in its entirety. In some embodiments, in addition to RNA payload (Cas mRNA and guide polynucleotides), lipid nanoparticles (LNPs) also include four components: cations or ionizable lipids, cholesterol, helper lipids and PEG-lipids. In some embodiments, cations or ionizable lipids include cKK-E12, C12-200, ALC-0315, DLin-MC3-DMA, DLin-KC2-DMA, FTT5, Moderna SM-102 and Intellia LP01. In some embodiments, PEG-lipids include PEG-2000-C-DMG, PEG-2000-DMG or ALC-0159. In some embodiments, helper lipids include DSPC. The components of LNPs are described in Paunovska et al., Nature Reviews Genetics 23:265-280 (2022), which is incorporated herein by reference in its entirety.

[0336] Lentiviral vectors

[0337] Another aspect of the present disclosure relates to a lentiviral vector comprising a CRISPR-Cas system herein, wherein the lentiviral vector comprises a guide polynucleotide herein and an mRNA encoding a Cas protein or fusion protein herein. In some embodiments, the lentiviral vector is pseudotyped with a homologous or heterologous envelope protein such as VSV-G. In some embodiments, the mRNA encoding the Cas protein or fusion protein is linked to an aptamer sequence.

[0338] RNP complex

[0339] Another aspect of the present disclosure relates to a ribonucleoprotein complex comprising a CRISPR-Cas system herein, wherein the ribonucleoprotein complex is formed by the guidance polynucleotides herein and Cas protein or fusion protein. In some embodiments, the ribonucleoprotein complex can be delivered to eukaryotic cells, mammalian cells or human cells by microinjection or electroporation. In some embodiments, the ribonucleoprotein complex can be packaged in virus-like particles and delivered to mammals or human subjects in vivo.

[0340] virus-like particles

[0341] Another aspect of the present disclosure relates to a virus-like particle (VLP) comprising the CRISPR-Cas system herein, wherein the virus-like particle comprises: a guide polynucleotide herein, and a Cas protein or fusion protein; or a ribonucleoprotein complex consisting of a guide polynucleotide and a Cas protein or fusion protein.

[0342] Banskota et al. Cell 185 (2): 250-265 (2022), Mangeot et al., Nature Communications 10 (1): 1-15 (2019), Campbell, et al., Molecular Therapy 27: 151-163 (2019), Campbell, et al., Molecular Therapy, 27 (2019): 151-163 and Mangeot et al. Molecular Therapy, 19 (9): 1656-1666 (2011) describe engineered VLPs, which are incorporated herein by reference in their entirety. In some embodiments, engineered virus-like particles (VLPs) are pseudotyped with homologous or heterologous envelope proteins, such as VSV-G. In some embodiments, the Cas protein is fused to a gag protein (e.g., MLV gag) via a cleavable linker, wherein the cleavage of the linker in the target cell exposes an NLS between the linker and the Cas protein. In some embodiments, the fusion protein comprises (e.g., from 5' to 3') a gag protein (e.g., MLVgag), one or more NESs, a cleavable linker, one or more NLSs, and Cas, as in Banskota et al. Cell 185 (2): 250-265 (2022). In some embodiments, the Cas protein is fused to a first dimerization domain that is capable of dimerizing or heterodimerizing with a second dimerization domain fused to a membrane protein, wherein the presence of a ligand promotes dimerization and enriches the Cas protein or fusion protein in a VLP, as in Campbell, et al., Molecular Therapy 27: 151-163 (2019).

[0343] cell

[0344] Another aspect of the present disclosure relates to cells comprising the CRISPR-Cas system herein.Cells (e.g., which can be used to produce a cell-free system) can be eukaryotic or prokaryotic.Examples of such cells include, but are not limited to, bacteria, archaebacteria, plants, fungi, yeast, insects, and mammalian cells, such as lactobacilli, lactococci, bacillus (e.g., bacillus subtilis), Escherichia (e.g., Escherichia coli), Clostridium, Saccharomyces, or Pichia (e.g., Saccharomyces cerevisiae or Pichia pastoris), Kluyveromyces lactis, Salmonella typhimurium, Drosophila cells, Caenorhabditis elegans cells, African clawed frog cells, SF9 cells, C129 cells, 293 cells, Neurospora, and immortalized mammalian cell lines (e.g., Hela cells, myeloid cell lines, and lymphoid cell lines).

[0345] In some embodiments, cell is a prokaryotic cell, such as a bacterial cell, such as escherichia coli. In some embodiments, cell is a eukaryotic cell, such as a mammalian cell or a human cell. In some embodiments, cell is a primary eukaryotic cell, stem cell, tumor / cancer cell, circulating tumor cell (CTC), blood cell (for example, T cell, B cell, NK cell, Tregs etc.), hematopoietic stem cell, specialized immune cell (such as tumor infiltrating lymphocyte or tumor suppressor lymphocyte), stromal cell (such as cancer associated fibroblast etc.) in tumor microenvironment. In some embodiments, cell is brain or neuronal cell (for example, neuron, astrocyte, microglia, retinal ganglion cell, rod / cone cell etc.) of central or peripheral nervous system.

[0346] Target nucleic acid

[0347] The CRISPR-Cas system, composition or kit herein can be used to target one or more target nucleic acid molecules, such as target nucleic acid molecules present in biological samples, environmental samples (such as soil, air or water samples), etc. In some embodiments, the target nucleic acid is a coding RNA, such as pre-mRNA or mature mRNA. In some embodiments, the target nucleic acid is a nuclear RNA. In some embodiments, the target nucleic acid is an RNA transcript located in a eukaryotic cell nucleus. In some embodiments, the target nucleic acid is a non-coding RNA, such as a functional RNA, siRNA, microRNA, snRNA, snoRNA, piRNA, scaRNA, tRNA, rRNA, lncRNA or lincRNA.

[0348] In some embodiments, in addition to targeting target nucleic acid molecules, the CRISPR-Cas system, composition or kit herein performs one or more of the following functions on the target nucleic acid: cutting one or more target nucleic acid molecules or nicking one or more target nucleic acid molecules, activating or upregulating one or more target nucleic acid molecules, activating or inhibiting the translation of one or more target nucleic acid molecules, inactivating one or more target nucleic acid molecules, visualizing, labeling or detecting one or more target nucleic acid molecules, binding one or more target nucleic acid molecules, editing one or more target nucleic acid molecules, transporting one or more target nucleic acid molecules, and masking one or more target nucleic acid molecules. In some examples, the CRISPR-Cas system, composition or kit herein modifies one or more target nucleic acid molecules, and the modification of one or more target nucleic acid molecules includes one or more of the following: RNA base substitution, RNA base deletion, RNA base insertion, breakage of target nucleic acid, RNA methylation and RNA demethylation. In some embodiments, the CRISPR-Cas system, composition or kit herein can target one or more target nucleic acid molecules. In some embodiments, the CRISPR-Cas system, composition or kit herein can bind to one or more target nucleic acid molecules. In some embodiments, the CRISPR-Cas system, composition or kit herein can cut one or more target nucleic acid molecules. In some embodiments, the CRISPR-Cas system, composition, or kit herein can activate the translation of one or more target nucleic acid molecules. In some embodiments, the CRISPR-Cas system, composition, or kit herein can inhibit the translation of one or more target nucleic acid molecules. In some embodiments, the CRISPR-Cas system, composition, or kit herein can detect one or more target nucleic acid molecules. In some embodiments, the CRISPR-Cas system, composition, or kit herein can edit one or more target nucleic acid molecules.

[0349] In some embodiments, the target nucleic acid is AQp1 RNA. Knocking down Aqp1 RNA levels using the CRISPR-Cas system described herein can reduce aqueous humor production and lower intraocular pressure, and can be used to treat diseases such as glaucoma. In some embodiments, the target nucleic acid is AQp1 RNA, and the guide sequence of the guide polynucleotide is SEQ ID NO: 380.

[0350] In some embodiments, the target nucleic acid is PTBP1 RNA. Knocking down the level of PTBP1 RNA using the CRISPR-Cas system herein can promote the transdifferentiation of brain astrocytes into neurons, which can be used to treat diseases such as Parkinson's disease. In some embodiments, the target nucleic acid is PTBP1 RNA, and the guide sequence of the guide polynucleotide is SEQ ID NO: 381.

[0351] In some embodiments, the target nucleic acid is VEGFA RNA. Reducing the level of VEGFA RNA using the CRISPR-Cas system herein can prevent choroidal neovascularization, which can be used to treat diseases such as age-related macular degeneration. In some embodiments, the target nucleic acid is VEGFA RNA and the guide sequence of the guide polynucleotide is SEQ ID NO: 406.

[0352] In some embodiments, the target nucleic acid is AR RNA. The CRISPR-Cas system herein is used to knock down AR RNA levels, thereby downregulating androgen receptor expression, ultimately for treating androgenic alopecia. In some embodiments, the target nucleic acid is AR RNA, and the guide sequence of the guide polynucleotide is selected from SEQ ID NO: 398.

[0353] In some embodiments, the target nucleic acid is ANGPTL3 RNA. Using the CRISPR-Cas system described herein to knock down ANGPTL3 RNA levels can reduce blood lipids such as low-density lipoprotein cholesterol (LDL-C), and can be used to treat atherosclerotic cardiovascular diseases such as hyperlipidemia and familial hypercholesterolemia. In some embodiments, the target nucleic acid is ANGPTL3 RNA.

[0354] In other embodiments, the CRISPR-Cas systems, compositions, and kits herein can also target other target nucleic acids for use in treating diseases. Targeting can be targeted knockdown, targeted knockout, single-base editing, homologous recombination, targeted enhancement of transcription, targeted inhibition of transcription, enhanced translation, inhibition of translation, and the like. These target nucleic acids and their corresponding diseases are shown in the following table:

[0355] Therapeutic applications

[0356] Another aspect of the present disclosure relates to a pharmaceutical composition comprising a CRISPR-Cas13 system herein, a Cas13 protein herein, a fusion protein herein, a guide polynucleotide herein, a nucleic acid herein, a vector system herein, a lipid nanoparticle herein, a lentiviral vector herein, a ribonucleoprotein complex herein, a virus-like particle herein, or a eukaryotic cell herein. The pharmaceutical composition may comprise, for example, an AAV vector encoding the Cas13 protein or fusion protein herein and a guide polynucleotide. The pharmaceutical composition may comprise, for example, a lipid nanoparticle comprising an mRNA encoding a guide polynucleotide herein and a Cas13 protein or fusion protein. The pharmaceutical composition may comprise, for example, a lentiviral vector comprising an mRNA encoding a guide polynucleotide herein and a Cas13 protein or fusion protein. The pharmaceutical composition may comprise, for example: a virus-like particle comprising a guide polynucleotide herein and a Cas13 protein or fusion protein; or a ribonucleoprotein complex formed by a guide polynucleotide and a Cas13 protein or fusion protein.

[0357] Another aspect of the present disclosure relates to the use of the CRISPR-Cas13 system herein, the Cas13 protein herein, the fusion protein herein, the guide polynucleotide herein, the nucleic acid herein, the vector system herein, the lipid nanoparticles herein, the lentiviral vector herein, the ribonucleoprotein complex herein, the virus-like particle herein, or the eukaryotic cell herein to cut or edit a target nucleic acid in a mammalian cell.

[0358] Another aspect of the present disclosure relates to the use of the CRISPR-Cas13 system herein, the Cas13 protein herein, the fusion protein herein, the guide polynucleotide herein, the nucleic acid herein, the vector system herein, the lipid nanoparticle herein, the lentiviral vector herein, the ribonucleoprotein complex herein, the virus-like particle herein, or the eukaryotic cell herein for any of the following: cutting or nicking one or more target nucleic acid molecules, activating or upregulating one or more target nucleic acid molecules, activating or inhibiting translation of one or more target nucleic acid molecules, inactivating one or more target nucleic acid molecules, visualizing, labeling or detecting one or more target nucleic acid molecules, binding one or more target nucleic acid molecules, transporting one or more target nucleic acid molecules, and masking one or more target nucleic acid molecules.

[0359] Another aspect of the present disclosure relates to the use of the CRISPR-Cas13 system herein, the Cas13 protein herein, the fusion protein herein, the guide polynucleotide herein, the nucleic acid herein, the vector system herein, the lipid nanoparticles herein, the lentiviral vector herein, the ribonucleoprotein complex herein, the virus-like particle herein, or the eukaryotic cell herein to modify one or more target nucleic acid molecules in mammalian cells, wherein the modification of the one or more target nucleic acid molecules includes one or more of the following: RNA base substitution, RNA base deletion, RNA base insertion, fragmentation of the target nucleic acid, RNA methylation and RNA demethylation.

[0360] Another aspect of the present disclosure relates to the use of the CRISPR-Cas13 system herein, the Cas13 protein herein, the fusion protein herein, the guidance polynucleotide herein, the nucleic acid herein, the vector system herein, the lipid nanoparticles herein, the lentiviral vector herein, the ribonucleoprotein complex herein, the virus-like particles herein, or the eukaryotic cells herein in diagnosing, treating, or preventing a disease or condition associated with a target nucleic acid. In some embodiments, the disease or condition is Parkinson's disease. In some embodiments, the disease or condition is Parkinson's disease, and the target nucleic acid is PTBP1 RNA. In some embodiments, the disease or condition is glaucoma. In some embodiments, the disease or condition is glaucoma, and the target nucleic acid is AQp1 RNA. In some embodiments, the disease or condition is amyotrophic lateral sclerosis. In some embodiments, the disease or condition is amyotrophic lateral sclerosis, and the target nucleic acid is superoxide dismutase 1 (SOD1) RNA. In some embodiments, the disease or condition is age-related macular degeneration, and the target nucleic acid is VEGFA RNA. In some embodiments, the disease or condition is age-related macular degeneration, and the target nucleic acid is VEGFA RNA or VEGFR1 RNA. In some embodiments, the disease or condition is elevated plasma LDL cholesterol levels. In some embodiments, the disease or condition is elevated plasma LDL cholesterol levels, and the target nucleic acid is PCSK9 RNA or ANGPTL3 RNA.

[0361] Another aspect of the present disclosure relates to the use of the CRISPR-Cas13 system herein, the Cas13 protein herein, the fusion protein herein, the guide polynucleotide herein, the nucleic acid herein, the vector system herein, the lipid nanoparticles herein, the lentiviral vector herein, the ribonucleoprotein complex herein, the virus-like particle herein, or the eukaryotic cell herein in the preparation of a medicament for diagnosing, treating or preventing a disease or condition associated with a target nucleic acid. In some embodiments, the disease or condition is Parkinson's disease. In some embodiments, the disease or condition is glaucoma. In some embodiments, the disease or condition is amyotrophic lateral sclerosis. In some embodiments, the disease or condition is age-related macular degeneration. In some embodiments, the disease or condition is elevated plasma LDL cholesterol levels. In some embodiments, the disease or condition is Parkinson's disease, and the target nucleic acid is PTBP1 RNA. In some embodiments, the disease or condition is glaucoma, and the target nucleic acid is AQp1 RNA. In some embodiments, the disease or condition is amyotrophic lateral sclerosis, and the target nucleic acid is superoxide dismutase 1 (SOD1) RNA. In some embodiments, the disease or condition is age-related macular degeneration, and the target nucleic acid is VEGFA RNA or VEGFR1 RNA. In some embodiments, the disease or condition is elevated plasma LDL cholesterol levels, and the target nucleic acid is PCSK9 RNA or ANGPTL3 RNA. In some embodiments, the disease or condition is androgenetic alopecia (AGA, also known as male pattern baldness or premature balding), and the target nucleic acid is AR RNA.

[0362] In some embodiments, the pharmaceutical composition is delivered to a human subject in vivo. The pharmaceutical composition can be delivered by any effective route. Exemplary routes of administration include, but are not limited to, intravenous infusion, intravenous injection, intraperitoneal injection, intramuscular injection, intratumoral injection, subcutaneous injection, intradermal injection, intraventricular injection, intravascular injection, intracerebellar injection, intraocular injection, subretinal injection, intravitreal injection, intracameral injection, intratympanic injection, intranasal administration, and inhalation.

[0363] In some embodiments, the method for targeting RNA results in editing the sequence of the target nucleic acid. For example, by using the Cas13 protein or fusion protein with non-mutated HEPN domains and the guidance polynucleotide comprising a guide sequence specific to the target nucleic acid, the target nucleic acid can be cut at a precise position or the target nucleic acid is formed into a nick (nick, for example, when the target nucleic acid exists as a double-stranded nucleic acid molecule, any single strand is cut). In some instances, this method is used to reduce the expression of the target nucleic acid, which will reduce the translation of the corresponding protein. This method can be used in cells that do not need to increase RNA expression. In one example, RNA is related to diseases such as cystic fibrosis, Huntington's disease, Tay-Sachs, fragile X syndrome, fragile X-related tremor / ataxia syndrome, muscular dystrophy, myotonic dystrophy, spinal muscular atrophy, spinocerebellar ataxia, age-related macular degeneration, or familial ALS. In another example, RNA is related to cancer (for example, lung cancer, breast cancer, colon cancer, liver cancer, pancreatic cancer, prostate cancer, bone cancer, brain cancer, skin cancer (for example, melanoma) or kidney cancer). Examples of target nucleic acids include, but are not limited to, those associated with cancer (e.g., PD-L1, BCR-ABL, Ras, Raf, p53, BRCA1, BRCA2, CXCR4, β-catenin, HER2, and CDK4). Editing such target nucleic acids can produce therapeutic effects.

[0364] In some embodiments, RNA is expressed in immune cells. For example, the target nucleic acid can encode a protein that causes suppression of a desired immune response (e.g., tumor infiltration). Knocking down this RNA can promote the immune response (e.g., PD1, CTLA4, LAG3, TIM3) required. In another example, the target nucleic acid encodes a protein that causes undesirable immune response activation, such as in the case of autoimmune diseases such as multiple sclerosis, Crohn's disease, lupus, or rheumatoid arthritis.

[0365] Diagnostic applications

[0366] Another aspect of the present disclosure relates to an in vitro composition comprising a CRISPR-Cas13 system described herein and a labeled detector RNA that is incapable of hybridizing to a guide polynucleotide described herein.

[0367] Another aspect of the present disclosure relates to the use of the CRISPR-Cas13 system described herein for detecting a target nucleic acid in a nucleic acid sample suspected of containing the target nucleic acid.

[0368] In some embodiments, the method for detecting a target nucleic acid comprises a Cas13 protein or fusion protein fused to a fluorescent protein or other detectable label and a guide polynucleotide comprising a guide sequence specific for the target nucleic acid. The binding of the Cas13 protein or fusion protein to the target nucleic acid can be visualized by a microscope or other imaging method. In another example, an RNA aptamer sequence can be attached to or inserted into a guide polynucleotide, such as MS2, PP7, Qβ and other aptamers. Introducing proteins that specifically bind to these aptamers, such as MS2 phage coat proteins fused to fluorescent proteins or other detectable labels can be used to detect target nucleic acids because the Cas13-guide-target nucleic acid complex will be labeled by aptamer interactions.

[0369] In some embodiments, the method for detecting a target nucleic acid in a cell-free system results in the production of a detectable label or enzyme activity. For example, by using a Cas13 protein, a guide polynucleotide comprising a guide sequence specific for the target nucleic acid, and a detectable label, the target nucleic acid will be recognized by Cas13. The binding of Cas13 to the target nucleic acid triggers its RNase activity, which results in the cleavage of the target nucleic acid and the detectable label.

[0370] In some embodiments, the detectable label is an RNA connected to a fluorescent probe and a quencher. Complete detectable RNA connects a fluorescent probe and a quencher to suppress fluorescence. After the detectable RNA is cut by Cas13, the fluorescent probe is released from the quencher and shows fluorescent activity. This method can be used to determine whether the target nucleic acid is present in a cracked cell sample, a cracked tissue sample, a blood sample, a saliva sample, an environmental sample (such as water, soil or air sample) or other cracked cells or cell-free samples. This method can also be used to detect pathogens, such as viruses or bacteria, or to diagnose disease states, such as cancer.

[0371] In some embodiments, the detection of target nucleic acids helps diagnose diseases and / or pathological conditions, or the presence of viral or bacterial infections. For example, Cas13-mediated detection of non-coding RNAs such as PCA3, if detected in patient urine, can be used to diagnose prostate cancer. In another example, Cas13-mediated detection of lncRNA-AA174084, a biomarker for gastric cancer, can be used to diagnose gastric cancer.

[0372] Example 1 Protein Screening

[0373] 1. CRISPR and gene annotation

[0374] The software was used to predict the proteins of the whole genome of microorganisms from the NCBI Genbank and CNGB (China National Gene Bank) databases, and then the CRISPR array on the genome was predicted by the software for analysis and annotation.

[0375] 2. Acquisition of CRISPR-related proteins

[0376] Protein sequences within 10 kb upstream and downstream of the CRISPR array were aligned with known Cas13 sequences, and proteins with an e-value greater than 1e-5 were filtered out. Proteins with high similarity were then compared with the NCBI NR library and the EBI patent library, and clustering was used to remove redundant proteins. Finally, candidate proteins were selected. Experimental verification revealed over 100 proteins with amino acid sequences shown in SEQ ID NOs: 1-184 and 426-428.

[0377] The guide polynucleotide (gRNA) that forms a CRISPR complex with the Cas protein mainly consists of a guide sequence and a direct repeat (DR) sequence or a scaffold sequence. The guide sequence can be located at the 3' end or the 5' end of the DR sequence.

[0378] With reference to Table 1, the nucleotide sequences of the DR sequences corresponding to the proteins are shown in SEQ ID NOs: 185-368, or the reverse complementary sequences of the sequences shown in SEQ ID NOs: 185-368. In the following Examples (Examples 3-6), unless otherwise specified, the DR sequences corresponding to the proteins use the sequences shown in SEQ ID NOs: 185-368 in Table 1, and the positional relationship between the guide sequence and the DR sequence is as shown in Table 1; however, if otherwise specified in Examples 3-6, Examples 3-6 shall prevail.

[0379] Table 1

[0380] Among the above-mentioned Cas13 proteins, several are small in size and have high editing activity, for example: C13-84 (914aa), C13-129 (520aa), C13-130 (520aa), C13-134 (749aa), C13-151 (732aa), C13-188 (953aa), C13-196 (346aa), and C13-198 (336aa).

[0381] Example 2 Preparation of protein

[0382] This example uses C13-82 as an example to demonstrate the expression, separation, and purification processes of the above proteins.

[0383] (1) Vector construction

[0384] 1. Take the pET28a vector plasmid, double-digest it with BamHI and XhoI, and then recover the linearized vector by agarose gel electrophoresis. Insert the DNA fragment containing the protein coding sequence (protein encoding and nuclear localization signal) into the cloning region of the pET28a vector by homologous recombination. Transform the reaction solution into Stbl3 competent cells, spread on LB plates containing kanamycin sulfate, and culture overnight at 37°C. Then, pick clones for sequencing and identification.

[0385] The constructed recombinant vector was named C13-82-pET28a (SEQ ID NO: 375), which was used to express the C13-82 recombinant protein. The C13-82 recombinant protein structure was Histag-NLS-Cas13-SV40NLS-nucleoplasmin NLS.

[0386] 2. Culture the positive clones with the correct sequence overnight, extract the plasmid and transform it into the expression strain RIPL-BL21 (DE3), spread it on LB plates containing kanamycin sulfate, and culture it at 37°C overnight.

[0387] (2) Protein expression

[0388] 1. Pick a single clone and inoculate it into 5 ml of LB culture medium containing kanamycin sulfate, and culture it at 37℃ overnight.

[0389] 2. Inoculate the cells into 500 ml of TB culture medium containing kanamycin sulfate at a volume ratio of 1:100, culture at 220 rpm and 37°C until the OD value reaches 0.6, add IPTG to a final concentration of 0.2 mM, and induce at 16°C for 24 h.

[0390] 3. Collect the cells by centrifugation. Rinse the cells with 15 ml of PBS and then collect the cells by centrifugation. Add lysis buffer and ultrasonically disrupt the cells. Centrifuge at 10,000 g for 30 min to obtain the supernatant containing the recombinant protein. Filter the supernatant through a 0.45 μm filter membrane and then apply it to the column for purification.

[0391] (3) Protein purification

[0392] The product was obtained by purification through MAC (Ni Sepharose 6 Fast Flow, CYTIVA) and HITRAP HEPARIN HP (CYTIVA).

[0393] Example 3 Verification of AQp1 and PTBP1 gene editing efficiency

[0394] This example uses C13-82 as an example to demonstrate the process of verifying the gene editing efficiency of each of the proteins described above.

[0395] 1. Construction of vectors targeting AQp1 and PTBP1

[0396] The expression vector C13-82-BsaI plasmid was synthesized, and its sequence is shown in SEQ ID NO: 378.

[0397] The target nucleic acids selected for the experiment were AQp1 (Aquaporin 1) and PTBP1 (Polypyrimidine Tract Binding Protein 1). AQp1 was verified using the 293T cell line that highly expresses AQp1, and PTBP1 was verified using the 293T cell line.

[0398] Method for constructing a 293T cell line overexpressing AQp1 (293T-AQp1 cells): A vector overexpressing the AQp1 and EGFP genes, Lv-AQp1-T2a-GFP (SEQ ID NO: 379), was constructed. AQp1 and EGFP were separated by a 2A peptide. The Lv-AQp1-T2a-GFP plasmid was packaged into lentiviral vectors and transduced into 293T cells to establish a cell line stably overexpressing the AQp1 gene.

[0399] The gRNA guide sequence targeting AQp1 is:

[0400] The gRNA guide sequence targeting PTBP1 is:

[0401] Primer annealing was used to obtain a fragment targeting the target site (i.e., a fragment containing the gRNA guide sequence). The primers are as follows:

[0402] Targeting PTBP1

[0403] Upstream primer: 5'-caccGTGGTTGGAGAACTGGATGTAGATGGGCTG-3' (SEQ ID NO: 382) Downstream primer: 5'-AAAACAGCCCATCTACATCCAGTTCTCCAACCAC-3' (SEQ ID NO: 383)

[0404] Targeting AQp1

[0405] Upstream primer: 5'-CACCAGGGCAGAACCGATGCTGATGAAGAC-3' (SEQ ID NO: 384)

[0406] Downstream primer: 5'-AAAAGTCTTCATCAGCATCGGTTCTGCCCT-3' (SEQ ID NO: 385)

[0407] The primer annealing reaction system is shown in the following table:

[0408] Oligo-F (10 μM) 2 μl, Oligo-R (10 μM) 2 μl, 10× endonuclease reaction buffer* 2 μl, deionized water up to 20 μl.

[0409] Incubate at 95°C in a PCR instrument for 5 minutes, then immediately remove and incubate on ice for 5 minutes to allow the primers to anneal to each other to form double-stranded DNA with sticky ends.

[0410] The synthesized C13-82-BsaI vector plasmid was digested with BsaI endonuclease. The annealed product and the purified backbone recovered after digestion were ligated using T4 DNA ligase. After transformation into E. coli, positive clones were selected and plasmids were extracted to obtain the C13-82 vector (C13-82 targeting AQP1 or PTBP1), which was used in the following C13-82 experimental group. The C13-82 vector construct is CMV-C13-2-U6-gRNA, which can be used to express the C13-82 protein and the gRNA targeting AQp1 or PTBP1.

[0411] The following control vectors were prepared using conventional methods:

[0412] The sequence of the CasRx-AQp1 plasmid (a positive control vector for CasRx targeting AQp1) is shown in SEQ ID NO: 386, and the plasmid structure is CMV-CasRx-U6-gRNA, which contains a sequence encoding CasRx (amino acid sequence shown in SEQ ID NO: 387).

[0413] The sequence of the CasRx-PTBP1 plasmid (positive control vector for CasRx targeting PTBP1) is shown in SEQ ID NO: 388. The difference from the CasRx-AQp1 plasmid is that the gRNA guide sequence targeting AQp1 is replaced with the gRNA guide sequence targeting PTBP1.

[0414] The sequence of the CasRx-blank plasmid (blank control vector, which can express CasRx and gRNA, but the gRNA does not target AQp1 and PTBP1) is shown in SEQ ID NO: 389, and the plasmid structure is CMV-CasRx-U6-gRNA.

[0415] 2. Transfect 293T cells and 293T-AQp1 cells with the vector to be verified

[0416] 293T-AQp1 cells were transfected with 500 ng of the control plasmid or the vector plasmid targeting AQP1 in a 24-well plate. 293T cells were transfected with 500 ng of the control plasmid or the vector plasmid targeting PTBP1 in a 24-well plate.

[0417] The transfection method is as follows:

[0418] 1. Digest the cells with trypsin (0.25%, EDTA, Thermo), count the cells, and add 2×10 5 Cells were plated in 24-well plates.

[0419] 2. For each transfection sample, prepare the complexes as follows:

[0420] a. Add 50 μL of serum-free Opti-MEMI (Thermo) reduced serum medium to each well of the 24-well plate containing cells and dilute the aforementioned plasmid DNA and mix gently.

[0421] b. Gently mix Lipofectamine 2000 (Thermo) before use, then dilute 1.8 μL of Lipofectamine 2000 in 50 μL of Opti-MEMI medium per well. Incubate at room temperature for 5 minutes. Note: Proceed to step c within 25 minutes.

[0422] c. After a 5-minute incubation, combine the diluted DNA with the diluted Lipofectamine 2000. Mix gently and incubate at room temperature for 20 minutes (the solution may appear turbid). Note: The complex is stable at room temperature for 6 hours.

[0423] The complexes were added to the cells and mixed.

[0424] 3. qPCR detection of target gene RNA

[0425] 48 hours after transfection, RNA was extracted from cells using the Steady Pure Universal RNA Extraction Kit AG21017, and RNA concentration was measured using an ultra-micro-spectrophotometer. RNA products were reverse transcribed using the Evo M-MLV Mix Kit with gDNA Clean for qPCR Reverse Transcription Kit and detected using the SYBR Green Premix ProTaq HS qPCR Kit.

[0426] The primers used in qPCR are as follows:

[0427] Detection of PTBP1:

[0428] Upstream primer: 5'-ATTGTCCCAGATATAGCCGTTG-3' (SEQ ID NO: 390)

[0429] Downstream primer: 5'-GCTGTCATTTCCGTTTGCTG-3' (SEQ ID NO: 391)

[0430] Detection of AQp1:

[0431] Upstream primer: 5'-GCTCTTCTGGAGGGCAGTGG-3' (SEQ ID NO: 392)

[0432] Downstream primer: 5'-CAGTGTGACAGCCGGGTTGAG-3' (SEQ ID NO: 393)

[0433] Detection of internal reference GAPDH:

[0434] Upstream primer: 5'-CCATGGGGAAGGTGAAGGTC-3' (SEQ ID NO: 394)

[0435] Downstream primer: 5'-GAAGGGGTCATTGATGGCAAC-3' (SEQ ID NO: 395)

[0436] The reaction system was prepared according to the instructions of SYBR Green Premix ProTaq HS qPCR Kit and the PCR products were analyzed using Quant Studio TM 5. Detection was performed using the Real-Time PCR System.

[0437] This experiment used the relative quantification method, 2 -△△Ct Calculate target RNA.

[0438] The RNA amounts of AQp1 and PTBP1 were calculated using the above calculation method. The results of the validation experiment targeting PTBP1 are shown in Table 4. The results of the validation experiment targeting AQp1 are shown in Table 5.

[0439] Table 4. PTBP1 RNA knockdown test results

[0440] Note: Different batches represent three independent biological replicates performed at the same or different time periods (transfection operations used cells from the same batch); CasRx-blank is a negative control that does not target PTBP1; CasRx-PTBP1 is a positive control that targets PTBP1; 293T-NC refers to a blank control of 293T cells that were not transfected with the plasmid.

[0441] Table 5. AQp1 RNA knockdown test results

[0442] Note: Different batches represent three independent biological replicates performed at the same or different time periods (transfection procedures used the same batch of cells); “-” indicates not tested; CasRx-blank is a negative control that does not target AQp1; CasRx-AQp1 is a positive control that targets AQp1.

[0443] The above experimental results show that multiple proteins showed obvious editing effects after combining with gRNA targeting AQp1 or PTBP1.

[0444] Example 4 AR gene editing efficiency verification

[0445] This example uses C13-134 as an example to demonstrate the process of verifying the AR gene editing efficiency of each protein.

[0446] 1. Construction of a validation vector targeting the endogenous gene AR

[0447] The C13-82 in Example 3 was replaced with C13-134 to synthesize the expression vector C13-134-BsaI plasmid.

[0448] The target nucleic acid selected for the experiment was AR (androgen receptor), and the 293T cell line was used to verify AR.

[0449] The gRNA guide sequence targeting AR is:

[0450] Primer annealing was used to obtain fragments targeting the target site. The primers are as follows:

[0451] Upstream primer: 5'-GTCAAGTACTGAATGACAGCCATCTGGTC-3' (SEQ ID NO: 399)

[0452] Downstream primer: 5'-AAAAGACCAGATGGCTGTCATTCAGTACT-3' (SEQ ID NO: 400)

[0453] The primer annealing reaction system is as follows:

[0454] Oligo-F (10 μM) 2 μl,

[0455] Oligo-R (10 μM) 2 μl,

[0456] 10× endonuclease reaction buffer*2μl,

[0457] Deionized water up to 20μl.

[0458] Incubate at 95°C in a PCR instrument for 5 minutes, then immediately remove and incubate on ice for 5 minutes to allow the primers to anneal to each other to form double-stranded DNA with sticky ends.

[0459] The C13-134-BsaI plasmid was digested with BsaI endonuclease, and the backbone was purified and recovered. The annealed product was ligated with T4 ligase to obtain the verification vector plasmid (CMV-C13-134-U6-gRNA). After transformation into E. coli, positive clones were selected and plasmids were extracted for subsequent experiments.

[0460] The following control vectors were prepared using conventional methods:

[0461] The sequence of the CasRx-AR plasmid (positive control vector for CasRx targeting AR) is shown in SEQ ID NO: 401, and the plasmid structure is CMV-CasRx-U6-gRNA. It contains a sequence encoding CasRx (amino acid sequence shown in SEQ ID NO: 387).

[0462] The sequence of the CasRx-blank plasmid (blank control vector, which can express CasRx and gRNA, but the gRNA does not target AR) is shown in SEQ ID NO: 389, and the plasmid structure is CMV-CasRx-U6-gRNA.

[0463] 2. Transfect 293T cells with the vector to be verified

[0464] 293T cells were transfected into 24-well plates with 500 ng of control plasmid or AR-targeting vector plasmid.

[0465] The transfection method was carried out according to the instructions of Lipofectamine 2000 (Thermo) to transfect 24-well plates.

[0466] 3. qPCR detection

[0467] 72 hours after transfection, RNA was extracted from cells using the Steady Pure Universal RNA Extraction Kit AG21017, and RNA concentration was determined using an ultra-micro-spectrophotometer. RNA products were reverse transcribed using the EvoM-MLV Mix Kit with gDNA Clean for qPCR AG11728 Reverse Transcription Kit, and the reverse transcription products were detected using the SYBR Green Premix ProTaq HSqPCR Kit (Low Rox Plus) qPCR Kit.

[0468] The primers used in qPCR are as follows:

[0469] Detecting AR:

[0470] Detection of internal reference GAPDH:

[0471] according to Green Premix ProTaq HS qPCR Kit (Rox Plus) instructions for preparing the reaction system and using Quant Studio TM 5. Detection was performed using the Real-Time PCR System.

[0472] This experiment used the relative quantification method, 2 -△△Ct The target RNA level was calculated by the method. The results are shown in Table 6 below.

[0473] Table 6. qPCR detection of edited AR RNA levels

[0474] Note: Different batches represent three independent biological replicates performed in the same time period or at different time periods (the same batch of cells was used for transfection); CasRx-blank is a negative control that does not target AR; CasRx-AR plasmid is a positive control that targets AR; 293T-NC refers to a blank control of 293T cells that were not transfected with the plasmid; DRrc in C13-134-DRrc-AR represents that the DR sequence used is the reverse complement of the sequence shown in SEQ ID NO: 289 in Table 1, and DRbrc in C13-134-DRbrc-VEGFA represents that the DR sequence used is the reverse complement of the sequence shown in SEQ ID NO: 289 and the guide sequence is located at the 5' end of the DR sequence.

[0475] The above experimental results show that after combining with gRNA targeting AR, multiple proteins can downregulate AR RNA expression levels.

[0476] Example 5 Verification of VEGFA gene editing efficiency

[0477] This example uses C13-134 as an example to demonstrate the process of verifying the VEGFA gene editing efficiency of each protein.

[0478] 1. Construction of a validation vector targeting the endogenous gene VEGFA

[0479] The expression vector C13-134-BsaI plasmid is provided (see Example 4).

[0480] The target nucleic acid selected for the experiment was VEGFA (Vascular Endothelial Growth Factors A), and 293T cell line was used for VEGFA validation.

[0481] The gRNA guide sequence targeting VEGFA is:

[0482] Primer annealing was used to obtain fragments targeting the target site. The primer information used is as follows:

[0483] Upstream primer: 5'-GTCATGGGTGCAGCCTGGGACCACTTGGCATGG-3' (SEQ ID NO: 407)

[0484] Downstream primer: 5'-AAAACCATGCCAAGTGGTCCCAGGCTGCACCCA-3' (SEQ ID NO: 408).

[0485] The primer annealing reaction system is as follows:

[0486] Oligo-F (10 μM) 2 μl,

[0487] Oligo-R (10 μM) 2 μl,

[0488] 10× endonuclease reaction buffer*2μl,

[0489] Deionized water up to 20μl.

[0490] Incubate at 95°C in a PCR instrument for 5 minutes, then immediately remove and incubate on ice for 5 minutes to allow the primers to anneal to each other to form double-stranded DNA with sticky ends.

[0491] The C13-134-BsaI plasmid was digested with BsaI endonuclease, and the backbone was purified and recovered. The annealed product was ligated with T4 ligase to obtain the verification vector plasmid (CMV-C13-134-U6-gRNA). After transformation into E. coli, positive clones were selected and plasmids were extracted for subsequent experiments.

[0492] The following control vectors were prepared using conventional methods:

[0493] The sequence of the CasRx-VEGFA plasmid (a positive control vector for CasRx targeting VEGFA) is shown in SEQ ID NO: 409, and the plasmid structure is CMV-CasRx-U6-gRNA. It contains a sequence encoding CasRx (amino acid sequence shown in SEQ ID NO: 387).

[0494] The sequence of the CasRx-blank plasmid (blank control vector, which can express CasRx and gRNA, but the gRNA does not target VEGFA) is shown in SEQ ID NO: 389, and the plasmid structure is CMV-CasRx-U6-gRNA.

[0495] 2. Transfect 293T cells with the vector to be verified

[0496] 293T cells were transfected into 24-well plates with 500 ng of control plasmid or VEGFA-targeted vector plasmid.

[0497] The transfection method was carried out according to the instructions of Lipofectamine 2000 (Thermo) to transfect 24-well plates.

[0498] 3. qPCR detection

[0499] 72 hours after transfection, RNA was extracted from cells using the Steady Pure Universal RNA Extraction Kit AG21017, and RNA concentration was determined using an ultra-micro-spectrophotometer. RNA products were reverse transcribed using the Evo M-MLV Mix Kit with gDNA Clean for qPCR AG11728 Reverse Transcription Kit, and the reverse transcription products were detected using the SYBR Green Premix Pro Taq HS qPCR Kit (LowRoxPlus).

[0500] The primers used in qPCR are as follows:

[0501] Detection of VEGFA:

[0502] Detection of internal reference GAPDH:

[0503] according to Green Premix Pro Taq HS qPCR Kit (Rox Plus) instructions for preparing the reaction system and using Quant Studio TM 5. Detection was performed using the Real-Time PCR System.

[0504] This experiment used the relative quantification method, 2 -△△Ct The target RNA level can be calculated by the following method:

[0505] ΔCt=Ct(VEGFA)-Ct(GAPDH)

[0506] △△Ct=△Ct(vector to be verified, such as C13-134-VEGFA)-△Ct(control C13-134-BsaI)

[0507] 2 -△△Ct =2^(-△△Ct)

[0508] The 2 of edited VEGFA RNA was calculated according to the above calculation method. -△△Ct The results are shown in Table 7 below.

[0509] Table 7. Target RNA levels after editing

[0510] Note: Different batches represent three independent biological replicates performed in the same time period or at different time periods (the same batch of cells was used for transfection); "-" indicates not tested; CasRx-blank is a negative control that does not target VEGFA; CasRx-VEGFA is a positive control that targets VEGFA; DRrc in C13-134-DRrc-AR represents that the DR sequence used is the reverse complementary sequence of the sequence shown in SEQ ID NO: 289 in Table 1, and DRbrc in C13-134-DRbrc-VEGFA represents that the DR sequence used is the reverse complementary sequence of the sequence shown in SEQ ID NO: 289 and the guide sequence is located at the 5' end of the DR sequence.

[0511] Example 6 Verification of EGFP gene editing efficiency

[0512] This example uses C13-67 as an example to demonstrate the process of verifying the EGFP gene editing efficiency of various proteins.

[0513] 1. Synthesize the EGFP-targeted vector to be verified and the control vector

[0514] The sequence of the synthetic exogenous EGFP expression vector is shown in SEQ ID NO: 413, and the plasmid structure is CMV-EGFP. The verification vector plasmid for synthesizing the C13-67 protein targeting EGFP is shown in SEQ ID NO: 414, and the plasmid structure is CMV-C13-67-U6-gRNA.

[0515] EGFP was used as an exogenous reporter gene, and its nucleic acid sequence (720 bp) is shown in SEQ ID NO: 415.

[0516] The guide sequence targeting EGFP is ugccguucuucugcuugucggccaugauau (SEQ ID NO: 416).

[0517] 2. Transfect 293T cells with the vector to be verified

[0518] The exogenous EGFP expression vector and the C13-67 protein targeting EGFP verification vector plasmid were transfected into 293T cells in a 24-well plate at a ratio of 1:2 (166ng:334ng).

[0519] The transfection method is as follows:

[0520] 1. Digest 293T cells with trypsin (0.25%, EDTA, Thermo, 25200056), count the cells, and add 2×10 5 Cells were plated in 24-well plates.

[0521] 2. For each transfection sample, prepare the complexes as follows:

[0522] a. Add 50 μL of serum-free Opti-MEMI (Thermo, 11058021) reduced serum medium to each well of the 24-well plate containing cells and gently mix.

[0523] b. Gently mix Lipofectamine 2000 (Thermo, 11668019) before use, then dilute 1 μL of Lipofectamine 2000 in 50 μL of Opti-MEMI medium per well. Incubate at room temperature for 5 minutes. Note: Proceed to step c within 25 minutes.

[0524] c. After a 5-minute incubation, combine the diluted DNA with diluted Lipofectamine 2000. Mix gently and incubate at room temperature for 20 minutes (the solution may appear turbid). Note: The complex is stable at room temperature for 6 hours. Add the complex to 293T cells and mix. Detect the complex by flow cytometry 48 hours later.

[0525] 3. Flow cytometry detection of the effect of each protein on downregulating EGFP expression

[0526] The cells and plasmids used are shown in Table 8 below:

[0527] Table 8. Transfected cell groups

[0528] Note: n represents the code of each protein.

[0529] 48 h after transfection, the cells were digested with trypsin (Trypsin 0.25%, EDTA, Thermo), and the supernatant was removed by centrifugation at 300g for 5 min. The cells in each well were resuspended in 500 μL of PBS, and EGFP fluorescence expression was detected by flow cytometry. After removing cell debris by FCS-A and SSC-A gating, the data were collected by flow cytometry.

[0530] Collect and record the Mean-FITC-A results of the FITC channel, and calculate the downregulation amplitude according to the following formula:

[0531] Let the GFP fluorescence of the EGFP group be a, and the GFP fluorescence of the other groups be x. Downregulation (%) = (ax) ÷ a × 100. The blank control group was not included in the comparison. The downregulation results are shown in Table 9. Each protein targeting EGFP downregulated EGFP expression to varying degrees, indicating effective RNA editing.

[0532] Table 9. GFP fluorescence detection results by flow cytometry

[0533] Note: Different batches represent three independent biological replicates performed in the same time period or at different time periods (the same batch of cells were used for transfection); "-" indicates not tested; C13-117-DRbrc-GFP indicates that the DR sequence used is the reverse complementary sequence of the sequence shown in SEQ ID NO: 300 in Table 1, and the guide sequence is located at the 5' end of the DR sequence; C13-132-DRrc-GFP indicates that the DR sequence used is the reverse complementary sequence of the sequence shown in SEQ ID NO: 301 in Table 1; C13-142-DRrc-GFP indicates that the DR sequence used is the reverse complementary sequence of the sequence shown in SEQ ID NO: 298 in Table 1; C13-143-DRb-GFP indicates that the guide sequence is located at the 5' end of the DR sequence; C13-159-DRbrc-GFP indicates that the DR sequence used is the reverse complementary sequence of the sequence shown in SEQ ID NO: 294 in Table 1, and the guide sequence is located at the 5' end of the DR sequence; C13-160-DRrc-GFP indicates that the DR sequence used is the reverse complementary sequence of the sequence shown in SEQ ID NO: NO:295 in Table 1; C13-160-DRb-GFP indicates that the guide sequence is located at the 5' end of the DR sequence; C13-160-DRbrc-GFP indicates that the DR sequence used is the reverse complementary sequence of the sequence shown in SEQ ID NO:295 in Table 1 and the guide sequence is located at the 5' end of the DR sequence; C13-181-DRrc-GFP indicates that the DR sequence used is the reverse complementary sequence of the sequence shown in SEQ ID NO:303 in Table 1.

[0534] Example 7 Identification of the consensus motif of Cas13 protein

[0535] In the above experiments on knockdown of AQp1, PTBP1, VEGFA, AR, and EGFP expression, different levels of knockdown were observed among the proteins. Further analysis of the protein isoforms of the 184 proteins identified above revealed that these proteins include multiple isoforms such as Cas13a, Cas13b, Cas13c, Cas13d, and Cas13x, as well as new isoforms.

[0536] 1. The crystal structure of RcCas13a reported in the reference (Leonhard M. Kick, et al. "Structure and mechanism of the RNA dependent RNase Cas13a from Rhodobacter capsulatus." Commun Biol. 2022 Jan 20; 5 (1): 71) discloses the amino acid residues that interact with crRNA in the RcCas13a protein, as well as the catalytic residues of the HEPN domain of the RcCas13a protein. A multiple sequence alignment of Cas13a protein and RcCas13a was performed (online MAFFT v7.504, E-INS-i algorithm, other default parameters) to identify the common motif of Cas13a protein.

[0537] Among them, the Cas13a proteins that were subjected to multiple sequence alignment with RcCas13a include: C13-82, C13-84, C13-89, C13-67, C13-68, C13-71, C13-74, C13-76, C13-78, C13-79, C13-81, C13-86, C13-139, C13-169, C13-207, etc.

[0538] As shown in Table 10 and Figures 1A-1D, motifs 1-9 appear more frequently in the Cas13a listed above, and motifs 26-34 are further clarifications of motifs 1-9. Figures 1A-1D are partial alignment results of multiple sequence alignments of Cas13a proteins and RcCas13a in an embodiment of the present invention.

[0539] Table 10. Recognition motifs

[0540] Note: Single-letter codes represent highly conserved amino acid residues, x represents any amino acid, and [] indicates that the position is an optional amino acid code within [].

[0541] 2. The crystal structure of PbuCas13b reported in the reference (Ian M Slaymaker, et al. "High-Resolution Structure of Cas13b and Biochemical Characterization of RNA Targeting and Cleavage." Cell Rep. 2019 Mar 26; 26(13): 3741-3751.e5). A multiple sequence alignment of Cas13b protein and PbuCas13b was performed (online MAFFT v7.504, E-INS-i algorithm, other default parameters) to identify the common motif of Cas13b protein.

[0542] Among them, the Cas13b proteins that underwent multiple sequence alignment with PbuCas13b included: C13-6, C13-12, C13-65, C13-70, C13-90, C13-116, C13-94, C13-49, C13-28, etc.

[0543] The results are shown in Table 11 and Figures 2A-2C. Motifs 10-13 appear more frequently in the Cas13b listed above. Figures 2A-2C are partial alignment results of multiple sequence alignments of the Cas13b protein and PbuCas13b in an embodiment of the present invention.

[0544] Table 11. Recognition motifs

[0545] Note: Single-letter codes represent highly conserved amino acid residues, x represents any amino acid, and [] indicates that the position is an optional amino acid code within [].

[0546] 3. The crystal structure of EsiCas13d reported in the reference (Cheng Zhang, et al. "Structural basis for the RNA-guided ribonuclease activity of CRISPR-Cas13d" Cell. 175(1): 212–223). A multiple sequence alignment (online MAFFT v7.504, E-INS-i algorithm, other default parameters) was performed between Cas13d protein and EsiCas13d to identify the consensus motif of Cas13a protein.

[0547] Among them, the Cas13d proteins that were subjected to multiple sequence alignment with EsiCas13d include: C13-151, C13-188, C13-148, etc., which were shown to be active in the above examples, and other Cas13d (for example: C13-1, C13-3, C13-14, C13-15, C13-16, C13-18, C13-22, C13-26, C13-56, C13-57, C13-58, C13-59, C13-60, C13-61, C13-62 , C13-93, C13-120, C13-136, C13-152, C13-153, C13-154, C13-161, C13-162, C13-163, C13-164, C13-165, C13-1 70, C13-171, C13-172, C13-183, C13-184, C13-187, C13-190, C13-191, C13-192, C13-204, C13-205, C13-208, etc.).

[0548] The results are shown in Table 12 and Figures 3A-3B. Motifs 14-25 appear more frequently in the Cas13d listed above, and motifs 35-46 are further clarifications of motifs 14-25. Figures 3A-3B are partial alignment results of multiple sequence alignments of the Cas13d protein and EsiCas13d in the embodiments of the present invention.

[0549] Table 12. Key motifs identified

[0550] Note: Single-letter codes represent highly conserved amino acid residues, x represents any amino acid, and [] indicates that the position is an optional amino acid code within [].

[0551] Example 8 Verification of Endogenous FTH1 Gene Editing Efficiency

[0552] Construction of a validation vector targeting the endogenous gene FTH1

[0553] All verification vectors in this experimental example use the same plasmid backbone, with a structure of CMV-NLS-Cas13-NLS-U6-crRNA. All Cas13s carry one NLS at the N-terminus and two NLSs at the C-terminus, for a total of 3×NLS structures. The C13-130-FTH1 verification vector sequence is SEQ ID NO: 417, the C13-188-FTH1 verification vector sequence is SEQ ID NO: 418, and the negative control vectors C13-130-GFP and C13-188-GFP that do not target FTH1 have sequences of SEQ ID NO: 419 and SEQ ID NO: 420, respectively. Except for the Cas13 protein coding sequence and the coding sequence of the direct repeat sequence DR, the verification plasmid has the same structure. The corresponding Cas13 encoding nucleic acid sequence was obtained by human codon optimization through the website https: / / www.vectorbuilder.cn / tool / codon-optimization.html. The positive control targeting FTH1 was named CasRx-FTH1 (SEQ ID NO: 421), and the negative control not targeting FTH1 was named CasRx-blank (the non-targeting guide sequence used was tgccgttcttctgcttgtcggccatgatat [SEQ ID NO: 422]).

[0554] The gRNA guide sequence targeting FTH1 is:

[0555] Transfect 293T cells with the vector to be verified

[0556] 500 ng of the validation vector and control plasmid were transfected into 293T cells in 24-well plates.

[0557] The 24-well plate was transfected according to the instructions of Lipofectamine 2000 (Thermo, 11668019).

[0558] qPCR detection of RNA changes in target genes

[0559] 72 hours after transfection, RNA was extracted from cells using the SteadyPure Universal RNA Extraction Kit AG21017, and RNA concentration was measured using an ultra-micro-spectrophotometer. RNA products were reverse transcribed using the Evo M-MLV Mix Kit with gDNA Clean for qPCR Reverse Transcription Kit and detected using the SYBR Green Premix Pro Taq HS qPCR Kit (Low Rox Plus).

[0560] The primers used in qPCR are as follows:

[0561] Detection of FTH1: CCCCCATTTGTGTGACTTCAT (SEQ ID NO: 424)

[0562] Detection internal reference GAPDH: CCATGGGGAAGGTGAAGGTC (SEQ ID NO: 394)

[0563] according to Green Premix Pro Taq HS qPCR Kit (Rox Plus) instructions for preparing the reaction system and using QuantStudio TM 5. Real-Time PCR System, 384-well plate. Changes in target RNA were calculated using the relative quantification method (2-ΔΔCt).

[0564] The 2-△△Ct value of FTH1 calculated according to the above calculation method is shown in the following table:

[0565] Table 13. Editing efficiency of targeting FTH1

[0566] The negative controls of each Cas13 group were normalized to make the average value corrected to 1, and the calculated 2-△△Ct values ​​are shown in the following table:

[0567] Table 14. Normalized editing efficiency

[0568] The positive control CasRx had a significant editing effect on FTH1; C13-130 had a higher editing effect on FTH1; C13-188 had a significant editing effect on FTH1, which was equivalent to the activity of CasRx.

[0569] Example 9. Screening of Cas13 protein

[0570] 1. CRISPR and gene annotation

[0571] For microbial genomes from NCBI Gebank and CNGB (China National Gene Bank) databases, the proteins of the whole genome are predicted, and then the CRISPR array on the genome is predicted using software for analysis and annotation.

[0572] 2. Preliminary screening of proteins

[0573] Cluster analysis was used to remove redundant proteins, and proteins with amino acid sequence lengths less than 800 AA (amino acids) or greater than 1400 AA were filtered out.

[0574] 3. Acquisition of CRISPR-related proteins

[0575] Protein sequences within 10kb upstream and downstream of the CRISPR Array were compared with known Cas13 proteins, and proteins with an evalue greater than 1*e-5 were filtered out. Proteins with high similarity were then compared with the NCBI NR library and the EBI patent library, and candidate proteins were selected.

[0576] The inventors further screened and experimentally verified a number of candidate proteins, ultimately obtaining the C13-52 protein (SEQ ID NO: 426), the C13-55 protein (SEQ ID NO: 427), and the C13-88 protein (SEQ ID NO: 428). The C13-52 protein is also known as CasRfg.7. The C13-55 protein is also known as CasRfg.5. The C13-88 protein is also known as CasRfg.6. The genomic sequence sources of these proteins are shown in Table 15.

[0577] Table 15. Sources of genomic sequences of Cas13 proteins.

[0578] The direct repeat (DR) sequence corresponding to C13-52 is:

[0579] The direct repeat (DR) sequence corresponding to C13-55 is:

[0580] The direct repeat (DR) sequence corresponding to C13-88 is:

[0581] The RNA secondary structures of the above-mentioned direct repeat sequences predicted using RNA fold are shown in Figures 4 to 6.

[0582] Example 10. Verification of editing efficiency of Cas13 protein

[0583] 1. Construction of an editing vector targeting AQp1 RNA

[0584] The test target nucleic acid selected in this experiment was AQp1 (Aquaporin 1) RNA, and the 293T cell line that highly expresses AQp1 was used.

[0585] The method for constructing a 293T cell line that highly expresses AQp1 (293T-AQp1 cells) is as follows:

[0586] The AQp1 gene-overexpressing vector, Lv-AQp1-T2a-GFP (sequence shown in SEQ ID NO: 379), was constructed using conventional methods. This vector, based on a lentiviral vector, inserted the AQp1 and EGFP genes, separated by a 2A peptide (T2a). The Lv-AQp1-T2a-GFP plasmid was packaged into lentivirus and then transduced into 293T cells, establishing a cell line stably overexpressing the AQp1 gene.

[0587] In the CRISPR-Cas13 system, the guide sequence of the gRNA targeting AQp1 is:

[0588] The C13-55-BsaI plasmid vector (sequence such as SEQ ID NO: 433) and the C13-88-BsaI plasmid vector (sequence such as SEQ ID NO: 434) with a universal gRNA backbone expression frame were synthesized by an outsourcing company.

[0589] When selecting a positive control protein, CasRx was considered to be the Cas13 protein with the highest editing efficiency currently disclosed in the art. It was used as a control and the following control vector was prepared using conventional methods:

[0590] A negative control vector, CasRx-blank (SEQ ID NO: 389), is used to express CasRx and a gRNA that does not target any gene in eukaryotic cells;

[0591] The positive control vector CasRx-AQp1 (SEQ ID NO: 386) is used to express CasRx and gRNA targeting AQp1 RNA.

[0592] Primer annealing was used to obtain fragments targeting the AQp1 target site. The primers are as follows:

[0593] C13-55-AQp1 corresponding primers:

[0594] C13-88-AQp1 corresponding primers:

[0595] The primer annealing reaction system is as follows:

[0596] Incubate at 95°C in a PCR instrument for 5 minutes, then immediately remove and incubate on ice for 5 minutes to allow the primers to anneal to each other to form double-stranded DNA with sticky ends.

[0597] The synthesized C13-55-BsaI and C13-88-BsaI plasmids were digested with an endonuclease (Bsa I). ​​The purified backbone was then ligated with the annealed product using a T4 ligation. After transformation into E. coli, positive clones were selected and the plasmids were extracted. The resulting C13-55-AQp1 and C13-88-AQp1 vectors, constructed with a CMV-Cas13-U6-gRNA architecture, express the C13-55 and C13-88 proteins (with an NLS-Cas13-SV40 NLS-nucleoplasmin NLS-HA architecture) and a gRNA targeting AQp1 containing the corresponding DR sequence.

[0598] 2. Transfect 293T-AQp1 cells with the vector to be verified

[0599] According to the instructions of Lipofectamine 2000 (Thermo), 500 ng of C13-55-AQp1, C13-88-AQp1, and control plasmids were transfected into 293T-AQp1 cells in 24-well plates. 293T-AQp1 cells without plasmid transfection served as blank controls.

[0600] 3. qPCR detection of RNA changes in target genes

[0601] 72 hours after transfection, RNA was extracted from cells using the SteadyPure Universal RNA Extraction Kit, and RNA concentration was measured using an ultra-micro-spectrophotometer. RNA products were reverse transcribed using the Evo M-MLV Mix Kit with gDNA Clean for qPCR Reverse Transcription Kit and detected using the SYBR Green Premix Pro Taq HS qPCR Kit (Low Rox Plus).

[0602] The primers used in qPCR are as follows:

[0603] Detection of AQp1: gctcttctggagggcagtgg (SEQ ID NO: 392)

[0604] Detection internal reference GAPDH: CCATGGGGAAGGTGAAGGTC (SEQ ID NO: 394)

[0605] according to Green Premix Pro Taq HS qPCR Kit (Rox Plus) instructions for preparing the reaction system and using QuantStudio TM 5. Detection was performed using the Real-Time PCR System.

[0606] The relative quantification method, 2-ΔΔCt method, was used to calculate the changes in target RNA.

[0607] The experiment was repeated three times independently, and the results were averaged. The results are shown in Table 16 and Figure 7.

[0608] Table 16. Results of knockdown test of AQp1 RNA by Cas13

[0609] The above experimental results showed that both C13-55 and C13-88 had a significant knockdown of AQp1 RNA (P<0.05), with editing efficiencies as high as 97% and 78%, respectively, and the editing efficiency of C13-55 protein was close to that of CasRx.

[0610] Example 11. Comparison with known Cas13 tools

[0611] 1. Construction of verification vector and control vector

[0612] When selecting the positive control protein, CasRx was considered to be the Cas13 protein with the highest editing efficiency currently disclosed in the field. At the same time, more reported Cas13 proteins such as PspCas13b, Cas13X.1, and Cas13Y.1 were set as controls.

[0613] The verification vectors targeting PTBP1 and VEGFA of CasRx, PspCas13b, Cas13X.1, Cas13Y.1, C13-52, C13-55, C13-88, and C13-115 shown in Table 3 were constructed. All verification vectors in this experimental example used the same plasmid backbone as Example 2, with a structure of CMV-Cas13-U6-crRNA. The expressed Cas13s all had one NLS at the N-terminus and two NLSs at the C-terminus, for a total of 3×NLS structures. The crRNA consists of a guide sequence and the corresponding DR sequences disclosed in the prior art. Exemplary, the plasmid sequences of CasRx-EGFP (SEQ ID NO: 444), CasRx-VEGFA (SEQ ID NO: 445), and PspCas13b-VEGFA (SEQ ID NO: 446) are given. The plasmid sequences of Cas13X.1-VEGFA, Cas13Y.1-VEGFA, C13-52-VEGFA, C13-55-VEGFA, C13-88-VEGFA, and C13-115-VEGFA differ from the CasRx-VEGFA sequence only in the corresponding replacement of the Cas protein coding sequence and the DR coding sequence.

[0614] The plasmid sequences targeting PTBP1 differ only in the gRNA coding sequence.

[0615] The amino acid sequence of C13-115 is shown in SEQ ID NO: 441, and the corresponding DR sequence is shown in SEQ ID NO: 442.

[0616] Because 293T cells do not contain EGFP sequences, a CasRx-targeted EGFP vector was used as a negative control.

[0617] The spacer targeting EGFP is tgccgttcttctgcttgtcggccatgatat (SEQ ID NO: 422).

[0618] The spacer targeting PTBP1 is GTGGTTGGAGAACTGGATGTAGATGGGCTG (SEQ ID NO: 443).

[0619] The spacer targeting VEGFA is TGGGTGCAGCCTGGGACCACTTGGCATGG (SEQ ID NO: 406).

[0620] The constructed verification vector is as follows:

[0621] Table 17. Verification vector and architecture

[0622] 2. Transfect 293T cells with the vector to be verified

[0623] The validation vector and the control vector were transfected into 293T cells.

[0624] Lipofectamine 2000 (Thermo) was used to transfect 24-well plates. 293T cells not transfected with the plasmid served as blank controls.

[0625] 3. qPCR detection of RNA changes in target genes

[0626] 48 hours after transfection, RNA was extracted from cells using the SteadyPure Universal RNA Extraction Kit AG21017, and RNA concentration was determined using an ultra-micro-spectrophotometer. RNA products were reverse transcribed using the Evo M-MLV Mix Kit with gDNA Clean for qPCR AG11728 Reverse Transcription Kit, and the reverse transcription products were detected using the SYBR Green Premix Pro Taq HS qPCR Kit (Low Rox Plus).

[0627] The primers used in qPCR are as follows:

[0628] Detection of VEGFA: ACCTCCACCATGCCAAGTGG (SEQ ID NO: 410)

[0629] Detection of PTBP1: ATTGTCCCAGATATAGCCGTTG (SEQ ID NO: 390)

[0630] Detection internal reference GAPDH: CCATGGGGAAGGTGAAGGTC (SEQ ID NO: 394)

[0631] according to Green Premix Pro Taq HS qPCR Kit (Rox Plus) instructions for preparing the reaction system and using QuantStudio TM 5. Detection was performed using the Real-Time PCR System.

[0632] This experiment used the relative quantitative method, 2-△△Ct method, to calculate the changes in target RNA.

[0633] The experiment was repeated four times independently, and the results were averaged over the four tests, as shown in Tables 18 and 19.

[0634] Table 18. Knockdown results of VEGFA targeting

[0635] In the VEGFA RNA-targeting test, the mean editing efficiency ranked as CasRx > C13-55 > PspCas13b > C13-52 > Cas13X.1 > C13-88 > Cas13Y.1 > C13-115. Compared to the CasRx-EGFP group, the C13-55, C13-52, and C13-88 groups all significantly knocked down VEGFA RNA, with statistically significant differences (P < 0.05).

[0636] Table 19. Knockdown results of targeted PTBP1

[0637] In tests targeting PTBP1 RNA, the mean editing efficiency ranked as CasRx > C13-55 > PspCas13b > Cas13X.1 > C13-52 > C13-88 > Cas13Y.1. Compared to the CasRx-EGFP group, the C13-55, C13-52, and C13-88 groups all significantly knocked down PTBP1 RNA, with statistically significant differences (P < 0.05).

[0638] 4. Off-target Analysis

[0639] RNAseq sequencing.

[0640] The total RNA samples of the experimental group targeting VEGFA and the control group were subjected to RNAseq sequencing (n=3 samples per group). The library type was LncRNA chain-specific library, the sequencing data volume was 16G, and the sequencing strategy was PE150.

[0641] RNAseq analysis

[0642] (1) Use fastqc and multiqc to perform quality control on the data, and use fastp to remove low-quality reads.

[0643] (2) Reads aligned to human rRNA sequences were removed, and the remaining reads were aligned to the hg38 reference genome using Hisat2 alignment software.

[0644] (3) After alignment, Kallisto software was used to quantify the gene expression levels, and then sleuth software was used to perform expression difference analysis. Genes with |b|>0.5, qval<0.05, and mean_obs>2 were considered differentially expressed genes (DEGs).

[0645] (4) The gRNA guide sequence was aligned to the reference cDNA using EMBOSS water software, and transcripts with the number of aligned bases >= 18, the number of mismatched bases <= 6, and the minimum number of consecutive paired bases >= 8 were considered as predicted targeted transcripts (on-target + off-target), and the corresponding genes were considered as targeted genes (on-target + off-target genes).

[0646] (5) The intersection of the differentially expressed genes with significantly downregulated expression and the targeted genes predicted to be off-target is taken, and the on-target genes are eliminated to obtain the off-target gene set.

[0647] The results are shown in Table 20. The number of off-target genes was CasRx > Cas13X.1 > C13-88 > C13-52 > C13-55. C13-55 had both high editing efficiency and low off-target effects.

[0648] Table 20. Comparison of the number of off-target genes

[0649] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0650] The above-described embodiments merely illustrate several implementations of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, and all such variations and improvements fall within the scope of protection of the present invention.

Claims

1. A Cas protein, characterized in that Its amino acid sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with the sequence shown in any one of SEQ ID NOs: 1-184, 426-428.

2. The Cas protein according to claim 1, characterized in that The Cas protein is capable of forming a CRISPR complex with a guide polynucleotide, the guide polynucleotide comprising a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to direct sequence-specific binding of the CRISPR complex to a target nucleic acid; Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with the sequence shown in any one of SEQ ID NOs: 185-368, 429-431, or the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with the reverse complement sequence of any one of SEQ ID NOs: 185-368, 429-431.

3. The Cas protein according to claim 2, characterized in that The Cas protein can be guided to the target nucleic acid by a guide polynucleotide; or The Cas protein can be guided to a target nucleic acid by a guide polynucleotide and target or modify the target nucleic acid.

4. The Cas protein according to claim 3, characterized in that The target nucleic acid is RNA or DNA.

5. The Cas protein according to claim 4, characterized in that The target nucleic acid is PTBP1 mRNA, AQp1 mRNA, VEGFA mRNA, VEGFR1 mRNA, VEGFR2 mRNA or AR mRNA.

6. The Cas protein according to any one of claims 1 to 5, characterized in that The Cas protein is derived from: the same kingdom, phylum, class, order, family, genus or species as the Cas protein having an amino acid sequence comprising the sequence shown in any one of SEQ ID NOs: 1-184, 426-428; and / or The Cas protein is a CRISPR-Cas effector protein and / or a CRISPR-Cas auxiliary protein.

7. A Cas protein, characterized in that The amino acid sequence of the Cas protein comprises the amino acid sequence shown in any of the following motifs 1-9: Motif 1: R-[NH]-x(2)-[FNI]-H, Motif 2: Px(4)-[VFL]-x(2)-[RK], Motif 3: [LIM]-[KR]-x(2)-Yx(3)-[FI], Motif 4: [RK]-x-[LVI]-x(6)-[GN]-x(4)-[LIV], Motif 5: [RL]-[ARK]-[NH]-x(2)-[RTE]-[LY], Motif 6: [LIK]-[MLI]-x(2)-[VLI]-x-[GSA]-[RKQ]-[LW], Motif 7: [LIR]-[WA]-ERD-[LRI]-x-[YF]-x(2)-[LK], Motif 8: [LI]-[RV]-[NGR]-x-[FIL]-x-[HS]-[FMC], Motif 9: [YHA]-[DR]-x-[KLN]-x(4)-V; or, The amino acid sequence of the Cas protein comprises the amino acid sequence shown in any one of the following motifs 10-13: Motif 10: RN-[YTF]-x-[SAT]-H, Motif 11: G-[LRA]-x-[YF]-[FL]-x-[SC]-[FL]-[FA]-L, Motif 12: [KQCG]-[NTQ]-xD-[IRS]-[VLI]-[LT], Motif 13: RNx(3)-H; or The amino acid sequence of the Cas protein comprises the amino acid sequence shown in any one of the following motifs 14-25: Motif 14: Rx(3)-[VA]-H, Motif 15: [KLN]-[NSMY]-x-[GP]-[FMYLV]-[STDN]-x-[KTER]-x-[LYIV]-[RF]-E, Motif 16: [DSN]-[STPKF]-x-[REK]-x-[KIR]-[LIMAVF], Motif 17: Kx(3)-Yx(2)-[EA]-Ax(2)-L, Motif 18: [FL]-[LMIQ]-[DYEN]-[GIDAS]-[KQ]- -[IVT]-[NS]-x-[LKFM]-x(4)-[ISTAM], Motif 19: [LIF]-x(4)-[SQGN]-[FVI]-x-[RK]-[MISF], Motif 20: [AS]-x(3)-[LAIF]-[GDN], Motif 21: [NS]-N-[VI]-[IVL], Motif 22: [NCS]-x(3)-[VIR]-x-[FSLY]-[VFAP]-[LGITF], Motif 23: K-[NSTQ]-[LMG]-[VMI]-x-[VWIT]-[NL]-[SAT], Motif 24: R-[NR]-x(3)-H, Motif 25: Yx(3)-[RS]-[YFK]-x-[NLS]-[LT]-[SAT]; 8. The Cas protein according to claim 7, characterized in that Among them, A, F, C, U, D, N, E, Q, G, H, L, I, K, O, M, P, R, S, T, V, W, and Y are standard amino acid codes, "x" is any amino acid, the number in the brackets after x represents multiple consecutive xs, "[]" is an optional amino acid code, and "-" is a separator. The amino acid sequence of the Cas protein comprises the amino acid sequences shown in the following motifs 26-34 from the N-terminus to the C-terminus: Motif 26: RH-[ARY]-[STL]-[FNI]-H, Motif 27: P-[RKS]-[LF]-[NMS]-[RSK]-[VF]-[IL]-[NRT]-[RK]-[AL]-[RK], Motif 28: [LI]-[KR]-[MHN]-[LI]-Y-[EAK]-[QTI]-[GD]-[FI], Motif 29: R-[SG]-[LI]-[RL]-[EQ]-x(2)-[RH]-x-[GN]-x(2)-[DS]-x-[LI], Motif 30: [RL-[AR]-[NH]-x(2)-[RT]-[LY], Motif 31: [LK]-[ML]-[MI]-xVx-[GSA]-[RK]-[LW], Motif 32: [LIR]-[WA]-ERD-[LR]-[YNH]-[YF]-[VL]-[TIL]-[LK], Motif 33: [LI]-RNx-[FI]-[SA]-H-[FM]-[NY], Motif 34: Y-[DR]-[RS]-[KL]-[LY]-[KN]-[NS]-[SA]-V; or, The amino acid sequence of the Cas protein comprises the amino acid sequences shown in the following motifs 35-46 from the N-terminus to the C-terminus: Motif 35: [LA]-R-[NQ]-x(2)-[VA]-Hx(2)-E, Motif 36: KNxGF-[ST]-x-[KT]-xLRE, Motif 37: D-[ST]-[IV]-RxKL, Motif 38: Kx-[KL]-[LI]-Yx-[DH]-[EA]-Ax(2)-L-[WY], Motif 39: GKEIN-[ED]-LLTT-[LC]-INKF-[DE]-NI, Motif 40: [EQ]-Lx-[LE]-x-[KN]-SF-[VA]-[KR]-[SM], Motif 41: Ax(2)-[IL]-LG, Motif 42: [NS]-NV-[VL], Motif 43: N-[RE]-A-[VI]-[VI]-[RD]-FVL, Motif 44: KN-[LM]-V-[NY]-[VI]-[N]-[SA], Motif 45: R-[N]-x(3)-H, Motif 46: Y-[NC]-x(2)-R-[YF]-KNL-[ST]-Ix(2)-LFD; Among them, A, F, C, U, D, N, E, Q, G, H, L, I, K, O, M, P, R, S, T, V, W, and Y are standard amino acid codes, "x" is any amino acid, the number in the brackets after x represents multiple consecutive xs, "[]" is an optional amino acid code, and "-" is a separator.

9. The Cas protein according to claim 7, characterized in that The Cas protein comprises a sequence as shown in any one of SEQ ID NOs: 1-15, SEQ ID NOs: 61-63, SEQ ID NOs: 103-119, or a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to a sequence shown in any one of SEQ ID NOs: 1-15, SEQ ID NOs: 61-63, SEQ ID NOs: 103-119; or Any amino acid residue in the amino acid sequence of the Cas protein, except the amino acids determined by motifs 1-25, is conservatively replaced with amino acids based on the wild-type sequence, and the wild-type sequence includes the sequences shown in SEQ ID NO: 1-15, SEQ ID NO: 61-63, and SEQ ID NO: 103-119.

10. The Cas protein according to any one of claims 7 to 9, characterized in that The Cas protein is capable of forming a CRISPR complex with a guide polynucleotide, the guide polynucleotide comprising a direct repeat sequence linked to a guide sequence, the guide sequence being engineered to direct sequence-specific binding of the CRISPR complex to a target nucleic acid; Optionally, the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with any one of SEQ ID NOs: 185-368, 429-431, or the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with the reverse complement of any one of SEQ ID NOs: 185-368, 429-431. At least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity.

11. The Cas protein according to claim 10, characterized in that The Cas protein can be guided to the target nucleic acid by a guide polynucleotide; or The Cas protein can be guided to a target nucleic acid by a guide polynucleotide and target or modify the target nucleic acid.

12. The Cas protein according to claim 11, characterized in that The target nucleic acid is PTBP1 mRNA, AQp1 mRNA, VEGFA mRNA, VEGFR1 mRNA, VEGFR2 mRNA or AR mRNA.

13. The Cas protein according to any one of claims 7 to 9, characterized in that The Cas protein is derived from: the same kingdom, phylum, class, order, family, genus or species as the Cas protein containing the sequence shown in any one of SEQ ID NOs: 1-15, SEQ ID NOs: 61-63, and SEQ ID NOs: 103-119; and / or The Cas protein is a CRISPR-Cas effector protein and / or a CRISPR-Cas auxiliary protein.

14. A fusion protein, recombinant protein or conjugate, characterized in that: It comprises: the Cas protein or a functional fragment thereof according to any one of claims 1 to 13.

15. The fusion protein, recombinant protein or conjugate according to claim 14, characterized in that: The Cas protein or its functional fragment is fused to a protein domain and / or a polypeptide tag, and the fusion does not change the original function of the Cas protein or its functional fragment.

16. The fusion protein, recombinant protein or conjugate according to claim 14, characterized in that: The Cas protein or its functional fragment is fused with a localization tag that provides subcellular localization; Optionally, the localization tag is selected from a nuclear localization signal (NLS) or a nuclear export signal (NES).

17. The fusion protein, recombinant protein or conjugate according to claim 14, characterized in that: The Cas protein or its functional fragment is fused to any one or more protein domains and / or polypeptide tags selected from the following: cytosine deaminase domain, adenosine deaminase domain, translation activation domain, translation inhibition domain, RNA methylation domain, RNA demethylation domain, nuclease domain, splicing factor domain, reporter domain, affinity domain, subcellular localization signal, reporter tag and affinity tag.

18. The fusion protein, recombinant protein or conjugate according to claim 14, characterized in that: The Cas protein or its functional fragment is covalently linked to the Cas protein domain; and / or The Cas protein or a functional fragment thereof is covalently or non-covalently linked to a conjugated molecule, the conjugated molecule does not change the original function of the Cas protein or a functional fragment thereof, and the conjugated molecule is selected from at least one of a polymer, a small molecule compound, an antibody, an oligopeptide, a detectable marker, and a polypeptide other than the Cas protein.

19. The fusion protein, recombinant protein or conjugate according to any one of claims 14 to 18, characterized in that: The structure of the fusion protein is NLS-Cas protein-SV40NLS-nucleoplasminNLS.

20. A guiding polynucleotide, characterized in that It comprises: (i) a direct repeat sequence, which has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to any one of SEQ ID NOs: 185-368, 429-431, or the direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the reverse complement of any one of SEQ ID NOs: 185-368, 429-431. At least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity, the direct repeat sequence is linked to (ii) a guide sequence engineered to hybridize to a target nucleic acid; the guide polynucleotide is capable of forming a CRISPR complex with a Cas protein and guiding the sequence-specific binding of the CRISPR complex to a target nucleic acid, the Cas protein being a CRISPR-Cas effector protein and / or a CRISPR-Cas accessory protein.

21. The guiding polynucleotide according to claim 20, characterized in that The Cas protein has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity compared to the amino acid sequence shown in SEQ ID NO: 1-184, 426-428; or The amino acid sequence of the Cas protein comprises, from the N-terminus to the C-terminus, the amino acid sequence shown in motifs 1-9, the amino acid sequence shown in motifs 10-13, or the amino acid sequence shown in motifs 14-25.

22. The guiding polynucleotide according to claim 20, characterized in that The direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity compared to any one of SEQ ID NOs: 185-199, 245-247, 287-303; or The direct repeat sequence has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity compared to the reverse complement of any one of SEQ ID NOs: 185-199, 245-247, 287-303.

23. The guiding polynucleotide according to claim 20, characterized in that The guide sequence is located at the 3' end of the direct repeat sequence; or The guide sequence is located at the 5' end of the direct repeat sequence.

24. The guiding polynucleotide according to claim 20, characterized in that The guide sequence comprises 15-35 nucleotides; and / or The guide sequence hybridizes to the target nucleic acid with no more than one nucleotide mismatch.

25. The guiding polynucleotide according to claim 20, characterized in that The direct repeat sequence comprises 25 to 40 nucleotides.

26. The guiding polynucleotide according to claim 20, characterized in that The guide polynucleotide further comprises an aptamer sequence.

27. The guiding polynucleotide according to claim 26, characterized in that The aptamer sequence is inserted into a loop of the guide polynucleotide; and / or The aptamer sequence includes an MS2 aptamer sequence, a PP7 aptamer sequence or a Qβ aptamer sequence.

28. The guiding polynucleotide according to claim 20, characterized in that The guide polynucleotide comprises modified nucleotides.

29. The guiding polynucleotide according to claim 28, characterized in that The modification comprises a 2'-O-methyl, a 2'-O-methyl-3'-phosphorothioate or a 2'-O-methyl-3'-thioPACE modification.

30. The guiding polynucleotide according to any one of claims 20 to 29, characterized in that The target nucleic acid is located in the nucleus of a eukaryotic cell.

31. The guiding polynucleotide according to any one of claims 20 to 29, characterized in that The target nucleic acid is optionally selected from TTR RNA, SOD1 RNA, PCSK9 RNA, VEGFA RNA, VEGFR1 RNA, PTBP1 RNA, AQp1 RNA, ANGPTL3 RNA or AR RNA.

32. The guiding polynucleotide according to any one of claims 20 to 29, characterized in that The guide sequence is optionally selected from the sequences shown in SEQ ID NO: 380-381, SEQ ID NO: 398 or SEQ ID NO:

406.

33. A CRISPR-Cas system, characterized in that It contains: The Cas protein according to any one of claims 1 to 13, or the fusion protein, recombinant protein or conjugate according to any one of claims 14 to 19, or a nucleic acid encoding the Cas protein or the fusion protein, and a guide polynucleotide, or a nucleic acid encoding the guide polynucleotide, comprising a direct repeat sequence linked to a guide sequence engineered to hybridize to a target nucleic acid; The guide polynucleotide is capable of forming a CRISPR complex with the Cas protein or fusion protein and guiding the sequence-specific binding of the CRISPR complex to the target nucleic acid.

34. The CRISPR-Cas system according to claim 33, characterized in that The direct repeat sequence has at least 70% sequence identity compared to any one of SEQ ID NOs: 185-368, 429-431.

35. The CRISPR-Cas system according to claim 33, characterized in that The CRISPR-Cas system is a CRISPR-Cas13 system.

36. The CRISPR-Cas system according to any one of claims 33 to 35, characterized in that The target nucleic acid is selected from TTR RNA, SOD1 RNA, PCSK9 RNA, VEGFA RNA, VEGFR1 RNA, PTBP1 RNA, AQp1 RNA, ANGPTL3 RNA or AR RNA; Optionally, the guide sequence is selected from the sequences shown in SEQ ID NO:380-381, SEQ ID NO:398 or SEQ ID NO:

406.

37. A CRISPR-Cas system, characterized in that It contains: The guiding polynucleotide or the nucleic acid encoding the same according to any one of claims 20 to 32, and CRISPR-Cas effector protein and / or CRISPR-Cas accessory protein, fusion protein thereof, recombinant protein thereof, conjugate thereof, or nucleic acid encoding the same.

38. A vector system comprising the CRISPR-Cas system of any one of claims 33-37, wherein: The vector system comprises one or more vectors, which comprise a polynucleotide sequence encoding the Cas protein, the fusion protein, the recombinant protein or the conjugate and a polynucleotide sequence encoding a guide polynucleotide.

39. An adeno-associated virus vector comprising the CRISPR-Cas system of any one of claims 33-37, wherein: The adeno-associated virus vector comprises a DNA sequence encoding the Cas protein, the fusion protein, the recombinant protein or the conjugate and a DNA sequence of the guide polynucleotide.

40. A lentiviral vector comprising the CRISPR-Cas system of any one of claims 33 to 37, wherein the lentiviral vector comprises a guide polynucleotide and an mRNA encoding the Cas protein or the fusion protein, recombinant protein or conjugate; Optionally, the lentiviral vector is pseudotyped with an envelope protein; Optionally, the mRNA encoding the Cas protein or fusion protein is linked to an aptamer sequence.

41. A ribonucleoprotein complex comprising the CRISPR-Cas system of any one of claims 33-37, wherein The ribonucleoprotein complex is formed by the guide polynucleotide and the Cas protein or fusion protein.

42. A virus-like particle comprising the CRISPR-Cas system of any one of claims 33-37, wherein The virus-like particle comprises a ribonucleoprotein complex formed by the guide polynucleotide and the Cas protein or fusion protein; Optionally, the Cas protein or fusion protein is fused to the gag protein.

43. An isolated nucleic acid, characterized in that It encodes the Cas protein described in any one of claims 1 to 13 or the fusion protein, recombinant protein or conjugate described in any one of claims 14 to 19.

44. An isolated nucleic acid, characterized in that It encodes the guiding polynucleotide of any one of claims 20-32.

45. A delivery composition, characterized in that It includes: a delivery system, which carries the Cas protein described in any one of claims 1-13, the fusion protein, recombinant protein or conjugate described in any one of claims 14-19, the guiding polynucleotide described in any one of claims 20-32, the CRISPR-Cas system described in any one of claims 33-37, the vector system described in claim 38, the adeno-associated virus vector described in claim 39, the lentiviral vector described in claim 40, the ribonucleoprotein complex described in claim 41, the virus-like particle described in claim 42, or the nucleic acid described in claim 43 or 44.

46. ​​The delivery composition of claim 44, wherein The delivery system is selected from at least one of lipid nanoparticles and extracellular vesicles.

47. A eukaryotic cell comprising the Cas protein of any one of claims 1-13, the fusion protein, recombinant protein or conjugate of any one of claims 14-19, the guide polynucleotide of any one of claims 20-32, or the CRISPR-Cas system of any one of claims 33-37; Optionally, the eukaryotic cell is a mammalian cell.

48. A pharmaceutical composition, characterized in that It comprises the Cas protein of any one of claims 1-13, the fusion protein, recombinant protein or conjugate of any one of claims 14-19, the guiding polynucleotide of any one of claims 20-32, or the CRISPR-Cas system of any one of claims 33-37, the vector system of claim 38, the adeno-associated virus vector of claim 39, the lentiviral vector of claim 40, the ribonucleoprotein complex of claim 41, the virus-like particle of claim 42, the nucleic acid of claim 43 or 44, the delivery composition of claim 45 or 46, or the eukaryotic cell of claim 47.

49. An in vitro composition, characterized in that It comprises the CRISPR-Cas system as described in any one of claims 33 to 37, and a labeled detector RNA that cannot hybridize with the guide polynucleotide.

50. Use of a Cas protein according to any one of claims 1-13, a fusion protein, a recombinant protein or a conjugate according to any one of claims 14-19, a guiding polynucleotide according to any one of claims 20-32, a CRISPR-Cas system according to any one of claims 33-37, or an isolated nucleic acid according to claim 46 or claim 47 in detecting a target nucleic acid in a nucleic acid sample suspected of containing the target nucleic acid or in preparing a reagent for detecting a target nucleic acid in a nucleic acid sample suspected of containing the target nucleic acid.

51. A use of a Cas protein according to any one of claims 1-13, a fusion protein according to any one of claims 14-19, a guide polynucleotide according to any one of claims 20-32, a CRISPR-Cas system according to any one of claims 33-37, a vector system according to claim 38, an adeno-associated virus vector according to claim 39, a lentiviral vector according to claim 40, a ribonucleoprotein complex according to claim 41, a virus-like particle according to claim 42, a nucleic acid according to claim 43 or 44, a delivery composition according to claim 45 or 46, or a eukaryotic cell according to claim 47 in any of the following or in the preparation of an agent for implementing any of the following schemes: Cut or nick one or more target nucleic acid molecules, activate or upregulate one or more target nucleic acid molecules, activate or inhibit translation of one or more target nucleic acid molecules, inactivate one or more target nucleic acid molecules, visualize, label or detect one or more target nucleic acid molecules, bind one or more target nucleic acid molecules, transport one or more target nucleic acid molecules, and mask one or more target nucleic acid molecules.

52. A method for diagnosing, treating or preventing a disease or condition associated with a target nucleic acid, characterized in that: The Cas protein of any one of claims 1-13, the fusion protein of any one of claims 14-19, the guide polynucleotide of any one of claims 20-32, the CRISPR-Cas system of any one of claims 33-37, the vector system of claim 38, the adeno-associated viral vector of claim 39, the lentiviral vector of claim 40, the ribonucleoprotein complex of claim 41, the virus-like particle of claim 42, the nucleic acid of claim 43 or 44, the delivery composition of claim 45 or 46, or the eukaryotic cell of claim 47 is administered to a sample of a subject in need or to a subject in need.

53. A use of a Cas protein according to any one of claims 1-13, a fusion protein according to any one of claims 14-19, a guide polynucleotide according to any one of claims 20-32, a CRISPR-Cas system according to any one of claims 33-37, a vector system according to claim 38, an adeno-associated virus vector according to claim 39, a lentiviral vector according to claim 40, a ribonucleoprotein complex according to claim 41, a virus-like particle according to claim 42, a nucleic acid according to claim 43 or 44, a delivery composition according to claim 45 or 46, or a eukaryotic cell according to claim 47 in the preparation of a medicament for diagnosing, treating or preventing a disease or disorder associated with a target nucleic acid.