Cas13 protein, CRISPR-Cas system and its applications
By providing a novel Cas13 protein and CRISPR-Cas system, the problems of large size and non-specific RNase activity of existing Cas13 proteins are solved, enabling efficient RNA-targeted cleavage and editing, suitable for gene editing and detection in various cell types.
Patent Information
- Application Number
- CN202211458493.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-11-21
AI Technical Summary
Existing Cas13 proteins suffer from large size and non-specific RNase activity defects in RNA-targeted cleavage, affecting editing performance. Furthermore, the functional and structural diversity of CRISPR/Cas13 systems from different microbial sources has not been fully explored.
A Cas13 protein is provided, the amino acid sequence of which is shown in SEQ ID NO.1-10, or a protein that retains its biological function by substituting, deleting or adding amino acids based on these sequences, and can be fused with a heterologous functional domain to form a CRISPR-Cas system with guide RNA for gene editing and detection.
It improves RNA degradation efficiency, reduces off-target potential, and is suitable for gene editing and RNA detection in eukaryotic and prokaryotic cells, showing broad application prospects.
Smart Images

Figure CN116355877B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the Cas13 protein, the CRISPR-Cas system, and their applications, and belongs to the field of biotechnology. Background Technology
[0002] The CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) system was developed by bacteria and archaea to defend against invading bacteriophage DNA. The CRISPR system comprises two families: family I is further divided into types I, III, and IV, and family II into types II, V, and VI. These six system types are further subdivided into 19 subtypes. Many prokaryotes contain multiple CRISPR-Cas systems, indicating that they are compatible and may share components.
[0003] The CRISPR / Cas13 system is a CRISPR system capable of targeting and cleaving RNA. This system consists of the Cas13 protein and CRISPR RNA (guide RNA) assembling into an RNA-targeting effector complex guided by the guide RNA. Currently, four isotypes of the Cas13 family have been identified: Cas13a, Cas13b, Cas13c, and Cas13d.
[0004] Compared to RNAi and CRISPRi technologies, the CRISPR / Cas13 system achieves higher RNA degradation rates and has a lower off-target potential. However, currently discovered Cas13 proteins also have drawbacks, such as large size and non-specific RNase activity, which affect the editing performance of the CRISPR / Cas13 system on target RNA. CRISPR / Cas13 systems from different microbial sources exhibit functional and structural diversity. Therefore, further exploration of more Cas13 proteins is needed to construct CRISPR / Cas13 systems with superior performance. Summary of the Invention
[0005] To address the above problems, the present invention provides a Cas13 protein, wherein the Cas13 protein is a protein as described in (a) or (b) below:
[0006] (a) A protein comprising the amino acid sequence shown in SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.3, SEQ ID NO.4, SEQ ID NO.5, SEQ ID NO.6, SEQ ID NO.7, SEQ ID NO.8, SEQ ID NO.9 or SEQ ID NO.10;
[0007] (b) A protein in which one or more amino acids have been substituted, deleted or added in the amino acid sequence defined in (a) and which substantially retains the biological function of the sequence from which it originated.
[0008] In one embodiment of the present invention, (a) is a protein composed of the amino acid sequence shown in SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.3, SEQ ID NO.4, SEQ ID NO.5, SEQ ID NO.6, SEQ ID NO.7, SEQ ID NO.8, SEQ ID NO.9 or SEQ ID NO.10.
[0009] In one embodiment of the invention, in (b), the amino acid sequence of the protein has at least 85% sequence identity compared to (a), and substantially retains the biological function of the sequence from which it originates.
[0010] In one embodiment of the present invention, in (b), the amino acid sequence of the protein has at least 95% sequence identity with that of (a), and substantially retains the biological function of the sequence from which it originates.
[0011] As is well known to those skilled in the art, Cas13 proteins can undergo amino acid changes, such as substitution, deletion, or addition, resulting in Cas13 proteins that retain their function or activity.
[0012] The term "substitution" refers to the replacement of an amino acid residue at a certain position in an amino acid sequence by another amino acid residue; wherein, "substitution" can be a conserved amino acid substitution.
[0013] "Conservative modification," "conservative substitution," or "conservative replacement" refers to the replacement of amino acids in a protein with other amino acids that have similar characteristics (such as charge, side chain size, hydrophobicity / hydrophilicity, main chain conformation, and rigidity), so that the protein can be frequently modified without changing its biological activity.
[0014] Those skilled in the art will understand that, in general, a single amino acid substitution in a non-essential region of a polypeptide does not substantially alter its biological activity (see Watson et al. (1987), Molecular Biology of the Gene, The Benjamin / Cummings Pub. Co., p. 224, (4th edition)). Furthermore, substitutions of structurally or functionally similar amino acids are unlikely to disrupt biological activity. Exemplary conserved substitutions are described in “Exemplary Conserved Amino Acid Substitutions” (see Table 1).
[0015] Table 1 Exemplary Conserved Substitutions of Amino Acids
[0016] Original residues Conservative replacement Ala(A) Gly(G); Ser(S) Arg(R) Lys(K); His(H) Asn(N) Gln(Q); His(H); Asp(D) Asp(D) Glu(E); Asn(N) Cys(C) Ser(S); Ala(A); Val(V) Gln(Q) Asn(N); Glu(E) Glu(E) Asp(D); Gln(Q) Gly(G) Ala(A) His(H) Asn(N); Gln(Q) Ile(I) Leu(L); Val(V) Leu(L) Ile(I); Val(V) Lys(K) Arg(R); His(H) Met(M) Leu(L); Ile(I); Tyr(Y) Phe(F) Tyr(Y); Met(M); Leu(L) Pro(P) Ala(A) Ser(S) Thr(T) Thr(T) Ser(S) Trp(W) Tyr(Y); Phe(F) Tyr(Y) Trp(W); Phe(F) Val(V) Ile(I); Leu(L)
[0017] The Cas13 protein provided by this invention also includes its functional fragments or derivatives, wherein the derivatives have the same sequence as the Cas13 protein in any one of SEQ ID NO.1-10 on the functional domain (HEPN) or RXXXXH motif.
[0018] The present invention also provides a fusion protein comprising the above-mentioned Cas13 protein and a heterologous functional domain.
[0019] In one embodiment of the present invention, the heterologous functional domain includes at least one of the following: deaminase (e.g., ADAR, APOBEC, AID, or TAD), reporter protein (e.g., glutathione S-transferase, horseradish peroxidase, chloramphenicol acetyltransferase, β-galactosidase, β-glucuronidase, or autofluorescent protein), detection tag, localization signal, protein targeting domain, DNA binding domain (e.g., methylation-binding protein, LexADBD, or Gal4 DBD), epitope tag (e.g., histidine tag, V5 tag, FLAG tag, influenza virus hemagglutinin tag, Myc tag, VSV-G tag, or thioredoxin tag), transcription activation domain (e.g., VP64), transcription repression domain (e.g., KRAB or SID domain), nuclease (e.g., FokI), methylation-active enzyme, demethylase, transcription release factor, RNA cleavage peptide (e.g., peptide with single-strand RNA cleavage activity, peptide with double-strand RNA cleavage activity), nucleic acid ligase, histone modification active factor, or nucleic acid binding active factor.
[0020] In one embodiment of the present invention, the heterologous functional domain is fused to the N-terminus or C-terminus of the Cas13 protein, or the heterologous functional domain is fused within the Cas13 protein.
[0021] In one embodiment of the present invention, the positioning signal is a nuclear positioning signal (NLS) or a nuclear output signal (NES).
[0022] The present invention also provides a CRISPR-Cas system comprising the above-mentioned Cas13 protein and guide RNA, or the CRISPR-Cas system comprising the above-mentioned fusion protein and guide RNA; wherein the guide RNA binds to the Cas13 protein to form a complex, guiding the complex to contact the target RNA.
[0023] In one embodiment of the invention, the guide RNA comprises a spacer sequence and a direct repeat (DR) sequence.
[0024] In one embodiment of the present invention, the spacer sequence is 15 to 60 nucleotides in size.
[0025] In one embodiment of the present invention, the spacer sequence is 25 to 55 nucleotides in size.
[0026] In one embodiment of the present invention, the spacer sequence is 30 nucleotides in size.
[0027] In one embodiment of the present invention, the size of the same-direction repeat sequence is 15 to 40 nucleotides.
[0028] In one embodiment of the present invention, the size of the same-direction repeat sequence is 20 to 38 nucleotides (e.g., 22, 36, 37, 38).
[0029] In one embodiment of the present invention, the spacer sequence is complementary to the target RNA by at least 15 nucleotides.
[0030] In one embodiment of the invention, the spacer sequence is at least 85% complementary to the target RNA in terms of nucleotides.
[0031] The present invention also provides a polynucleotide that encodes the Cas13 protein described above, or the fusion protein described above, or the CRISPR-Cas system described above.
[0032] In one embodiment of the invention, the polynucleotide is codon-optimized to facilitate expression in host cells.
[0033] In one embodiment of the present invention, the host cell includes a eukaryotic cell or a prokaryotic cell.
[0034] In one embodiment of the present invention, the prokaryotic cells include bacteria.
[0035] In one embodiment of the present invention, the eukaryotic cells include animal cells or plant cells.
[0036] In one embodiment of the present invention, the animal cells include human cells or non-human mammalian cells, insect cells, bird cells, and rodent cells.
[0037] In one embodiment of the present invention, the nucleotide sequence of the polynucleotide is as shown in SEQ ID NO. 11, SEQ ID NO. 12, SEQ ID NO. 13, SEQ ID NO. 14, SEQ ID NO. 15, SEQ ID NO. 16, SEQ ID NO. 17, SEQ ID NO. 18, SEQ ID NO. 19 or SEQ ID NO. 20.
[0038] The present invention also provides a carrier, characterized in that the carrier carries the above-mentioned polynucleotides.
[0039] In one embodiment of the invention, the vector carries polynucleotides that are operatively linked to regulatory elements.
[0040] In one embodiment of the present invention, the control element includes a promoter.
[0041] In one embodiment of the present invention, the regulatory element further includes at least one of an enhancer, a polyadenylation signal (PAS), a terminator, a protein degradation signal, a transposon, a leader sequence, a polyadenylation sequence, or a marker gene.
[0042] In one embodiment of the present invention, the promoter is at least one of a constitutive promoter or an inducible promoter.
[0043] In one embodiment of the present invention, the promoter is at least one of a broad-spectrum expression promoter or a tissue-specific promoter.
[0044] In one embodiment of the present invention, the carrier includes at least one of a viral carrier or a non-viral carrier.
[0045] In one embodiment of the present invention, the viral vector includes at least one of a retroviral vector, a bacteriophage vector, adenovirus vector, adeno-associated virus vector, vaccinia virus vector, hybrid virus vector, baculovirus vector, herpes simplex virus vector, or lentiviral vector.
[0046] In one embodiment of the present invention, the non-viral vector includes a plasmid vector.
[0047] The present invention also provides a delivery system that carries the above-mentioned Cas13 protein, the above-mentioned fusion protein, the above-mentioned CRISPR-Cas system, the above-mentioned polynucleotide, or the above-mentioned vector.
[0048] In one embodiment of the present invention, the delivery carrier of the delivery system includes at least one of liposome carrier, chitosan carrier, polymer carrier, nanoparticle carrier, exosome carrier, exogenous body carrier, microbubble carrier, viral carrier, gene gun or electroporation device.
[0049] The present invention also provides a host cell carrying the above-mentioned Cas13 protein, the above-mentioned fusion protein, the above-mentioned CRISPR-Cas system, the above-mentioned polynucleotide or the above-mentioned vector.
[0050] In one embodiment of the present invention, the host cell includes a eukaryotic cell or a prokaryotic cell.
[0051] In one embodiment of the present invention, the prokaryotic cells include bacteria.
[0052] In one embodiment of the present invention, the eukaryotic cells include animal cells or plant cells.
[0053] In one embodiment of the present invention, the animal cells include human cells, non-human mammalian cells, insect cells, bird cells, and rodent cells.
[0054] The present invention also provides a model, which is a plant model or an animal model; the plant model or animal model contains the aforementioned host cells.
[0055] The present invention also provides a method for gene editing, the method comprising contacting target RNA with the Cas13 protein, the fusion protein, the CRISPR-Cas system, the polynucleotide, the vector, or the host cell.
[0056] In one embodiment of the present invention, the guide RNA binds to the Cas13 protein to form a complex, and guides the complex to contact the target RNA to perform gene editing on the target RNA.
[0057] In one embodiment of the present invention, the target RNA is located inside or outside the cell.
[0058] In one embodiment of the present invention, the target RNA is encoded by prokaryotic DNA, eukaryotic DNA, or is derived from a virus; the eukaryotic DNA is non-human mammalian DNA, human DNA, plant DNA, insect DNA, bird DNA, reptile DNA, fish DNA, yeast DNA, or bacterial DNA.
[0059] In one embodiment of the present invention, the target RNA is a disease-specific RNA, mRNA, tRNA, rRNA, non-coding RNA, lncRNA, or nuclear RNA.
[0060] In one embodiment of the present invention, the gene editing of the target RNA is to convert the adenine of the target RNA into hypoxanthine, or to convert the cytosine of the target RNA into uracil.
[0061] The present invention also provides a gene-editing product, which is obtained by gene editing of target RNA using the above method.
[0062] The present invention also provides the application of the above-mentioned Cas13 protein or the above-mentioned fusion protein or the above-mentioned CRISPR-Cas system or the above-mentioned method or the above-mentioned polynucleotide or the above-mentioned vector or the above-mentioned delivery system or the above-mentioned host cell or the above-mentioned method in gene editing.
[0063] The present invention also provides a kit comprising the above-described Cas13 protein, the above-described fusion protein, the above-described CRISPR-Cas system, the above-described polynucleotide, the above-described vector, the above-described delivery system, or the above-described host cell.
[0064] In one embodiment of the present invention, the kit is used for the detection of target RNA in a sample to be tested.
[0065] The present invention also provides a nucleic acid detection method, wherein the method involves using the above-mentioned kit to perform nucleic acid detection on the sample to be tested.
[0066] In one embodiment of the present invention, the method includes contacting the sample to be tested with the above-mentioned kit. If the sample to be tested contains target RNA, the guide RNA will bind to the Cas13 protein to form a complex and guide the complex to contact the target RNA.
[0067] The present invention also provides the application of the above-mentioned Cas13 protein or the above-mentioned fusion protein or the above-mentioned CRISPR-Cas system or the above-mentioned polynucleotide or the above-mentioned vector or the above-mentioned delivery system or the above-mentioned host cell or the above-mentioned kit or the above-mentioned method in nucleic acid detection.
[0068] The present invention also provides the use of the above-mentioned Cas13 protein or the above-mentioned fusion protein or the above-mentioned CRISPR-Cas system or the above-mentioned polynucleotide or the above-mentioned vector or the above-mentioned delivery system or the above-mentioned host cell in the preparation of a medicament for treating a condition or disease of a subject in need.
[0069] In some embodiments of the present invention, the condition or disease includes cancer or infectious diseases.
[0070] In some embodiments of the present invention, the cancer includes at least one of Wilms' tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid carcinoma, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin's lymphoma, non-Hodgkin's lymphoma, or bladder cancer.
[0071] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. As used herein, the following terms have the meanings assigned to them as described below, unless otherwise stated.
[0072] "Polynucleotide" and "nucleic acid" refer to polymeric forms of nucleotides (such as RNA or DNA) of any length. The term includes, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derived nucleotide bases.
[0073] "DNA sequence encoding a specific RNA" refers to the DNA nucleotide sequence transcribed into RNA.
[0074] "Regulatory elements," "regulatory components," or "DNA regulatory sequences" or "control elements" refer to transcriptional and translation regulatory sequences, such as promoters, enhancers, and terminators.
[0075] A promoter (also known as a promoter sequence) is a DNA regulatory region that binds to RNA polymerase and initiates transcription of downstream (3' direction) coding or non-coding sequences. Generally, various promoters, including inducible promoters, can be used to drive the expression of the various vectors of this invention.
[0076] "Codon optimization" refers to modifying nucleic acid sequences to make them better expressed in host cells. Generally, this involves replacing at least one codon in the original nucleic acid sequence (e.g., one or more, such as 10, 20 or more) with a codon that is more frequently or most frequently used in the host cell gene, while maintaining the natural amino acid sequence expressed.
[0077] The "target RNA" can be any target RNA molecule, either naturally occurring or engineered (e.g., mRNA, tRNA, ribosomal RNA (rRNA), microRNA (miRNA), interfering RNA (siRNA), ribozymes, satellite RNA, or viral RNA). The target RNA can be associated with a symptom or disease (e.g., an infectious disease or cancer). Perfect complementarity between the target RNA and the guide RNA is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of the CRISPR-Cas complex.
[0078] “Target RNA” includes disease-associated RNA. “Disease-associated” RNA refers to any RNA that produces translational products at abnormal levels or in abnormal forms in tissues or cells derived from disease-affected tissues or cells compared to non-disease control tissues or cells: for example, it could be RNA transcribed from a gene that has become abnormally highly expressed, or RNA that has become abnormally low in expression, wherein the altered expression is associated with the occurrence and / or progression of the disease. Disease-associated RNA also refers to RNA transcribed from a gene with a mutation or genetic variation that is directly responsible for or in linkage disequilibrium with the gene responsible for the cause of the disease. Translational products can be known or unknown and can be at normal or abnormal levels. Target RNA can be any RNA that is endogenous or exogenous to the cell. For example, target RNA can be RNA present in the nucleus of a eukaryotic cell.
[0079] "Naturally occurring (also known as unmodified, wild-type (wt))" refers to nucleic acids, polypeptides, cells, or organisms that exist in nature. For example, polypeptide or polynucleotide sequences that can be isolated from natural sources and exist in organisms are naturally occurring.
[0080] A "fusion protein" is a hybrid polypeptide that contains protein domains from at least two different proteins. One protein may be located at the N-terminal (N-terminal) portion or the C-terminal (C-terminal) portion of the fusion protein, thus forming an N-terminal fusion protein or a C-terminal fusion protein, respectively.
[0081] "Deaminase" refers to an enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is "cytidine deaminase (or cytosine deaminase)," which catalyzes the hydrolysis and deamination of cytidine or deoxycytidine to uracil or deoxyuracil, respectively. In other embodiments, the deaminase is "adenosine deaminase," which refers to a protein, or polypeptide, or one or more functional domains of a protein, or one or more functional domains of a polypeptide, capable of catalyzing the hydrolysis and deamination reaction that converts adenine to hypoxanthine.
[0082] A “CRISPR-Cas system” (also known as a “CRISPR system” or “CRISPR / Cas system”) typically contains transcripts or other elements that are associated with the expression of a CRISPR-related (“Cas”) gene, or that are capable of directing the activity of the Cas gene. In some embodiments, components of a CRISPR system may include nucleic acids (e.g., vectors), protein components, or combinations thereof that encode one or more components of the system.
[0083] "Cas protein" refers to a CRISPR-associated protein (Cas) (also known as "CRISPR-associated protein," "CRISPR effector," "effector," "Cas protein," "Cas enzyme," or "CRISPR enzyme"), which is a protein that performs enzymatic activity and / or binds to target sites on nucleic acids specified by an RNA guide. In some embodiments, the Cas protein has endonuclease activity, nicking enzyme activity, exonuclease activity, transposase activity, and / or excision activity; in other embodiments, the Cas protein may be nuclease-inactivated or partially inactivated.
[0084] "Guide RNA (also known as directing RNA, gRNA or crRNA)" refers to any RNA molecule that facilitates the targeting of the Cas13 protein of the present invention to a target nucleic acid (such as DNA and / or RNA). "Guide RNA" includes one or more guide RNAs and their equivalents known to those skilled in the art, including but not limited to RNA-based molecules (e.g., direct repeat (DR) sequences) capable of forming a complex with the Cas protein, and containing a sequence (e.g., a spacer sequence) that is sufficiently complementary to the target nucleic acid sequence to hybridize with the target nucleic acid sequence and guide the complex to bind specifically to the target nucleic acid sequence.
[0085] "Vector" refers to a polynucleotide composition used to transfer, deliver, or introduce nucleic acids into host cells. Suitable vectors include plasmid vectors, bacteriophage vectors, viral vectors (e.g., retroviral vectors, adeno-associated virus vectors, herpes simplex virus vectors, AAV vectors, lentiviral vectors, baculovirus vectors), etc.
[0086] "Delivery system" comprising a delivery vector, the delivery vector including one or more liposomes, nanoparticles, exosomes, exosomes, microvesicles, viral vectors, gene guns or electroporation devices, etc.
[0087] "Functional fragment" refers to a protein or polypeptide sequence that contains fewer amino acids than the original sequence of the protein or polypeptide, but the remaining amino acid sequence still retains a certain proportion (e.g., 10%, 20%, 30%, 40%, 50%, or 60-99%, 100%) of the functional activity relative to the original reference sequence (e.g., the protein or polypeptide can be modified by substitution, insertion, deletion, and / or addition of one or more amino acids while retaining a certain proportion of enzyme activity).
[0088] "Identity" refers to the matching sequences between two polypeptides or two nucleic acids. The molecules are considered identical at that position when positions in two compared sequences are occupied by the same base or amino acid monomer subunit (e.g., if each position in two DNA molecules is occupied by adenine, or each position in two polypeptides is occupied by lysine). The "percentage identity" between two sequences is a function of the number of common matching positions divided by the number of positions being compared, multiplied by 100. For example, if six out of ten positions in two sequences match, then the two sequences have 60% identity. For instance, the DNA sequences CTGACT and CAGGTT have 50% identity (three out of six positions match). Typically, two sequences are compared to produce the maximum possible identity.
[0089] "Hybridization," "complementary," or "substantially complementary" refers to nucleic acids (e.g., RNA, DNA) containing nucleotide sequences that enable the nucleic acid to non-covalently bind (i.e., form Watson-Crick base pairs and / or G / U base pairs) with another nucleic acid under appropriate in vitro and / or in vivo temperature and solution ionic strength conditions in a sequence-specific, antiparallel (i.e., nucleic acid-specific binding to complementary nucleic acids), "annealing," or "hybridization." Hybridization requires both nucleic acids to contain complementary sequences, but mismatches between bases are possible. Typically, hybridizable nucleic acids are 8 nucleotides or longer (e.g., 10 nucleotides or longer, 12 nucleotides or longer, 15 nucleotides or longer, 20 nucleotides or longer, 22 nucleotides or longer, 25 nucleotides or longer, or 30 nucleotides or longer).
[0090] "Directional repeat (DR) sequence" refers to the DNA-coding sequence in a CRISPR locus, or the RNA encoded by it in a guide RNA. When described at the RNA level, each T should be understood as representing a U.
[0091] "Host cell" includes cells or cell lines or their progeny that are in vitro, ex vivo, or in vivo, and the cells or cell lines or their progeny include: the Cas13 protein, fusion protein, CRISPR-Cas system, polynucleotide, vector, or delivery system described in this invention.
[0092] "Subject," "individual," and "patient" refer to an individual organism, such as mammals, including but not limited to rats, apes, humans, non-human mammals, ungulates, felines, and canines. It also includes tissues, cells, and their progeny of biological entities obtained in vivo or cultured in vitro.
[0093] The "sample to be tested" may contain whole cells, live cells, or cell fragments, and may contain or be derived from body fluids.
[0094] "Operationally ligated" refers to linking a nucleotide sequence of interest to a regulatory sequence in a manner that allows for the expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a target cell when a vector is introduced into a target cell).
[0095] The technical solution of this invention has the following advantages:
[0096] This invention provides a Cas13 protein, the amino acid sequence of which is shown in SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.3, SEQ ID NO.4, SEQ ID NO.5, SEQ ID NO.6, SEQ ID NO.7, SEQ ID NO.8, SEQ ID NO.9 or SEQ ID NO.10, and named Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, Cas13T_8; Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, Cas13T_8, respectively. Cas13T6, Cas13T7, and Cas13T8 have low homology with the previously reported Cas13 protein and exhibit RNA nuclease editing activity on target RNA. Therefore, Cas13a2, Cas13b1, Cas13b2, Cas13d2, Cas13T1, Cas13T4, Cas13T5, Cas13T6, Cas13T7, and Cas13T8 all show great promise for gene editing.
[0097] The Cas proteins involved in this invention can also be used for RNA detection. The Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8 proteins of this invention exhibit non-specific / incidental RNase activity (paracleavage activity) and can be used to detect the presence of specific RNA.
[0098] In addition, Cas13d_2 has only 511 amino acids, while Cas13T_5 has only 447 amino acids. The smaller Cas13 protein is beneficial for encapsulation, packaging and delivery in the gene editing process. Attached Figure Description
[0099] Figure 1 A schematic diagram of the genomic locus structure of the CRISPR-Cas system is shown.
[0100] Figure 2 The secondary structure predicted by the directional repeat (DR) sequence associated with Cas13a_2 is shown.
[0101] Figure 3 The secondary structure predicted by the directional repeat (DR) sequence associated with Cas13b_1 is shown.
[0102] Figure 4 The secondary structure predicted by the directional repeat (DR) sequence associated with Cas13b_2 is shown.
[0103] Figure 5 The secondary structure predicted by the directional repeat (DR) sequence associated with Cas13d_2 is shown.
[0104] Figure 6 The secondary structure predicted by the directional repeat (DR) sequence associated with Cas13T_1 is shown.
[0105] Figure 7 The secondary structure predicted by the directional repeat (DR) sequence associated with Cas13T_4 is shown.
[0106] Figure 8 The secondary structure predicted by the directional repeat (DR) sequence associated with Cas13T_5 is shown.
[0107] Figure 9 The secondary structure predicted by the directional repeat (DR) sequence associated with Cas13T_6 is shown.
[0108] Figure 10 The secondary structure predicted by the directional repeat (DR) sequence associated with Cas13T_7 is shown.
[0109] Figure 11 The secondary structure predicted by the directional repeat (DR) sequence associated with Cas13T_8 is shown.
[0110] Figure 12 The effector protein phylogenetic tree of Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, Cas13T_8 and previously identified Cas13 proteins is shown.
[0111] Figure 13 The structural domains of Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8 are shown. Figure 13 In the diagram, the vertical line represents the position of the RXXXXH motif.
[0112] Figure 14 The predicted three-dimensional structure of the protein by Cas13a_2 is shown.
[0113] Figure 15 The predicted three-dimensional structure of the protein by Cas13b_1 is shown.
[0114] Figure 16 The predicted three-dimensional structure of the protein by Cas13b_2 is shown.
[0115] Figure 17 The predicted three-dimensional structure of the protein by Cas13d_2 is shown.
[0116] Figure 18 The predicted three-dimensional structure of the protein by Cas13T_1 is shown.
[0117] Figure 19 The predicted three-dimensional structure of the protein by Cas13T_4 is shown.
[0118] Figure 20 The predicted three-dimensional structure of the protein by Cas13T_5 is shown.
[0119] Figure 21 The predicted three-dimensional structure of the protein by Cas13T_6 is shown.
[0120] Figure 22 The predicted three-dimensional structure of the protein by Cas13T_7 is shown.
[0121] Figure 23 The predicted three-dimensional structure of the protein by Cas13T_8 is shown.
[0122] Figure 24 The flowchart for the cleavage activity verification experiment is shown.
[0123] Figure 25 The relationship between guide RNAs with different structures (5'DR-sgRNA_1~5'DR-sgRNA_4, 5'DR-sgRNA_7, 5'DR-sgRNA_10) and the editing activity of Cas13a_2 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 25 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0124] Figure 26 The relationship between guide RNAs with different structures (3'DR-sgRNA_1~3'DR-sgRNA_3, 3'DR-sgRNA_5~3'DR-sgRNA_7) and the editing activity of Cas13b_1 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 26 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0125] Figure 27 The relationship between guide RNAs with different structures (3'DR-sgRNA_8~3'DR-sgRNA_10, 5'DR-sgRNA_1~5'DR-sgRNA_5) and the editing activity of Cas13b_1 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 27 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0126] Figure 28 The relationship between guide RNAs with different structures (5'DR-sgRNA_6~5'DR-sgRNA_8, 5'DR-sgRNA_10) and the editing activity of Cas13b_1 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells) was shown. Figure 28 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0127] Figure 29 The relationship between guide RNAs with different structures (5'DR-sgRNA_1~5'DR-sgRNA_4) and the editing activity of Cas13b_2 on target RNA was shown (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 29In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0128] Figure 30 The relationship between guide RNAs with different structures (5'DR-sgRNA_5~5'DR-sgRNA_8, 5'DR-sgRNA_10) and the editing activity of Cas13b_2 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells) was shown. Figure 30 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0129] Figure 31 The relationship between guide RNAs with different structures (3'DR-sgRNA_1~3'DR-sgRNA_4) and the editing activity of Cas13d_2 on target RNA was shown (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 31 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0130] Figure 32 The relationship between guide RNAs with different structures (3'DR-sgRNA_5~3'DR-sgRNA_8, 3'DR-sgRNA_10) and the editing activity of Cas13d_2 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells) was shown. Figure 32 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0131] Figure 33 The relationship between guide RNAs with different structures (3'DR-sgRNA_1~3'DR-sgRNA_2, 3'DR-sgRNA_4~3'DR-sgRNA_7) and the editing activity of Cas13T_1 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 33 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0132] Figure 34 The relationship between guide RNAs with different structures (3'DR-sgRNA_8~3'DR-sgRNA_10, 5'DR-sgRNA_1~5'DR-sgRNA_5) and the editing activity of Cas13T_1 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 34In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0133] Figure 35 The relationship between guide RNAs with different structures (5'DR-sgRNA_6~5'DR-sgRNA_8, 5'DR-sgRNA_10) and the editing activity of Cas13T_1 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells) was shown. Figure 35 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0134] Figure 36 The relationship between the editing activities of guide RNAs with different structures (3'DR-sgRNA_1~3'DR-sgRNA_4) and Cas13T_4 on target RNA was shown (i.e. the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 36 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0135] Figure 37 The relationship between guide RNAs with different structures (3'DR-sgRNA_5~3'DR-sgRNA_8, 3'DR-sgRNA_10) and the editing activity of Cas13T_4 on target RNA (i.e. the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells) was shown. Figure 37 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0136] Figure 38 The relationship between guide RNAs with different structures (5'DR-sgRNA_1~5'DR-sgRNA_4) and the editing activity of Cas13T_5 on target RNA was shown (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 38 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0137] Figure 39 The relationship between guide RNAs with different structures (5'DR-sgRNA_5~5'DR-sgRNA_8, 5'DR-sgRNA_10) and the editing activity of Cas13T_5 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells) was shown. Figure 39 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0138] Figure 40 The relationship between guide RNAs with different structures (5'DR-sgRNA_1~5'DR-sgRNA_3, 5'DR-sgRNA_5) and the editing activity of Cas13T_6 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells) was shown. Figure 40 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0139] Figure 41 The relationship between guide RNAs with different structures (5'DR-sgRNA_6~5'DR-sgRNA_8, 5'DR-sgRNA_10) and the editing activity of Cas13T_6 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 41 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0140] Figure 42 The relationship between guide RNAs with different structures (5'DR-sgRNA_1~5'DR-sgRNA_4) and the editing activity of Cas13T_7 on target RNA was shown (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 42 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0141] Figure 43 The relationship between guide RNAs with different structures (5'DR-sgRNA_5~5'DR-sgRNA_6, 5'DR-sgRNA_8, 5'DR-sgRNA_10) and the editing activity of Cas13T_7 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 43 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0142] Figure 44 The relationship between guide RNAs with different structures (5'DR-sgRNA_1~5'DR-sgRNA_4) and the editing activity of Cas13T_8 on target RNA was shown (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells). Figure 44 In the graph, the vertical axis represents cell count, and the horizontal axis represents relative fluorescence intensity.
[0143] Figure 45The relationship between the editing activities of guide RNAs with different structures (5'DR-sgRNA_5, 5'DR-sgRNA_7~5'DR-sgRNA_8, 5'DR-sgRNA_10) and Cas13T_8 on target RNA (i.e., the knockdown of the reporter gene mCherry expression in transfected human HEK293T cells) was shown. Figure 45 In the graph, the vertical axis represents the count, and the horizontal axis represents the relative fluorescence intensity.
[0144] Figure 46 The relationship between different spacer sequences and the editing activity of the Cas13a_2 protein on target RNA was shown (i.e., the mCherry knockdown efficiency of transfected human HEK293T cells represents the Cas13a_2 protein-specific nuclease activity).
[0145] Figure 47 The relationship between different spacer sequences and different crRNA (sgRNA) structures and the editing activity of Cas13b_1 protein on target RNA was shown (i.e., the mCherry knockdown efficiency of transfected human HEK293T cells represents the Cas13b_1 protein-specific nuclease activity).
[0146] Figure 48 The relationship between different spacer sequences and the editing activity of the Cas13b_2 protein on target RNA was shown (i.e., the mCherry knockdown efficiency of transfected human HEK293T cells represents the Cas13b_2 protein-specific nuclease activity).
[0147] Figure 49 The relationship between different spacer sequences and the editing activity of the Cas13d_2 protein on target RNA was shown (i.e., the mCherry knockdown efficiency of transfected human HEK293T cells represents the Cas13d_2 protein-specific nuclease activity).
[0148] Figure 50 The relationship between different spacer sequences and different crRNA (sgRNA) structures and the editing activity of Cas13T_1 protein on target RNA was shown (i.e., the mCherry knockdown efficiency of transfected human HEK293T cells represents the Cas13T_1 protein-specific nuclease activity).
[0149] Figure 51 The relationship between different spacer sequences and the editing activity of the Cas13T_4 protein on target RNA was shown (i.e., the mCherry knockdown efficiency of transfected human HEK293T cells represents the Cas13T_4 protein-specific nuclease activity).
[0150] Figure 52 The relationship between different spacer sequences and the editing activity of the Cas13T_5 protein on target RNA was shown (i.e., the mCherry knockdown efficiency of transfected human HEK293T cells represents the Cas13T_5 protein-specific nuclease activity).
[0151] Figure 53 The relationship between different spacer sequences and the editing activity of the Cas13T_6 protein on target RNA was shown (i.e., the mCherry knockdown efficiency of transfected human HEK293T cells represents the Cas13T_6 protein-specific nuclease activity).
[0152] Figure 54 The relationship between different spacer sequences and the editing activity of Cas13T_7 protein on target RNA was shown (i.e., the mCherry knockdown efficiency of transfected human HEK293T cells represents the Cas13T_7 protein-specific nuclease activity).
[0153] Figure 55 The relationship between different spacer sequences and the editing activity of the Cas13T_8 protein on target RNA was shown (i.e., the mCherry knockdown efficiency of transfected human HEK293T cells represents the Cas13T_8 protein-specific nuclease activity).
[0154] Figure 56 The plasmid map of the mCherry reporter gene plasmid is shown.
[0155] Figure 57 The plasmid map of U6-BbSI-PUC57 used to construct the guide RNA recombinant plasmid is shown.
[0156] Figure 58 The relationship between different spacer sequences and the paracleavage activity of Cas13a_2 on the GFP reporter gene was shown (i.e., the GFP knockdown efficiency of transfected human HEK293T cells represents the non-specific nuclease activity of the Cas13a_2 protein).
[0157] Figure 59 The relationship between different spacer sequences and the paracleavage activity of Cas13b_1 on the GFP reporter gene was shown (i.e., the GFP knockdown efficiency of transfected human HEK293T cells represents the non-specific nuclease activity of the Cas13b_1 protein).
[0158] Figure 60The relationship between different spacer sequences and the paracleavage activity of Cas13b_2 on the GFP reporter gene was shown (i.e., the GFP knockdown efficiency of transfected human HEK293T cells represents the non-specific nuclease activity of the Cas13b_2 protein).
[0159] Figure 61 The relationship between different spacer sequences and the paracleavage activity of Cas13d_2 on the GFP reporter gene was shown (i.e., the GFP knockdown efficiency of transfected human HEK293T cells represents the non-specific nuclease activity of the Cas13d_2 protein).
[0160] Figure 62 The relationship between different spacer sequences and the paracleavage activity of Cas13T_1 on the GFP reporter gene was shown (i.e., the GFP knockdown efficiency of transfected human HEK293T cells represents the non-specific nuclease activity of the Cas13T_1 protein).
[0161] Figure 63 The relationship between different spacer sequences and the paracleavage activity of Cas13T_4 on the GFP reporter gene was shown (i.e., the GFP knockdown efficiency of transfected human HEK293T cells represents the non-specific nuclease activity of the Cas13T_4 protein).
[0162] Figure 64 The relationship between different spacer sequences and the paracleavage activity of Cas13T_5 on the GFP reporter gene was shown (i.e., the GFP knockdown efficiency of transfected human HEK293T cells represents the non-specific nuclease activity of the Cas13T_5 protein).
[0163] Figure 65 The relationship between different spacer sequences and the paracleavage activity of Cas13T_6 on the GFP reporter gene was shown (i.e., the GFP knockdown efficiency of transfected human HEK293T cells represents the non-specific nuclease activity of the Cas13T_6 protein).
[0164] Figure 66 The relationship between different spacer sequences and the paracleavage activity of Cas13T_7 on the GFP reporter gene was shown (i.e., the GFP knockdown efficiency of transfected human HEK293T cells represents the non-specific nuclease activity of the Cas13T_7 protein).
[0165] Figure 67 The relationship between different spacer sequences and the paracleavage activity of Cas13T_8 on the GFP reporter gene was shown (i.e., the GFP knockdown efficiency of transfected human HEK293T cells represents the non-specific nuclease activity of the Cas13T_8 protein). Detailed Implementation
[0166] The following embodiments are provided to better understand the present invention and are not limited to the preferred embodiments described. They do not constitute a limitation on the content and scope of protection of the present invention. Any product that is the same as or similar to the present invention, derived by any person under the guidance of the present invention or by combining the features of the present invention with other prior art, falls within the protection scope of the present invention.
[0167] For any experimental steps or conditions not specified in the following examples, the procedures or conditions described in the literature in this field can be followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional reagent products.
[0168] Example 1-1: Cas13 protein and its encoding polynucleotide
[0169] This embodiment provides a Cas13 protein, the amino acid sequence of which is shown in SEQ ID NO.1. It belongs to type 2 VI Cas protein and is named Cas13a_2. The nucleotide sequence of the polynucleotide encoding Cas13a_2 is shown in SEQ ID NO.11 (the polynucleotide has been codon optimized).
[0170] Examples 1-2: Cas13 protein and its encoding polynucleotide
[0171] This embodiment provides a Cas13 protein, the amino acid sequence of which is shown in SEQ ID NO.2. It belongs to type 2 VI Cas protein and is named Cas13b_1. The nucleotide sequence of the polynucleotide encoding Cas13b_1 is shown in SEQ ID NO.12 (the polynucleotide has been codon optimized).
[0172] Examples 1-3: Cas13 protein and its encoding polynucleotide
[0173] This embodiment provides a Cas13 protein, the amino acid sequence of which is shown in SEQ ID NO.3. It belongs to type 2 VI Cas protein and is named Cas13b_2. The nucleotide sequence of the polynucleotide encoding the Cas13b_2 is shown in SEQ ID NO.13 (the polynucleotide has been codon optimized).
[0174] Examples 1-4: Cas13 protein and its encoding polynucleotide
[0175] This embodiment provides a Cas13 protein, the amino acid sequence of which is shown in SEQ ID NO.4. It belongs to type 2 VI Cas protein and is named Cas13d_2. The nucleotide sequence of the polynucleotide encoding the Cas13d_2 is shown in SEQ ID NO.14 (the polynucleotide has been codon optimized).
[0176] Examples 1-5: Cas13 protein and its encoding polynucleotide
[0177] This embodiment provides a Cas13 protein, the amino acid sequence of which is shown in SEQ ID NO.5. It belongs to type 2 VI Cas protein and is named Cas13T_1. The nucleotide sequence of the polynucleotide encoding Cas13T_1 is shown in SEQ ID NO.15 (the polynucleotide has been codon optimized).
[0178] Examples 1-6: Cas13 protein and its encoding polynucleotide
[0179] This embodiment provides a Cas13 protein, the amino acid sequence of which is shown in SEQ ID NO.6. It belongs to type 2 VI Cas protein and is named Cas13T_4. The nucleotide sequence of the polynucleotide encoding Cas13T_4 is shown in SEQ ID NO.16 (the polynucleotide has been codon optimized).
[0180] Examples 1-7: Cas13 protein and its encoding polynucleotide
[0181] This embodiment provides a Cas13 protein, the amino acid sequence of which is shown in SEQ ID NO.7. It belongs to type 2 VI Cas protein and is named Cas13T_5. The nucleotide sequence of the polynucleotide encoding Cas13T_5 is shown in SEQ ID NO.17 (the polynucleotide has been codon optimized).
[0182] Examples 1-8: Cas13 protein and its encoding polynucleotide
[0183] This embodiment provides a Cas13 protein, the amino acid sequence of which is shown in SEQ ID NO.8. It belongs to type 2 VI Cas protein and is named Cas13T_6. The nucleotide sequence of the polynucleotide encoding Cas13T_6 is shown in SEQ ID NO.18 (the polynucleotide has been codon optimized).
[0184] Examples 1-9: Cas13 protein and its encoding polynucleotide
[0185] This embodiment provides a Cas13 protein, the amino acid sequence of which is shown in SEQ ID NO.9. It belongs to type 2 VI Cas protein and is named Cas13T_7. The nucleotide sequence of the polynucleotide encoding Cas13T_7 is shown in SEQ ID NO.19 (the polynucleotide has been codon optimized).
[0186] Examples 1-10: Cas13 protein and its encoding polynucleotide
[0187] This embodiment provides a Cas13 protein, the amino acid sequence of which is shown in SEQ ID NO.10. It belongs to type 2 VI Cas protein and is named Cas13T_8. The nucleotide sequence of the polynucleotide encoding Cas13T_8 is shown in SEQ ID NO.20 (the polynucleotide has been codon optimized).
[0188] Example 2-1: CRISPR-Cas System
[0189] This embodiment provides a CRISPR-Cas system, which consists of the Cas13 protein and guide RNA as described in Examples 1-1. The guide RNA comprises a direct repeat (DR) sequence capable of forming a complex with the Cas13 protein and a spacer sequence sufficiently complementary to the target RNA. The guide RNA binds to the Cas13 protein to form a complex and guides the complex to specifically bind to the target RNA. The genomic locus structure of the CRISPR-Cas system is shown in [reference needed]. Figure 1 The Cas13 protein coding sequence is connected to multiple directional repeat (DR) sequences and spacers. The long bar with one pointed end represents the coding sequence of the Cas protein, the short elliptical vertical bar represents the directional repeat (DR) sequence, and the diamond shape represents the spacer sequence.
[0190] Examples 2-2 to 2-10: CRISPR-Cas System
[0191] This embodiment provides a CRISPR / Cas system, which is based on Example 2-1, except that the Cas13 protein in Example 1-1 is replaced with the Cas13 protein in Examples 1-2 to 1-10. The Cas13 proteins in Examples 1-2 to 1-10 are Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8, respectively.
[0192] Experimental Example 1: Classification and Structural Analysis of Cas13 Protein
[0193] Metagenomic analysis, redundancy removal (removal of identical protein sequences), and protein clustering analysis identified 10 Cas13 proteins (Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8) from Examples 1-1 to 1-10. Comparison revealed that these 10 Cas13 proteins have low similarity to previously identified Cas13 proteins. To further determine the category and structure of these 10 Cas13 proteins, this experimental example provides the classification and structural analysis process for these 10 Cas13 proteins:
[0194] Sequence analysis identified direct repeat (DR) sequences associated with Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8. The DNA sequences encoding these DR sequences are shown in Table 1.
[0195] The secondary structure of the homologous repeat (DR) sequences associated with Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8 was analyzed using the RNA secondary structure prediction software RNAfold. The results are shown in […]. Figures 2-11 .
[0196] Multiple sequence alignment of Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8 with previously identified Cas13 family proteins was performed using sequence alignment software, and a phylogenetic tree was constructed. The results are shown in [link to phylogenetic tree]. Figure 12 .
[0197] Functional annotation and multiple sequence alignment were performed on Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8 using publicly available Cas13 protein databases. Conserved regions of these 10 Cas13 protein sequences were analyzed, and the conserved RXXXXH motif positions of these 10 Cas13 proteins were identified. The identification results are shown below. Figure 13 .
[0198] The three-dimensional structures of Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8 were predicted using the protein structure prediction software Alphafold. The prediction results are shown below. Figures 14-23 .
[0199] Depend on Figures 2-11 It can be seen that all DR sequences clearly possess relatively conserved secondary structures. For example, the secondary structures of the DR sequences Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_6, Cas13T_7, and Cas13T_8 are as follows: a 5-base pair (5'-GUUGU-3') stem (with the exception of Cas13T_6, which has an extra unpaired G base), followed by a 6 / 10 nucleotide protrusion (excluding the aforementioned 5-base pair), then a 4 / 5 / 7-base pair stem, and finally a 6 / 8 nucleotide protrusion at the end. The DR sequences associated with Cas13a_2, Cas13b_1, and Cas13T_4 have similar secondary structures, with 3 protrusion structures, while the DR sequence associated with Cas13T_5 lacks a stem structure.
[0200] Depend on Figure 12It can be seen that Cas13a_2 belongs to the Cas13a protein family, Cas13b_1 and Cas13b_2 belong to the Cas13b protein family, Cas13d_2 belongs to the Cas13d protein family, while Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8 are not close to the known Cas13 protein family in the phylogenetic tree, so they are classified as the Cas13T protein family.
[0201] Depend on Figure 13It can be seen that Cas13a_2, Cas13b_1, Cas13b_2, Cas13T_1, Cas13T_4, and Cas13T_8 all contain two RXXXXH motifs. Among them, the two RXXXXH motifs of Cas13a_2, Cas13b_1, Cas13T_4, and Cas13T_8 are located closer to the N-terminus and C-terminus of the Cas protein compared to Cas13b_2 and Cas13T_1. The two RXXXXH motifs of Cas13b_2 are located at the C-terminus of the Cas protein, while the two RXXXXH motifs of Cas13T_1 are located at... The C-terminus of Cas13a_2 is closer to the C-terminus of the Cas protein; Cas13d_2, Cas13T_5, Cas13T_6, and Cas13T_7 each have only one RXXXXH motif. Specifically, the RXXXXH motifs of Cas13d_2 and Cas13T_5 are closer to the C-terminus of the Cas protein, while the RXXXXH motifs of Cas13T_6 and Cas13T_7 are relatively closer to the N-terminus of the Cas protein (the two RXXXXH motifs of Cas13a_2 are located at amino acids 482–487 and 1076–1081, respectively, and the two RXXXXH motifs of Cas13b_1 are located at amino acids 482–487 and 1076–1081, respectively). The two RXXXXH motifs of Cas13b_2 are located at amino acids 115–120 and 1033–1038, respectively; the two RXXXXH motifs of Cas13T_1 are located at amino acids 672–677 and 686–691, respectively; the two RXXXXH motifs of Cas13T_1 are located at amino acids 26–31 and 103–108, respectively; the two RXXXXH motifs of Cas13T_4 are located at amino acids 138–143 and 1214–1219, respectively; and the two RXXXXH motifs of Cas13T_8 are located at amino acids 98–103 and 1118–1123, respectively. The RXXXXH motif of Cas13d_2 is located at amino acids 406-411, the RXXXXH motif of Cas13T_5 is located at amino acids 331-336, the RXXXXH motif of Cas13T_6 is located at amino acids 98-103, and the RXXXXH motif of Cas13T_7 is located at amino acids 113-118. The RXXXXH motif is the catalytic active site of the Cas13 protein. Mutations in this domain may produce nuclease-inactive proteins of the Cas13 protein, while essentially maintaining its ability to bind to guide RNA and target RNA.
[0202] Depend on Figures 14-23It can be seen that although the two RXXXXH motifs of Cas13a_2, Cas13b_1, and Cas13T_8 are located relatively at the N-terminus and C-terminus of the protein, the predicted three-dimensional structures of each protein show that the two RXXXXH motifs of Cas13a_2, Cas13b_1, and Cas13T_8 are very close in three-dimensional structure, while the two RXXXXH motifs of Cas13T_4 are relatively far apart. The two RXXXXH motifs of Cas13b_2 are located at the C-terminus of the Cas protein, while the two RXXXXH motifs of Cas13T_1 are located relatively at the N-terminus of the Cas protein. From a three-dimensional structural perspective, the two RXXXXH motifs of Cas13b_2 and Cas13T_1 are also relatively close.
[0203] Table 1. Coding sequences of direct repeat (DR) sequences associated with different Cas13 proteins.
[0204]
[0205]
[0206] Experimental Example 2: Verification of the cleavage activity of Cas13 protein
[0207] To verify whether the 10 Cas13 proteins (Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8) in Examples 1-1 to 1-10 possess RNase activity, this experimental example provides a verification experiment on the cleavage activity of these 10 Cas13 proteins (see schematic diagram of the experimental procedure). Figure 24 The specific experimental procedure is as follows:
[0208] The codon-optimized polynucleotides encoding Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8 (SEQ ID NO.11–SEQ ID NO.20) were digested with enzymes and then cloned into the p23_puro_NPLS_Cas13_msfGFP_NLS_3xFlag_C plasmid containing the green fluorescent protein (GFP) gene (SEQ ID NO.31). (The nucleotide sequence of the p23_puro_NPLS_Cas13_msfGFP_NLS_3xFlag_C plasmid is shown in SEQ ID NO.11.) As shown in NO.44, the restriction enzyme site is NheI:XbaI, which yields the Cas13 protein recombinant plasmid (in the Cas13 protein recombinant plasmid, the polynucleotides of the Cas13 protein replace the sequences at positions 3472 to 6369 in the p23_puro_NPLS_Cas13_msfGFP_NLS_3xFlag_C plasmid).
[0209] The mCherry reporter gene recombinant plasmid (the nucleotide sequence of the mCherry reporter gene recombinant plasmid is shown in SEQ ID NO. 45, and the plasmid map of the mCherry reporter gene recombinant plasmid is shown in [link to plasmid map]) was obtained. Figure 56 ).
[0210] Spacer_1 to spacer_10 are spacer sequences targeting different positions of mCherry (SEQ ID NO.32), and NT-spacer is a spacer sequence not targeting mCherry. DR_1 to DR_10 are ligated to the 3' or 5' end of spacer_1 to spacer_10, respectively, to obtain guide RNAs targeting the reporter gene mCherry; DR_1 to DR_10 are ligated to the 3' or 5' end of NT-spacer, respectively, to obtain guide RNAs not targeting the reporter gene mCherry. The guide RNAs targeting and not targeting the reporter gene mCherry are cloned into the U6-BbSI-PUC57 plasmid (purchased from Shanghai Sangon Biotech Co., Ltd.; plasmid map of U6-BbSI-PUC57 plasmid can be found in...). Figure 57The guide RNA recombinant plasmid was obtained by digestion at the BbSI restriction site (nucleotide sequence of the guide RNA recombinant plasmid is shown in SEQ ID NO. 46); both the mCherry-targeting and non-mCherry-targeting guide RNAs consist of a direct repeat (DR) sequence and a spacer sequence. Theoretically, the guide RNA precursor in the CRISPR / Cas13 system can produce two types of guide RNA structures during maturation: a direct repeat (DR) sequence + spacer sequence (i.e., 5'DR) and a spacer sequence + direct repeat (DR) sequence (i.e., 3'DR). The structures of the mCherry-targeting and non-mCherry-targeting guide RNAs used for the 10 Cas13 proteins are shown in Table 3.
[0211] Human HEK293T cells (purchased from ATCC) were used at a rate of 1×10⁻⁶. 4 After seeding the cells in 48-well plates, the Cas13 protein recombinant plasmid, guide RNA recombinant plasmid (containing guide RNA targeting the reporter gene mCherry), and mCherry reporter gene plasmid were co-transfected into human HEK293T cells using LIPOFECTAMINE3000 and P3000™ reagents (purchased from ThermoFisher, L3000015) to obtain the experimental group of transfected human HEK293T cells. The Cas13 protein recombinant plasmid, the guide RNA recombinant plasmid (containing guide RNA that does not target the reporter gene mCherry), and the mCherry reporter gene plasmid were co-transfected into human HEK293T cells to obtain the negative control group of transfected human HEK293T cells; human HEK293T cells that were not transfected with any plasmid served as the blank control group. The experimental group of transfected human HEK293T cells, the negative control group of transfected human HEK293T cells, and the blank control group were cultured in a cell culture incubator at 37°C and 5% (v / v) CO2.
[0212] After 48 hours of culture, microscopic observation revealed that the experimental group of transfected human HEK293T cells showed similar growth and morphology to the negative control group. The transfected human HEK293T cells were observed under a fluorescence microscope, and finally, flow cytometry was used for detection and quantitative analysis. The results are shown below. Figures 25-55Flow cytometry analysis revealed that in the negative control group, guide RNA targeting the non-reporter gene mCherry failed to knock down mCherry fluorescence signal expression. Compared with the negative control group, the experimental group showed a significant decrease in mCherry fluorescence signal intensity in human HEK293T cells transfected with mCherry. This indicates that all 10 Cas13 proteins can effectively knock down mCherry mRNA levels using guide RNA targeting mCherry, thereby reducing mCherry protein expression (see...). Figures 25-55 ).
[0213] Furthermore, flow cytometry (FCM) analysis revealed that the degree of fluorescence signal intensity reduction varied when spacer sequences targeting different positions on the mCherry plasmid were used. This indicates that the spacer sequences targeting different positions on the target RNA affect the editing effect of the Cas13 protein (see [link to FCM analysis]). Figures 25-55 ).
[0214] In addition, by Figures 46-55 It can be seen that Cas13b_1 and Cas13T_1 exhibit a clear preference for the two guide RNA structures. Specifically, when Cas13b_1 and Cas13T_1 proteins are guided by 3'DR crRNA, the mCherry corresponding to certain spacer sequences is knocked down more effectively. However, for other Cas13 proteins, the guide RNAs for 3'DR and 5'DR are not significantly different.
[0215] Table 2. Spacer sequences targeting and not targeting mCherry.
[0216]
[0217]
[0218] Table 3 Guide RNA
[0219]
[0220] Experiment Example 3: Verification of the paracleavage activity (i.e., non-specific cleavage activity) of Cas13 protein
[0221] To verify whether the 10 Cas13 proteins (Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8) in Examples 1-1 to 1-10 possess paracleavage activity (i.e., non-targeted nuclease activity), this experimental example provides a paracleavage activity verification experiment for these 10 Cas13 proteins. The experimental procedure is as follows:
[0222] Using the same guide RNA targeting mCherry as in Experiment 2, and repeating the experiment of Experiment 2, the knockdown efficiency of the GFP reporter gene in cells 48 hours after transfection was analyzed by flow cytometry. The knockdown efficiency of GFP represents the non-specific cleavage activity of Cas13. The experimental results are shown in [Figure 1]. Figures 58-67 .
[0223] from Figures 58-67 It can be seen that all 10 Cas13 proteins, namely Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8, exhibit non-specific cleavage activity. This indicates that these 10 Cas13 proteins can be used for RNA molecule detection. Among them, for Cas13b_1 and Cas13T_1, the 3'DR and 5'DR affect the paracleavage activity of the above two Cas13 proteins. Furthermore, for the 10 Cas13 proteins, namely Cas13a_2, Cas13b_1, Cas13b_2, Cas13d_2, Cas13T_1, Cas13T_4, Cas13T_5, Cas13T_6, Cas13T_7, and Cas13T_8, the guide RNA corresponding to different spacer sequences also affects the level of paracleavage activity.
[0224] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A Cas13 protein, characterized in that, The amino acid sequence of the Cas13 protein is shown in SEQ ID NO.
1.
2. A CRISPR-Cas system, characterized in that, The CRISPR-Cas system comprises the Cas13 protein of claim 1 and a guide RNA; the guide RNA is capable of binding to the Cas13 protein to form a complex and guiding the complex to contact the target RNA.
3. The CRISPR-Cas system as described in claim 2, characterized in that, The guide RNA comprises a spacer sequence and a direct repeat sequence; the spacer sequence is 15–60 nucleotides in size; the direct repeat sequence is 15–40 nucleotides in size.
4. The CRISPR-Cas system as described in claim 3, characterized in that, The spacer sequence is complementary to the target RNA by at least 15 nucleotides.
5. A polynucleotide, characterized in that, The polynucleotide encodes the Cas13 protein of claim 1, or the polynucleotide encodes the CRISPR-Cas system of any one of claims 2 to 4.
6. A carrier, characterized in that, The vector carries the polynucleotide of claim 5.
7. The carrier as described in claim 6, characterized in that, The vector carries a polynucleotide linker regulatory element.
8. The carrier as described in claim 7, characterized in that, The regulatory element includes at least one of a promoter, enhancer, terminator, protein degradation signal, leader sequence, or polyadenylate sequence.
9. The carrier as described in claim 7, characterized in that, The regulatory element includes a polyadenylation signal.
10. The carrier according to any one of claims 6 to 9, characterized in that, The vector includes at least one of a viral vector or a non-viral vector.
11. The carrier as described in claim 10, characterized in that, The viral vector includes at least one of the following: retroviral vector, bacteriophage vector, adenovirus vector, adeno-associated virus vector, vaccinia virus vector, hybrid virus vector, baculovirus vector, or herpes simplex virus vector.
12. The carrier as described in claim 10, characterized in that, The non-viral vectors include plasmid vectors.
13. A delivery system, characterized in that, The delivery system carries the Cas13 protein of claim 1, the CRISPR-Cas system of any one of claims 2 to 4, the polynucleotide of claim 5, or the vector of any one of claims 6 to 12.
14. The delivery system as claimed in claim 13, characterized in that, The delivery carrier of the delivery system includes at least one of nanoparticle carriers or viral carriers.
15. The delivery system as claimed in claim 14, characterized in that, The nanoparticle carrier includes at least one of liposome carriers, chitosan carriers, polymer carriers, exosome carriers, or microbubble carriers.
16. A host cell, characterized in that, The host cell carries the Cas13 protein of claim 1, the CRISPR-Cas system of any one of claims 2 to 4, the polynucleotide of claim 5, or the vector of any one of claims 6 to 12; the host cell does not possess totipotency.
17. A method for constructing an animal model, characterized in that, The method for constructing the animal model comprises: performing gene editing on the animal using the Cas13 protein of claim 1, the CRISPR-Cas system of any one of claims 2 to 4, the polynucleotide of claim 5, the vector of any one of claims 6 to 12, or the host cell of claim 16.
18. A method for gene editing, characterized in that, The method is not for the purpose of disease diagnosis and treatment. The method includes contacting the target RNA with the Cas13 protein of claim 1, the CRISPR-Cas system of any one of claims 2 to 4, the polynucleotide of claim 5, the vector of any one of claims 6 to 12, or the host cell of claim 16.
19. The method as described in claim 18, characterized in that, The guide RNA binds to the Cas13 protein to form a complex, and guides the complex to contact the target RNA for gene editing.
20. The application of the Cas13 protein of claim 1, or the CRISPR-Cas system of any one of claims 2-4, or the polynucleotide of claim 5, or the vector of any one of claims 6-12, or the delivery system of any one of claims 13-15, or the host cell of claim 16, or the method of claim 18 or 19, in gene editing, characterized in that, The application is for purposes other than disease diagnosis and treatment.
21. A reagent kit, characterized in that, The kit comprises the Cas13 protein of claim 1, the CRISPR-Cas system of any one of claims 2 to 4, the polynucleotide of claim 5, the vector of any one of claims 6 to 12, the delivery system of any one of claims 13 to 15, or the host cell of claim 16.
22. A nucleic acid detection method, characterized in that, The method is not for the purpose of disease diagnosis and treatment; the method is to perform nucleic acid detection on the sample to be tested using the kit described in claim 21.
23. The application of the Cas13 protein of claim 1, or the CRISPR-Cas system of any one of claims 2-4, or the polynucleotide of claim 5, or the vector of any one of claims 6-12, or the delivery system of any one of claims 13-15, or the host cell of claim 16, or the kit of claim 21, or the method of claim 22 in nucleic acid detection, characterized in that, The application is for purposes other than disease diagnosis and treatment.
24. The use of the Cas13 protein of claim 1, or the CRISPR-Cas system of any one of claims 2 to 4, or the polynucleotide of claim 5, or the vector of any one of claims 6 to 12, or the delivery system of any one of claims 13 to 15, or the host cell of claim 16, in the preparation of a medicament for treating a condition or disease of a subject in need.
Citation Information
Patent Citations
VI-B type CRISPR / Cas13 gene editing system and application thereof
CN112430586A