Type VI-E and VI-F CRISPR-Cas systems and their use

Smaller and less active Cas13e and Cas13f proteins address the packaging and collateral RNase issues of larger Cas proteins, providing efficient and precise RNA editing for gene therapy.

JP7709511B2Active Publication Date: 2025-07-16HUIDAGENE THERAPEUTICS (SINGAPORE) PTE LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023218259
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-07-16
Estimated Expiration
2040-02-28

AI Technical Summary

Technical Problem

Current Class 2, type VI CRISPR-Cas proteins, such as Cas13a, Cas13b, and Cas13d, are too large for efficient packaging into low-capacity gene therapy vectors like AAV and exhibit non-specific/collateral RNase activity, posing risks for gene therapy applications.

Method used

Development of smaller Class 2, type VI CRISPR-Cas proteins, Cas13e and Cas13f, with reduced non-specific RNase activity, allowing efficient packaging into AAV vectors and enabling precise RNA targeting and editing.

Benefits of technology

Cas13e and Cas13f proteins are more potent in RNA targeting and single-base editing with minimal off-target effects, suitable for gene therapy applications, reducing the risk of collateral damage to cellular RNA.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007709511000025
    Figure 0007709511000025
  • Figure 0007709511000026
    Figure 0007709511000026
  • Figure 0007709511000027
    Figure 0007709511000027
Patent Text Reader

Abstract

To provide novel CRISPR / Cas compositions and uses thereof for targeting nucleic acids.SOLUTION: The invention provides non-naturally occurring or engineered RNA-targeting systems, comprising a novel RNA-targeting Cas13e or Cas13f effector protein, and at least one targeting nucleic acid component, such as a guide RNA (gRNA) or crRNA. The novel Cas effector protein is one of the smallest known Cas effector proteins, at about 800 amino acids in size, and are thus uniquely suitable for delivery using vectors of small capacity, such as an AAV vector.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] CRISPR (clustered regularly interspaced short palindromic repeats) is a family of DNA sequences found within the genomes of prokaryotes, such as bacteria and archaea. These sequences are understood to be derived from DNA fragments of bacteriophages that previously infected the prokaryote, and are used to detect and destroy DNA from similar bacteriophages during subsequent infections of the prokaryote.

[0002] The CRISPR-associated system is a set of homologous genes, or Cas genes, some of which encode Cas proteins with helicase and nuclease activities. The Cas protein is an enzyme that uses RNA derived from the CRISPR sequence (crRNA) as a guide sequence to recognize and cleave a specific strand of a polynucleotide (e.g., DNA) that is complementary to the crRNA.

[0003] Together, the CRISPR-Cas system constitutes a primitive prokaryotic "immune system" that confers resistance or acquired immunity against foreign pathogenic genetic elements, such as extrachromosomal DNA (e.g., plasmids) and bacteriophages, or genetic elements present within foreign RNA encoded by foreign DNA.

[0004] Originally, the CRISPR / Cas system appears to be a broad-spectrum prokaryotic defense mechanism against foreign genetic material and is found in approximately 50% of sequenced bacterial genomes and nearly 90% of sequenced archaea. This prokaryotic system has since been developed to form the basis of a technology known as CRISPR-Cas, which has found widespread use in a variety of applications, including basic biological research, the development of biotechnology products, and the treatment of diseases, in a number of eukaryotes, including humans.

[0005] Prokaryotic CRISPR-Cas systems include a very diverse group of protein effectors, non-coding elements, and locus architectures, and some examples of these have been engineered and configured to yield important biotechnologies.

[0006] CRISPR locus structures have been studied in many systems. In such systems, the CRISPR array in genomic DNA typically contains short DR sequences separated by unique spacer sequences, following an AT-rich leader sequence. This CRISPR DR sequence typically ranges in size from 28 to 37 bp, but can range from 23 to 55 bp. Some DR sequences exhibit dyad symmetry, which implies the formation of secondary structures such as stem-loops (“hairpins”) in RNA. There also appear to be some that are not structured. The size of spacers within various CRISPR arrays typically ranges from 32 to 38 bp (in the range of 21 to 72 bp). Usually, there are less than 50 repeat-spacer sequences within a CRISPR array.

[0007] A small cluster of cas genes is often found adjacent to such CRISPR repeat-spacer arrays. To date, the 93 identified cas genes have been classified into 35 families based on the sequence similarity of the encoded proteins. Eleven of the 35 families form the so-called cas core, which includes the protein families Cas1 to Cas9. A complete CRISPR-Cas locus has at least one gene belonging to the cas cluster.

[0008] The CRISPR-Cas system can be roughly divided into two classes - Class 1 systems use a complex of multiple Cas proteins to degrade foreign nucleic acids, while Class 2 systems use a single large Cas protein for the same purpose. The single-subunit effector compositions of Class 2 systems provide a simpler set of components for engineering and application translation, and have thus far been an important source for the discovery, engineering, and optimization of new and powerful programmable technologies for genome engineering and beyond.

[0009] Class 1 systems are further divided into types I, III, and IV; Class 2 systems are divided into types II, V, and VI. These six system types are in addition divided into 19 subtypes. The classification is also based on the complement of cas genes present. Most CRISPR-Cas systems have a Cas1 protein. Many prokaryotes contain multiple CRISPR-Cas systems. This suggests that multiple CRISPR-Cas systems can be compatible and share components.

[0010] One of the best-characterized class 2, type II Cas proteins, Cas9 is a typical member of class 2, type II, and is derived from Streptococcus pyogenes (SpCas9). Cas9 is a DNA endonuclease activated by a short crRNA molecule that complements the target DNA sequence and a separate trans-activating CRISPR RNA (tracrRNA). The crRNA consists of a direct repeat (DR) sequence that bears the protein binding to the crRNA and a spacer sequence, and these sequences can be engineered to be complementary to any desired nucleic acid target sequence. Thus, the CRISPR system can be programmed to target DNA or RNA targets by modifying the spacer sequence of the crRNA. The crRNA and tracrRNA have been fused to form a single guide RNA (sgRNA) for better practical utility. When combined with Cas9, the sgRNA hybridizes to its target DNA, guiding Cas9 to cleave the target DNA. Other Cas9 effector proteins from other species, including Cas9 from the S. thermophilus CRISPR system, have been identified and are similarly used. These CRISPR / Cas9 systems are widely used in a number of eukaryotes, including the budding yeast Saccharomyces cerevisiae, the opportunistic pathogen Candida albicans, zebrafish (Danio rerio), fruit flies (Drosophila melanogaster), ants (Harpegnathos saltator and Ooceraea biroi), mosquitoes (Aedes aegypti), nematodes (Caenorhabditis elegans), plants, mice, monkeys, and human embryos.

[0011] Another recently characterized Cas effector protein is Cas12a (formerly known as Cpf1). Cas12a, along with C2c1 and C2c3, is a member of the class 2, type V Cas proteins that lack the HNH nuclease but have RuvC nuclease activity. Cas12a was first characterized in the CRISPR / Cpf1 system of the bacterium Francisella novicida. Its original name reflects the prevalence of that CRISPR-Cas subtype in Prevotella and Francisella lineages. Cas12a has shown several major differences from Cas9: it causes "staggered" cuts within double-stranded DNA, as opposed to the "blunt" cuts brought about by Cas9; it targets a "T-rich" PAM sequence (which provides an alternative targeting site to Cas9); and it requires only the CRISPR RNA (crRNA) for successful targeting and not the tracrRNA. The smaller crRNA of Cas12a is more suitable for multiplex genome editing than Cas9, as more can be packaged within one vector than the sgRNA of Cas9. Additionally, the sticky 5' overhangs left by Cas12a can be used for DNA assembly that is much more target-specific than conventional restriction enzyme cloning. Finally, Cas12a cuts DNA 18 - 23 base pairs downstream from its PAM site. This means that, unlike Cas9 cleavage sequences which have only 3 base pairs upstream of the PAM site and the NHEJ pathway typically prevents further rounds of cleavage by causing indel mutations that disrupt the recognition sequence, Cas12a allows for multiple rounds of DNA cleavage after DNA repair following the creation of double-strand breaks (DSBs) by the NHEJ system because there are no interruptions in the nuclease recognition sequence after DNA repair. Theoretically, repeated rounds of DNA cleavage are associated with an increased likelihood of the desired genome editing occurring. showed several major differences: it causes "staggered" cuts within double-stranded DNA, as opposed to the "blunt" cuts brought about by Cas9; it targets a "T-rich" PAM sequence (which provides an alternative targeting site to Cas9); and it requires only the CRISPR RNA (crRNA) for successful targeting and not the tracrRNA. The smaller crRNA of Cas12a is more suitable for multiplex genome editing than Cas9, as more can be packaged within one vector than the sgRNA of Cas9. Additionally, the sticky 5' overhangs left by Cas12a can be used for DNA assembly that is much more target-specific than conventional restriction enzyme cloning. Finally, Cas12a cuts DNA 18 - 23 base pairs downstream from its PAM site. This means that, unlike Cas9 cleavage sequences which have only 3 base pairs upstream of the PAM site and the NHEJ pathway typically prevents further rounds of cleavage by causing indel mutations that disrupt the recognition sequence, Cas12a allows for multiple rounds of DNA cleavage after DNA repair following the creation of double-strand breaks (DSBs) by the NHEJ system because there are no interruptions in the nuclease recognition sequence after DNA repair. Theoretically, repeated rounds of DNA cleavage are associated with an increased likelihood of the desired genome editing occurring.

[0012] More recently, several class 2, type VI Cas proteins, including Cas13 (also known as C2c2), Cas13b, Cas13c, and Cas13d, have been identified. Each is an RNA-guided RNase (i.e., these Cas proteins recognize target RNA sequences rather than target DNA sequences using crRNA in Cas9 and Cas12a). Overall, the CRISPR / Cas13 system can achieve higher RNA digestion efficiency compared to conventional RNAi and CRISPRi technologies, while showing considerably less off-target cleavage compared to RNAi.

[0013] One drawback of these currently identified Cas13 proteins is their relatively large size. Cas13a, Cas13b, and Cas13c each have more than 1100 amino acid residues. Therefore, packaging the coding sequence (about 3.3 kb) and sgRNA, plus any required promoter and translation regulatory sequences, into a particular small-capacity gene therapy vector, such as the currently most effective and safest gene therapy vector based on adeno-associated virus (AAV) with a packaging capacity of about 4.7 kb, is difficult if possible at all. Only Cas13d, the smallest Cas13 protein to date, has about 920 amino acids (i.e., a coding sequence of about 2.8 kb) and can theoretically be packaged into an AAV vector, but its use in single-base editing-based gene therapy, which relies on using Cas13d-based fusion proteins with single-base editing functions such as dCas13d-ADAR2DD (having a coding sequence of about 3.9 kb), is restricted.

[0014] Furthermore, all currently known Cas13 proteins / systems have non-specific / collateral RNase activity immediately after activation by crRNA-based target sequence recognition. This activity is particularly strong in Cas13a and Cas13b and is still detectably present in Cas13d. While this property can be advantageously used in nucleic acid detection methods, the non-specific / collateral RNase activity of these Cas13 proteins poses a significant potential risk for gene therapy use.

Summary of the Invention

Means for Solving the Problems

[0015] One aspect of the present invention is: (1) an RNA guide sequence comprising a spacer sequence capable of hybridizing to a target RNA and a direct repeat (DR) sequence on the 3'-side of the spacer sequence; and (2) a CRI having an amino acid sequence of any one of SEQ ID NOs: 1 to 7 associated with a SPR protein (Cas), or a derivative or functional fragment of said Cas; Cas, the derivative and functional fragment of said Cas are capable of (i) binding to the RNA guide sequence and (ii) targeting the target RNA, provided that the spacer sequence is not 100% complementary to a naturally occurring bacteriophage nucleic acid when the complex contains any one of the Cas of SEQ ID NOs: 1 to 7, or the target RNA is encoded by eukaryotic DNA, to provide a clustered regularly interspaced short palindromic repeat (CRISPR)-Cas complex.

[0016] In certain embodiments, the DR sequence has a secondary structure substantially the same as any one of the secondary structures of SEQ ID NOs: 8 to 14.

[0017] In certain embodiments, the DR sequence is encoded by any one of SEQ ID NOs: 8 to 14.

[0018] In certain embodiments, the target RNA is encoded by eukaryotic DNA.

[0019] In certain embodiments, the eukaryotic DNA is non-human mammalian DNA, non-human primate DNA, human DNA, plant DNA, insect DNA, avian DNA, reptilian DNA, rodent DNA, fish DNA, worm / nematode DNA, yeast DNA.

[0020] In certain embodiments, the target RNA is mRNA.

[0021] In certain embodiments, the spacer sequence is 15 - 55 nucleotides, 25 - 35 nucleotides, or about 30 nucleotides.

[0022] In certain embodiments, the spacer sequence is 90 - 100% complementary to the target RNA.

[0023] In certain embodiments, the derivative comprises a conservative amino acid substitution of one or more residues of any one of SEQ ID NOs: 1 - 7.

[0024] In certain embodiments, the derivative comprises only conservative amino acid substitutions.

[0025] In certain embodiments, the derivative has the same sequence as any one of the wild-type Cas of SEQ ID NOs: 1 - 7 within the HEPN domain or the RXXXXH motif.

[0026] In certain embodiments, the derivative can bind to an RNA guide sequence that hybridizes to the target RNA, but does not have RNase catalytic activity due to mutations within the RNase catalytic site of Cas.

[0027] In certain embodiments, the derivative has an N-terminal deletion of 210 residues or less and / or a C-terminal deletion of 180 residues or less.

[0028] In certain embodiments, the derivative has an N-terminal deletion of about 180 residues and / or a C-terminal deletion of about 150 residues.

[0029] In certain embodiments, the derivative further comprises an RNA base-editing domain.

[0030] In certain embodiments, the RNA base-editing domain is an adenosine deaminase, such as a double-stranded RNA-specific adenosine deaminase (e.g., ADAR1 or ADAR2) ; apolipoprotein B mRNA editing enzyme; catalytic polypeptide-like (APOBEC); or activation-induced cytidine deaminase (AID).

[0031] In certain embodiments, ADAR has an E488Q / T375G double mutation or is ADAR2DD.

[0032] In certain embodiments, the base-editing domain is further fused to an RNA-binding domain, such as MS2.

[0033] In certain embodiments, the derivative further comprises an RNA methyltransferase, an RNA demethylase, an RNA splicing modifier, a localization factor, or a translation modifier.

[0034] In certain embodiments, Cas, the derivative, or the functional fragment comprises a nuclear localization signal (NLS) sequence or a nuclear export signal (NES).

[0035] In certain embodiments, targeting of the target RNA results in modification of the target RNA.

[0036] In certain embodiments, the modification of the target RNA is cleavage of the target RNA.

[0037] In certain embodiments, the modification of the target RNA is deamination of adenosine (A) to inosine (I).

[0038] In certain embodiments, the CRISPR-Cas complex of the invention further comprises a target RNA comprising a sequence capable of hybridizing to the spacer sequence.

[0039] Another aspect of the present invention provides a fusion protein comprising (1) a Cas of the present invention, a derivative thereof, or a functional fragment thereof, and (2) a heterologous functional domain.

[0040] In certain embodiments, the heterologous functional domain is: a nuclear localization signal (NLS), a reporter protein or detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP), a localization signal, a protein targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD), an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc.), a transcriptional activation domain (e.g., VP64 or VPR), a transcriptional repression domain (e.g., the KRAB moiety or the SID moiety), a nuclease (e.g., FokI), a deaminase domain (e.g., ADAR1, ADAR2, APOBEC, AID, or TAD), a methylase, a demethylase, a transcriptional release factor, an HDAC, a polypeptide having ssRNA cleavage activity, a polypeptide having dsRNA cleavage activity, a polypeptide having ssDNA cleavage activity, a polypeptide having dsDNA cleavage activity, a DNA or RNA ligase, or any combination thereof.

[0041] In certain embodiments, the heterologous functional domain is fused N-terminally, C-terminally, or internally within the fusion protein.

[0042] Another aspect of the present invention provides a conjugate comprising (1) a Cas of the present invention, a derivative thereof, or a functional fragment thereof, conjugated to (2) a heterologous functional moiety.

[0043] In certain embodiments, the heterologous functional moiety is: a nuclear localization signal (NLS), a reporter protein or detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP), localization signals, protein targeting moieties, DNA binding domains (e.g., MBP, Lex A DBD, Gal4 DBD), epitope tags (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc.), transcriptional activation domains (e.g., VP64 or VPR), transcriptional repression domains (e.g., KRAB moiety or SID moiety), nucleases (e.g., FokI), deamination domains (e.g., ADAR1, ADAR2, APOBEC, AID, or TAD), methylases, demethylases, transcriptional release factors, HDACs, polypeptides having ssRNA cleavage activity, polypeptides having dsRNA cleavage activity, polypeptides having ssDNA cleavage activity, polypeptides having dsDNA cleavage activity, DNA or RNA ligases, or any combination thereof.

[0044] In certain embodiments, the heterologous functional moiety is conjugated to Cas, its derivatives, or its functional fragments at the N-terminus, C-terminus, or internally.

[0045] Another aspect of the present invention provides a polynucleotide encoding any one of SEQ ID NOs: 1-7, or its derivative, or its functional fragment, or its fusion protein, provided that the polynucleotide is not any of SEQ ID NOs: 15-21.

[0046] In certain embodiments, the polynucleotide is codon-optimized for expression in cells.

[0047] In certain embodiments, the cell is a eukaryotic cell.

[0048] Another aspect of the present invention provides an unnatural polynucleotide comprising a derivative of any one of SEQ ID NOs: 8-14, wherein the derivative has (i) an addition, deletion, or substitution of one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) nucleotides compared to any one of SEQ ID NOs: 8-14; (ii) at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 97% sequence identity to any one of SEQ ID NOs: 8-14; (iii) hybridizes to any one of SEQ ID NOs: 8-14, or any one of (i) and (ii) under stringent conditions; or (iv) is a complement of any one of (i)-(iii), provided that the derivative is not any of SEQ ID NOs: 8-14 and the derivative encodes (or is) an RNA that maintains a substantially the same secondary structure as any of the RNAs encoded by SEQ ID NOs: 8-14.

[0049] In certain embodiments, the derivative functions as a DR sequence of any one of the Cas, its derivatives, or its functional fragments of the present invention.

[0050] Another aspect of the present invention provides a vector comprising the polynucleotide of the present invention.

[0051] In certain embodiments, the polynucleotide is operably linked to a promoter and optionally an enhancer.

[0052] In certain embodiments, the promoter is a constitutive promoter, an inducible promoter, a ubiquitous promoter, or a tissue-specific promoter.

[0053] In certain embodiments, the vector is a plasmid.

[0054] In certain embodiments, the vector is a retroviral vector, a phage vector, an adenoviral vector, a herpes simplex virus (HSV) vector, an AAV vector, or a lentiviral vector.

[0055] In certain embodiments, the AAV vector is a recombinant AAV vector of serotype AAV1, AAV2, AAV4, AAV5, AAV6, AAV7, AAVrh74, AAV8, AAV9, AAV10, AAV11, AAV12, or AAV13.

[0056] Another aspect of the invention provides a delivery system comprising (1) a delivery vehicle and (2) a CRISPR-Cas complex of the invention, a fusion protein of the invention, a conjugate of the invention, a polynucleotide of the invention, or a vector of the invention.

[0057] In certain embodiments, the delivery vehicle is a nanoparticle, liposome, exosome, microvesicle, or gene gun.

[0058] Another aspect of the invention provides a cell or its progeny comprising a CRISPR-Cas complex of the invention, a fusion protein of the invention, a conjugate of the invention, a polynucleotide of the invention, or a vector of the invention.

[0059] In certain embodiments, the cell or its progeny is a eukaryotic cell (e.g., a non-human mammalian cell, a human cell, or a plant cell), or a prokaryotic cell (e.g., a bacterial cell).

[0060] Another aspect of the invention provides a non-human multicellular eukaryote comprising a cell of the invention.

[0061] In certain embodiments, the non-human multicellular eukaryote is an animal (e.g., a rodent or a primate) model for a human genetic disorder.

[0062] Another aspect of the present invention provides a method for modifying a target RNA, the method comprising contacting the target RNA with a CRISPR-Cas complex of the present invention, wherein the spacer sequence is complementary to at least 15 nucleotides of the target RNA; the Cas, derivative, or functional fragment binds to the RNA guide sequence to form a complex; the complex binds to the target RNA; and binding of the complex to the target RNA causes the Cas, derivative, or functional fragment to modify the target RNA.

[0063] In certain embodiments, the target RNA is modified by cleavage by Cas.

[0064] In certain embodiments, the target RNA is modified by deamination by a derivative comprising a double-stranded RNA-specific adenosine deaminase.

[0065] In certain embodiments, the target RNA is mRNA, tRNA, rRNA, non-coding RNA, lncRNA, or nuclear RNA.

[0066] In certain embodiments, binding of the complex to the target RNA causes the Cas, derivative, and functional fragment to exhibit no substantial (or detectable) collateral RNase activity.

[0067] In certain embodiments, the target RNA is intracellular.

[0068] In certain embodiments, the cell is a cancer cell.

[0069] In certain embodiments, the cell is infected with an infectious agent.

[0070] In certain embodiments, the infectious agent is a virus, prion, protozoan, fungus, or parasite worm.

[0071] In certain embodiments, the CRISPR-Cas complex is encoded by a first polynucleotide encoding any one of SEQ ID NOs: 1-7, or a derivative or functional fragment thereof, and a second polynucleotide comprising any one of SEQ ID NOs: 8-14 and a sequence encoding a spacer RNA capable of binding to a target RNA, and the first and second polynucleotides are introduced into a cell.

[0072] In certain embodiments, the first and second polynucleotides are introduced into a cell by the same vector.

[0073] In certain embodiments, the method causes one or more of the following: (i) in vitro or in vivo induction of cell senescence; (ii) in vitro or in vivo cell cycle arrest; (iii) in vitro or in vivo inhibition of cell growth and / or cell proliferation; (iv) in vitro or in vitro induction of anergy; (v) in vitro or in vitro induction of apoptosis; and (vi) in vitro or in vitro induction of necrosis.

[0074] Another aspect of the present invention provides a method for treating a receptor or a disease in a subject in need thereof, the method comprising administering to the subject a composition comprising the CRISPR-Cas complex of the present invention, or a polynucleotide encoding the same; the spacer sequence is complementary to at least 15 nucleotides of a target RNA associated with the receptor or the disease; the Cas, derivative, or functional fragment binds to the RNA guide sequence to form a complex; the complex binds to the target RNA; and binding of the complex to the target RNA causes the Cas, derivative, or functional fragment to cleave the target RNA to treat the receptor or the disease in the subject.

[0075] In certain embodiments, the receptor or disease is cancer or an infectious disease.

[0076] In certain embodiments, the cancer is Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, or bladder cancer.

[0077] In certain embodiments, the method is an in vitro method, an in vivo method, or an ex vivo method.

[0078] Another aspect of the invention provides a cell or its progeny obtained by the method of the invention, wherein the cell and its progeny comprise a modification not found in nature (e.g., a modification not found in nature in the transcribed RNA of the cell / progeny).

[0079] Another aspect of the invention provides a method for detecting the presence of a target RNA, the method comprising contacting the target RNA with a composition comprising the fusion protein of the invention, or the conjugate of the invention, or a polynucleotide encoding the fusion protein, wherein the fusion protein or conjugate comprises a detectable label (e.g., one detectable by fluorescence, Northern blot, or FISH) and a synthetic spacer sequence capable of binding to the target RNA.

[0080] Another aspect of the invention provides a eukaryotic cell comprising a clustered regularly interspaced short palindromic repeat (CRISPR)-Cas complex, wherein the CRISPR-Cas complex is: (1 A spacer sequence capable of hybridizing to a target RNA, and an RNA guide sequence containing a direct repeat (DR) sequence on the 3' side with respect to the spacer sequence; and (2) a CRISPR-associated protein (Cas) having any one of the amino acid sequences of SEQ ID NOs: 1 to 7, or a derivative or functional fragment of the Cas; the Cas, the derivative and the functional fragment of the Cas can (i) bind to the RNA guide sequence and (ii) target the target RNA.

[0081] Any embodiment of the present invention described herein that is mentioned only in the examples or claims, or only in the following one aspect / section, can be combined with any other one or more embodiments of the present invention, provided that it is not explicitly negated or inappropriate. It should be understood.

Brief Description of Drawings

[0082]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

[0083] 1. General Introduction The present invention described herein provides novel class 2, type VI Cas effector proteins, sometimes referred to herein as Cas13e and Cas13f. The novel Cas13 proteins of the present invention are much smaller than previously discovered Cas13 effector proteins (Cas13a-Cas13d) such that they can be readily packaged, along with their crRNA coding sequences, into low-capacity gene therapy vectors such as AAV vectors. Further, compared to Cas13a, Cas13b, and Cas13d effector proteins, the newly discovered Cas13e and Cas13f effector proteins are more potent in knocking down RNA target sequences and are more efficient in single base editing, while showing very little non-specific / collateral RNase activity after activation by crRNA-based target recognition, except when the spacer sequence is within a specific narrow range (e.g., about 30 nucleotides). Thus, these new Cas proteins are ideally suited for gene therapy. While being more efficient in single base editing, they show very little non-specific / collateral RNase activity after activation by crRNA-based target recognition, except when the spacer sequence is within a specific narrow range (e.g., about 30 nucleotides). Thus, these new Cas proteins are ideally suited for gene therapy.

[0084] Accordingly, in a first aspect, the present invention provides Cas13e and Cas13f effector proteins, such as those having the amino acid sequences of SEQ ID NOs: 1-7, or orthologs, homologs, various derivatives (described later herein), functional fragments (described later herein) thereof, wherein said orthologs, homologs, derivatives and functional fragments maintain at least one function of any one of the proteins of SEQ ID NOs: 1-7. Such functions include, but are not limited to, the ability to bind to the guide RNA / crRNA (described later herein) of the present invention to form a complex, RNase activity, and the ability to bind to and cleave a target RNA at a specific site under the guidance of a crRNA that is at least partially complementary to the target RNA.

[0085] In certain embodiments, the Cas13e or Cas13f effector protein of the invention can be (i) any one of SEQ ID NOs: 1-7; (ii) a derivative having one or more amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 residues) of addition, deletion, and / or substitution (e.g., conservative substitution) of any one of SEQ ID NOs: 1-7; or (iii) a derivative having at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity compared to any one of SEQ ID NOs: 1-7.

[0086] In certain embodiments, the Cas13e and Cas13f effector proteins, their orthologs, homologs, derivatives, and functional fragments do not occur in nature and have, for example, at least one amino acid difference compared to naturally occurring sequences.

[0087] In related aspects, the invention provides further derivative Cas13e and Cas13f effector proteins based on any one of SEQ ID NOs: 1-7, or their orthologs, homologs, derivatives, and functional fragments as described above, comprising another covalently or non-covalently linked protein or polypeptide or other molecule (such as a detection reagent or a drug / chemical moiety). Such other protein / polypeptide / other molecule can be linked, for example, via a chemical bond, gene fusion, or other non-covalent bond (such as biotin-streptavidin binding). Such derivative proteins do not affect the functions of the original protein, such as the ability to bind to the guide RNA / crRNA (described later herein) of the invention to form a complex, RNase activity, and the ability to bind to and cleave a target RNA at a specific site under the guidance of a crRNA that is at least partially complementary to the target RNA.

[0088] Using such an induction, for example, a nuclear localization signal (NLS such as the NLS of SV40 large T antigen) can be added to enhance the ability of the Cas13e and Cas13f effector proteins of the present invention to enter the cell nucleus. Also, using such an induction, a target molecule or moiety can be added to direct the Cas13e and Cas13f effector proteins of the present invention to specific cells or intracellular locations. Further, using such an induction, a detectable label can be added to facilitate the detection, monitoring, or purification of the Cas13e and Cas13f effector proteins of the present invention. Moreover, using such an induction, a deaminase moiety (such as one having adenine or cytosine deaminase activity) can be added to facilitate RNA base editing.

[0089] The induction can be carried out by adding any of the additional moieties to the N-terminus or C-terminus, or internally (e.g., via an internal fusion or a side chain of an internal amino acid), of the Cas13e and Cas13f effector proteins of the present invention.

[0090] In a related second aspect, the present invention provides a conjugate of a Cas13e and Cas13f effector protein of the present invention based on any one of SEQ ID NOs: 1 to 7, or an ortholog, homolog, derivative, and functional fragment thereof, which is conjugated to a moiety such as another protein or polypeptide, a detectable label, or a combination thereof. Such conjugate moieties include, but are not limited to, localization signals, reporter genes (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP), labels (e.g., fluorescent dyes such as FITC or DAPI), NLS, targeting moieties, DNA binding domains (e.g., MBP, Lex A DBD, Gal4 DBD), epitope tags (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc.), transcriptional activation domains (e.g., VP64 or VPR), transcriptional repression domains (e.g., KRAB moiety or SID moiety), nucleases (e.g., FokI), deamination domains (e.g., ADAR1, ADAR2, APOBEC, AID, or TAD), methylases, demethylases, transcriptional release factors, HDAC, ssRNA cleavage activity, dsRNA cleavage activity, ssDNA cleavage activity, dsDNA cleavage activity, DNA or RNA ligases, any combination thereof, and the like.

[0091] For example, the conjugate may include one or more NLSs that can be located at the N-terminus, C-terminus or near it, internally, or a combination thereof. The conjugation can be carried out via an amino acid (such as D or E, or S or T), an amino acid derivative (such as Ahx, β-Ala, GABA or Ava), or a PEG bond.

[0092] In certain embodiments, the conjugation does not affect the function of the original protein, such as the ability to bind to the guide RNA / crRNA of the present invention (described later herein) to form a complex, RNase activity, and the ability to bind to and cleave the target RNA at a specific site under the guidance of a crRNA that is at least partially complementary to the target RNA.

[0093] In a third related aspect, the present invention provides a fusion of a Cas13e and Cas13f effector protein of the present invention based on any one of SEQ ID NOs: 1-7, or an ortholog, homolog, derivative, and functional fragment thereof as described above, the fusion being a localization signal, a reporter gene (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP), NLS, a protein targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD), an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc.), a transcriptional activation domain (e.g., VP64 or VPR), a transcriptional repression domain (e.g., the KRAB moiety or the SID moiety), a nuclease (e.g., FokI), a deamination domain (e.g., ADAR1, ADAR2, APOBEC, AID, or TAD), a methylase, a demethylase, a transcriptional release factor, HDAC, ssRNA cleavage activity, dsRNA cleavage activity, ssDNA cleavage activity, dsDNA cleavage activity, a DNA or RNA ligase, a fusion with a moiety such as any combination thereof.

[0094] For example, the fusion may include one or more NLSs that may be located at the N-terminus, C-terminus or near it, internally, or a combination thereof. In certain embodiments, the conjugation does not affect the function of the original protein, e.g., the ability to bind to a guide RNA / crRNA of the present invention (described later herein) to form a complex, RNase activity, and the ability to bind to and cleave a target RNA at a specific site under the guidance of a crRNA that is at least partially complementary to the target RNA.

[0095] In a fourth aspect, the present invention provides a polynucleotide that is (i) any one of SEQ ID NOs: 8 to 14; (ii) a polynucleotide having 1, 2, 3, 4, or 5 nucleotides of deletion, addition, and / or substitution as compared to any one of SEQ ID NOs: 8 to 14; (iii) a polynucleotide sharing at least 80%, 85%, 90%, 95% sequence identity with any one of SEQ ID NOs: 8 to 14; (iv) a polynucleotide that hybridizes under stringent conditions to any one of the polynucleotides of (i) to (iii) or its complement; (v) an isolated polynucleotide comprising a complementary sequence of any of the polynucleotides of (i) to (iii).

[0096] (ii) to (iv) Any of the polynucleotides maintains the function of the original SEQ ID NOs: 8 to 14, which is for encoding the direct repeat (DR) sequence of crRNA in the Cas13e or Cas13f system of the present invention.

[0097] As used herein, "direct repeat sequence" can refer to a DNA coding sequence at a CRISPR locus, or the RNA encoded thereby in a crRNA. Thus, when any of SEQ ID NOs: 8 to 14 is referred to in the context of an RNA molecule such as crRNA, it is understood that each T represents U.

[0098] Thus, in certain embodiments, the isolated polynucleotide is DNA encoding the DR sequence for the crRNA of the Cas13e and Cas13f systems of the present invention.

[0099] In certain other embodiments, the isolated polynucleotide is RNA that is the DR sequence for the crRNA of the Cas13e and Cas13f systems of the present invention.

[0100] In a fifth aspect, the present invention provides a complex comprising: (i) a protein composition that can be any one of the Cas13e or Cas13f effector proteins of the present invention, or an ortholog, homolog, derivative, conjugate, functional fragment, conjugate thereof, or fusion thereof; and (ii) an isolated polynucleotide described in the fourth aspect of the present invention (e.g., a DR sequence), and a polynucleotide composition comprising a spacer sequence complementary to at least a portion of the target RNA. In certain embodiments, the DR sequence is at the 3' end of the spacer sequence.

[0101] In some embodiments, the polynucleotide composition is a guide RNA / crRNA of the Cas13e or Cas13f system of the present invention that does not contain tracrRNA.

[0102] In certain embodiments, for use with Cas13e and Cas13f effector proteins, homologs, orthologs, derivatives, fusions, conjugates, or functional fragments having RNase activity, the spacer sequence is at least about 10 nucleotides, or 10 - 60, 15 - 50, 20 - 50, 25 - 40, 25 - 50, or 19 - 50 nucleotides. In certain embodiments, for use with Cas13e and Cas13f effector proteins, or homologs, orthologs, derivatives, fusions, conjugates, or functional fragments that do not have RNase activity but have the ability to bind to a guide RNA and a target RNA complementary to the guide RNA, the spacer sequence is at least about 10 nucleotides, or about 10 - 200, 15 - 180, 20 - 150, 25 - 125, 30 - 110, 35 - 100, 40 - 80, 45 - 60, 50 - 55, or about 50 nucleotides.

[0103] In certain embodiments, the DR sequence is 15 - 36, 20 - 36, 22 - 36, or about It is 36 nucleotides. In certain embodiments, the DR sequence in the guide RNA has a secondary structure (including stems, bulges, and loops) that is substantially the same as any one of the RNA forms of SEQ ID NOs: 8-14.

[0104] In certain embodiments, the guide RNA is about 36 nucleotides longer than any of the above spacer sequence lengths, such as 45-96, 55-86, 60-86, 62-86, or 63-86 nucleotides.

[0105] In a sixth aspect, the present invention provides an isolated polynucleotide comprising: (i) a polynucleotide encoding any one of the Cas13e or Cas13f effector proteins of SEQ ID NOs: 1-7, or an ortholog, homolog, derivative, functional fragment, fusion thereof; (ii) a polynucleotide of any one of SEQ ID NOs: 8-14; or (iii) a polynucleotide comprising (i) and (ii).

[0106] In some embodiments, the polynucleotide is non-natural / naturally occurring, excluding, for example, SEQ ID NOs: 15-21.

[0107] In some embodiments, the polynucleotide is codon-optimized for prokaryotic expression. In some embodiments, the polynucleotide is codon-optimized for expression in eukaryotes, such as humans or human cells.

[0108] In a seventh aspect, the present invention provides a vector comprising or containing any of the polynucleotides of the sixth aspect. The vector can be a cloning vector or an expression vector. The vector can be, by way of example only, a plasmid, a phagemid, or a cosmid. In certain embodiments, the vector can be used to express a polynucleotide, any one of the Cas13e or Cas13f effector proteins of SEQ ID NOs: 1-7, or an ortholog, homolog, derivative, functional fragment, fusion thereof in mammalian cells such as human cells; or any of the polynucleotides of the fourth aspect; or any of the complexes of the fifth aspect.

[0109] In an eighth aspect, the present invention provides a host cell comprising any of the polynucleotides of the fourth or sixth aspect and / or the vector of the seventh aspect of the present invention. The host cell can be a prokaryote such as Escherichia coli (E. coli), or a cell derived from a eukaryote such as yeast, insect, plant, animal (e.g., mammals including humans and mice). The host cell can be an isolated primary cell (such as bone marrow cells for ex vivo therapy), or an established cell line such as a tumor cell line, 293T cell, or a stem cell, iPC.

[0110] In a related aspect, the present invention provides a eukaryotic cell comprising a clustered regularly interspaced short palindromic repeat (CRISPR)-Cas complex, wherein the CRISPR-Cas complex comprises: (1) an RNA guide sequence comprising a spacer sequence capable of hybridizing to a target RNA and a direct repeat (DR) sequence on the 3'-side of the spacer sequence; and (2) a CRISPR-associated protein (Cas) having any one of the amino acid sequences of SEQ ID NOs: 1-7, or a derivative or functional fragment of the Cas, wherein the Cas, the derivative of the Cas, and the functional fragment are capable of (i) binding to the RNA guide sequence and (ii) targeting the target RNA.

[0111] In a ninth aspect, the present invention provides a composition comprising: (i) a first (protein) composition selected from any one of the Cas13e or Cas13f effector proteins of SEQ ID NOs: 1-7, or an ortholog, homolog, derivative, conjugate, functional fragment, or fusion thereof; and (ii) an RNA comprising a guide RNA / crRNA, particularly a spacer sequence therefor, or a second (nucleotide) composition comprising a coding sequence. The guide RNA may comprise a DR sequence and a spacer sequence that can complement or hybridize to the target RNA. The guide RNA can form a complex with the first (protein) composition of (i). In certain embodiments, the DR sequence can be the polynucleotide of the fourth aspect of the present invention. In certain embodiments, the DR sequence can be at the 3' end of the guide RNA. In some embodiments, the composition (such as (i) and / or (ii)) is non-natural or modified from a natural composition. In some embodiments, at least some components of the composition are non-natural or modified from the natural components of the composition. In some embodiments, the target sequence is an RNA from a prokaryote or eukaryote, such as an RNA that does not occur naturally. The target RNA can be present inside the cell, such as in the cytosol or inside an organelle. In some embodiments, the protein composition can have an NLS located at its N-terminus or C-terminus, or internally.

[0112] In a tenth aspect, the present invention provides a composition comprising one or more vectors of the seventh aspect of the present invention, wherein the one or more vectors comprise: (i) optionally, a first polynucleotide encoding any one of the Cas13e or Cas13f effector proteins of SEQ ID NOs: 1-7, or an ortholog, homolog, derivative, functional fragment, or fusion thereof, operably linked to a first regulatory element; and (ii) optionally, a second polynucleotide encoding the guide RNA of the present invention, operably linked to a second regulatory element. The first and second polynucleotides can be on different vectors or on the same vector. The guide RNA can form a complex with the protein product encoded by the first polynucleotide and can include a DR sequence (such as any one of the fourth aspect) and a spacer sequence that can bind to the target RNA / complement the target RNA. In some embodiments, the first regulatory element is a promoter such as an inducible promoter. In some embodiments, the second regulatory element is a promoter such as an inducible promoter. In some embodiments, the composition (such as (i) and / or (ii)) is non-natural or modified from a natural composition. In some embodiments, at least some components of the composition are non-natural or modified from the natural components of the composition. In some embodiments, the target sequence is an RNA derived from a prokaryote or eukaryote, such as an RNA that does not occur naturally. The target RNA can be present inside the cell, such as in the cytosol or inside an organelle. In some embodiments, the protein composition can have an NLS that can be located at its N-terminus or C-terminus, or internally.

[0113] In some embodiments, the vector is a plasmid. In certain embodiments, the vector is a viral vector based on a retrovirus, a replication-deficient retrovirus, an adenovirus, a replication-deficient adenovirus, or AAV. In some embodiments, the vector is self-replicable within the host cell (e.g., has a bacterial origin of replication sequence). In some embodiments, the vector can integrate into the host genome and be replicated by the host genome. In certain embodiments, the vector is a cloning vector. In certain embodiments, the vector is an expression vector.

[0114] The present invention further provides a delivery composition for delivering the Cas13e or Cas13f effector protein of SEQ ID NOs: 1-7 of the first to third aspects of the present invention, or an ortholog, homolog, derivative, conjugate, functional fragment, fusion thereof; the polynucleotide of the fourth and / or sixth aspects of the present invention; the complex of the fifth aspect of the present invention; the vector of the seventh aspect of the present invention; the cell of the eighth aspect of the present invention, and the composition of the ninth and / or tenth aspects of the present invention. The delivery can be carried out by any one known in the art, such as transfection, lipofection, electroporation, gene gun, microinjection, sonication, calcium phosphate transfection, cationic transfection, viral vector delivery, etc.

[0115] The present invention further provides a kit comprising any one or more of the following: any one of the Cas13e or Cas13f effector proteins of SEQ ID NOs: 1-7 of the first to third aspects of the present invention, or an ortholog, homolog, derivative, conjugate, functional fragment, or fusion thereof; the polynucleotide of the fourth and / or sixth aspects of the present invention; the complex of the fifth aspect of the present invention; the vector of the seventh aspect of the present invention; the cell of the eighth aspect of the present invention, and the composition of the ninth and / or tenth aspects of the present invention. In some embodiments, the kit may further comprise instructions on how to use the components of the kit and / or how to obtain additional components from a third party for use with the components of the kit. Any component of the kit may be stored in any suitable container.

[0116] The present invention has been generally described above in this specification, but more detailed descriptions of various aspects of the present invention are provided in the following separate sections. However, for the sake of brevity and to reduce redundancy, it should be understood that specific embodiments of the present invention may be described in only one section, or only in the claims or examples. Thus, it should also be understood that any one embodiment of the present invention, including those described in only one aspect, section, or only in the claims or examples, may be combined with any other embodiment of the present invention, unless specifically negated or the combination is inappropriate.

[0117] 2. Novel Class 2, Type VI CRISPR RNA-guided RNases, and Derivatives Thereof In one aspect, the present invention described herein provides two novel families of CRISPR Class 2, Type VI effectors having two strictly conserved RX4-6H (RXXXXH) motifs characteristic of the nucleotide-binding (HEPN) domains of higher eukaryotes and prokaryotes. Similar CRISPR Class 2, Type VI effectors containing two HEPN domains have been previously characterized and include, for example, CRISPR Cas13a (C2c2), Cas13b, Cas13c, and Cas13d.

[0118] The HEPN domain is an RNase domain that has been shown to bind to and cleave target RNA molecules. The target RNA can be any suitable form of RNA including, but not limited to, mRNA, tRNA, ribosomal RNA, non-coding RNA, lncRNA (long non-coding RNA), and nuclear RNA. For example, in some embodiments, the Cas protein recognizes and cleaves an RNA target located on the coding strand of an open reading frame (ORF).

[0119] In one embodiment, the present disclosure provides two families of class 2, type VI CRISPR-Cas effector proteins, generally referred to herein as type VI-E and VI-F CRISPR-Cas effector proteins, Cas13e or Cas13f. A direct comparison of type VI-E and VI-F CRISPR-Cas effector proteins with effector proteins of these other systems shows that the type VI-E and VI-F CRISPR-Cas effector proteins are significantly smaller (e.g., about 20% fewer amino acids) than the previously identified minimal type VI-D / Cas13d effector (see FIG. 4) and have less than 30% sequence similarity in a one-to-one sequence alignment with other previously described effector proteins including the phylogenetically closest relative Cas13b (see FIG. 3).

[0120] These two newly identified families of class 2, type VI CRISPR-Cas effectors can be used for a variety of applications and are particularly suitable for therapeutic applications because they are significantly smaller than other effectors (e.g., CRISPR Cas13a, Cas13b, Cas13c, and Cas13d effectors), thereby enabling packaging into AAV vectors This is because it enables the packaging of effectors and nucleic acids encoding their guide RNA coding sequences into delivery systems with size limitations. Furthermore, the absence of detectable collateral / non-specific RNase activity in a selected range of spacer sequence lengths (such as about 30 nucleotides, see Figure 11) makes it less likely for these Cas effectors to cause potentially dangerous general off-target RNA digestion in target cells (in the absence of immunity) after activation of specific RNase activity. On the other hand, at other selected spacer lengths such as about 30 nucleotides, significant collateral RNase activity exists for these Cas effectors, and thus the Cas effectors of the present invention can also be used in applications that rely on such collateral RNase activity.

[0121] In bacteria, type VI-E and VI-F CRISPR-Cas systems contain a single effector (about 775 residues and 790 residues, respectively) in close proximity to the CRISPR array (see Figure 1). The CRISPR array typically contains direct repeat (DR) sequences that are 36 nucleotides long, which are generally well conserved both in sequence and secondary structure (see Figure 2).

[0122] The data presented herein demonstrate that crRNAs are processed from the 5' end such that the DR sequence ends at the 3' end of the mature crRNA.

[0123] The spacers contained in the Cas13e and Cas13f CRISPR arrays are most commonly 30 nucleotides in length, and most of the length variations are contained within the range of 29 - 30 nucleotides. However, a wide range of spacer lengths can be tolerated. For example, for use in a functional Cas13e or Cas13f effector protein, or a homolog, ortholog, derivative, fusion, conjugate, or functional fragment thereof, the spacer can be 10 - 60 nucleotides, 20 - 50 nucleotides, 25 - 45 nucleotides, 25 - 35 nucleotides, or about 27, 28, 29, 30, 31, 32, or 33 nucleotides. However, for use in any of the above dCas forms, the spacer can be 10 - 200 nucleotides, 20 - 150 nucleotides, 25 - 100 nucleotides, 25 - 85 nucleotides, 35 - 75 nucleotides, 45 - 60 nucleotides, or about 46, 47, 48, 49, 50, 51, 52, 53, 54, or 55 nucleotides.

[0124] Exemplary type VI-E and VI-F CRISPR-Cas effector proteins are shown in the following table.

[0125] [Table 1]

[0126] [Table 2]

[0127] [Table 3]

[0128] In the above array, two RX4-6H (RXXXXH) motifs in each effector are underlined. In Cas13e.1, the C-terminal motif can have two possibilities depending on the RR and HH sequences adjacent to the motif. Mutations in one or both of such domains can give rise to the RNase dead form (or "dCas") of the Cas13e and Cas13f effector proteins, their homologs, orthologs, fusions, conjugates, derivatives, or functional fragments, while substantially maintaining their ability to bind to the guide RNA and the target RNA complementary to the guide RNA.

[0129] The corresponding DR coding sequences of the Cas effectors are listed below:

[0130]

Table 4

[0131] Since the secondary structure of the DR sequence, including the positions and sizes of the steps, bulges, and loop structures, is likely to be more important than the specific nucleotide sequence forming such secondary structure, alternative or derivative DR sequences can also be used in the systems and methods of the present invention as long as these derivative or alternative DR sequences have a secondary structure substantially similar to the secondary structure of the RNA encoded by any one of SEQ ID NOs: 8-14. For example, the derivative DR sequence can have ±1 or 2 base pairs in one or both of the stems (see Figure 2), can have ±1, 2, or 3 bases in either or both of the single strands in the bulge, and / or can have ±1, 2, 3, or 4 bases in the loop region.

[0132] In some embodiments, type VI-E and VI-F CRISPR-Cas effector proteins include "derivatives" having an amino acid sequence with at least about 80% sequence identity (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) to any one of the amino acid sequences of SEQ ID NOs: 1-7 above. Such derivative Cas effectors that share significant protein sequence identity to any one of SEQ ID NOs: 1-7 retained the ability to bind to and form a complex with at least one crRNA containing at least one of the Cas functions of SEQ ID NOs: 8-14 (see below). For example, a Cas13e.1 derivative may share 85% amino acid sequence identity to each of SEQ ID NOs: 1, 2, 3, 4, 5, 6, or 7 and retain the ability to bind to and form a complex with a crRNA having each of the DR sequences of SEQ ID NOs: 8, 9, 10, 11, 12, 13, or 14.

[0133] In some embodiments, the derivative includes conservative amino acid residue substitutions. In some embodiments, the derivative includes only conservative amino acid residue substitutions (i.e., all amino acid substitutions in the derivative are conserved substitutions and there are no non-conserved substitutions).

[0134] In some embodiments, the derivative includes 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or fewer amino acid insertions or deletions into any one of the wild-type sequences of SEQ ID NOs: 1-7. The insertions and / or deletions may be grouped together or separated throughout the length of the sequence so long as at least one of the functions of the wild-type sequence is conserved. Such functions may include the ability to bind to a guide / crRNA, RNase activity, the ability to bind to and / or cleave a target RNA complementary to the guide / crRNA. In some embodiments, the insertions and / or deletions are not present within or 5, 10, 15, or 20 residues from the RXXXXH motif.

[0135] In some embodiments, the derivative retained the ability to bind to the guide RNA / crRNA.

[0136] In some embodiments, the derivative retained the guide / crRNA-activated RNase activity.

[0137] In some embodiments, the derivative retained the ability to bind to and / or cleave target RNA in the presence of guide / crRNA bound to at least a portion of the target RNA in the sequence.

[0138] In other embodiments, the derivative had completely or partially lost the guide / crRNA-activated RNase activity, for example, due to mutations in one or more catalytic residues of the RNA-guided RNase. Such derivatives may be referred to as dCas, for example, dCas13e.1, etc.

[0139] Thus, in certain embodiments, the derivative can be modified to have reduced nuclease / RNase activity, for example, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or 100% nuclease inactivation compared to the corresponding wild-type protein. The nuclease activity can be reduced by several methods known in the art, for example, by introducing mutations into the nuclease (catalytic) domain of the protein. In some embodiments, the catalytic residues for nuclease activity are identified and these amino acid residues can be substituted with various amino acid residues (e.g., glycine or alanine) to reduce the nuclease activity. In some embodiments, the amino acid substitution is a conservative amino acid substitution. In some embodiments, the amino acid substitution is a non-conservative amino acid substitution.

[0140] ​In some embodiments, the modification comprises one or more mutations (e.g., amino acid deletions, insertions, or substitutions) in at least one HEPN domain. In some embodiments, there are 1, 2, 3, 4, 5, 6, 7, 8, 9, or more amino acid substitutions in at least one HEPN domain. For example, in some embodiments, one or more mutations include substitutions (e.g., alanine substitutions) at the amino acid residues corresponding to R84, H89, R739, H744, R740, H745 of SEQ ID NO: 1, or R97, H102, R770, H775 of SEQ ID NO: 2, or R77, H82, R764, H769 of SEQ ID NO: 3, or R79, H84, R766A, H771 of SEQ ID NO: 4, or R79, H84, R766, H771 of SEQ ID NO: 5, or R89, H94, R773, H778 of SEQ ID NO: 6, or R89, H94, R777, H782 of SEQ ID NO: 7.

[0141] In certain embodiments, one or more mutations or two or more mutations can be in the catalytic active domain of an effector protein that includes a HEPN domain or a catalytic active domain homologous to the HEPN domain. In certain embodiments, the effector protein comprises one or more of the following mutations: R84A, H89A, R739A, H744A, R740A, H745A (where the amino acid positions correspond to the amino acid positions of Cas13e.1). One of ordinary skill in the art will understand that the corresponding amino acid positions in different Cas13e and Cas13f proteins can be mutated to have the same effect. In certain embodiments, one or more mutations completely or partially abolish the catalytic activity of the protein (e.g., altered cleavage rate, altered specificity, etc.).

[0142] Other exemplary (catalytic) residue mutations are R97A, H102A, R770A, H775A of Cas13e.2, or R77A, H82A, R764A, H769A of Cas13f.1, or R79A, H84A, R766A, H771A of Cas13f.2, or Cas It comprises R79A, H84A, R766A, H771A of Cas13f.3, or R89A, H94A, R773A, H778A of Cas13f.4, or R89A, H94A, R777A, H782A of Cas13f.5. In certain embodiments, any of the R and / or H residues herein may be substituted with G, V, or I rather than A.

[0143] The presence of at least one of these mutations results in derivatives having reduced or diminished RNase activity as compared to the corresponding wild-type protein lacking the mutation.

[0144] In certain embodiments, the effector proteins described herein are "dead" effector proteins such as dead Cas13e or Cas13f effector proteins (i.e., dCas13e and dCas13f). In certain embodiments, the effector protein has one or more mutations in HEPN domain 1 (N-terminal). In certain embodiments, the effector protein has one or more mutations in HEPN domain 2 (C-terminal). In certain embodiments, the effector protein has one or more mutations in both HEPN domain 1 and HEPN domain 2.

[0145] The inactivated Cas or derivative or a functional fragment thereof can be fused or associated with one or more heterologous / functional domains (e.g., via a fusion protein, a linker peptide, a "GS" linker, etc.). These functional domains can have various activities, such as methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, DNA cleavage activity, nucleic acid binding activity, base editing activity, and switch activity (e.g., photoinducible). In some embodiments, the functional domains are Krüppel-associated box (KRAB), SID (e.g., SID4X), VP64, VPR, VP16, Fok1, P65, HSF1, MyoD1, adenosine deaminases acting on RNA such as ADAR1, ADAR2, APOBEC, cytidine deaminase (AID), TAD, mini-SOG, APEX, and biotin-APEX.

[0146] In some embodiments, the functional domain is a base editing domain, such as ADAR1 (including wild-type with or without E1008Q or its ADAR1DD form), ADAR2 (including wild-type with or without the E488Q mutation or its ADAR2DD form), APOBEC, or AID.

[0147] In some embodiments, the functional domain can include one or more nuclear localization signal (NLS) domains. The one or more heterologous functional domains can include at least two or more NLS domains. The one or more NLS domains may be located at or near the terminus of an effector protein (e.g., a Cas13e / Cas13f effector protein), and in the case of two or more NLSs, each of the two may be located at or near the terminus of the effector protein (e.g., a Cas13e / Cas13f effector protein).

[0148] In some embodiments, at least one or more non-homologous functional domains can be at or near the amino terminus of the effector protein and / or at least one or more non-homologous functional domains are at or near the carboxy terminus of the effector protein. One or more non-homologous functional domains can be fused to the effector protein. One or more non-homologous functional domains can be tethered to the effector protein. One or more non-homologous functional domains can be linked to the effector protein by a linker portion.

[0149] In some embodiments, there are multiple (e.g., 2, 3, 4, 5, 6, 7, 8, or more) identical or different functional domains.

[0150] In some embodiments, a functional domain (e.g., a base editing domain) is further fused to an RNA binding domain (e.g., MS2).

[0151] In some embodiments, a functional domain associates with or is fused through a linker sequence (e.g., a flexible linker sequence or a rigid linker sequence). Exemplary linker sequences and functional domain sequences are provided in the following table.

[0152]

Table 5

[0153] The positioning of one or more functional domains in an inactivated Cas protein causes It enables the appropriate spatial positioning of functional domains to target the resulting functional effects. For example, when the functional domain is a transcriptional activator (e.g., VP16, VP64, or p65), the transcriptional activator is arranged in a spatial positioning that enables it to affect the transcription of the target. Similarly, a transcriptional repressor is positioned to affect the transcription of the target, and a nuclease (e.g., Fok1) is positioned to cleave or partially cleave the target. In some embodiments, the functional domain is positioned at the N-terminus of Cas / dCas. In some embodiments, the functional domain is positioned at the C-terminus of Cas / dCas. In some embodiments, the inactivated CRISPR-associated protein (dCas) is modified to include a first functional domain at the N-terminus and a second functional domain at the C-terminus.

[0154] Various examples of inactivated CRISPR-associated proteins fused with one or more functional domains and methods of using them are described, for example, in particular, with respect to the features described herein, in WO 2017 / 219027 pamphlet, which is hereby incorporated by reference in its entirety.

[0155] In some embodiments, type VI-E and VI-F CRISPR-Cas effector proteins include any one of the amino acid sequences of SEQ ID NOs: 1-7 above. In some embodiments, type VI-E and VI-F CRISPR-Cas effector proteins exclude any one of the natural amino acid sequences of SEQ ID NOs: 1-7 above.

[0156] In some embodiments, instead of using the full-length wild-type (SEQ ID NOs: 1-7) or derivatives of type VI-E and VI-F Cas effectors, their "functional fragments" can be used.

[0157] As used herein, "functional fragment" refers to a fragment of any one of the wild-type proteins of SEQ ID NOs: 1-7, or a derivative thereof having a sequence less than full length. The deleted residues in the functional fragment can be at the N-terminus, C-terminus, and / or internally. The functional fragment retains at least one function of wild-type VI-E or VI-F Cas, or at least one function of its derivative. Thus, the functional fragment is specifically defined with respect to the function. For example, a functional fragment whose function is the ability to bind to crRNA and target RNA may not be a functional fragment with respect to the RNase function, because losing the RXXXXH motif at both ends of Cas does not affect its ability to bind to crRNA and target RNA, but can abolish and destroy the RNase activity.

[0158] In some embodiments, compared to SEQ ID NOs: 1-7 of the full-length sequence, the type VI-E or VI-F CRISPR-Cas effector protein or its derivative or its functional fragment lacks about 30, 60, 90, 120, 150, or about 180 residues from the N-terminus.

[0159] In some embodiments, compared to SEQ ID NOs: 1-7 of the full-length sequence, the type VI-E or VI-F CRISPR-Cas effector protein or its derivative or its functional fragment lacks about 30, 60, 90, 120, or about 150 residues from the C-terminus.

[0160] In some embodiments, compared to SEQ ID NOs: 1-7 of the full-length sequence, the type VI-E or VI-F CRISPR-Cas effector protein or its derivative or its functional fragment lacks about 30, 60, 90, 120, 150, or about 180 residues from the N-terminus and lacks about 30, 60, 90, 120, or about 150 residues from the C-terminus.

[0161] In some embodiments, the type VI-E or VI-F CRISPR-Cas effec The turret protein or its derivative or its functional fragment has RNase activity, for example, guide / crRNA-activated specific RNase activity.

[0162] In some embodiments, the type VI-E or VI-F CRISPR-Cas effector protein or its derivative or its functional fragment has no substantial / detectable collateral RNase activity.

[0163] As used herein, "collateral RNase activity" refers to the non-specific RNase activity observed in certain other class 2, type VI RNA-guided RNases, such as Cas13a. A complex containing Cas13a undergoes a conformational change, for example, after activation by binding to a target nucleic acid (e.g., a target RNA), such that the complex then acts as a non-specific RNase and cleaves and / or degrades RNA molecules (e.g., ssRNA or dsRNA molecules) in the vicinity (i.e., the "collateral" effect).

[0164] In certain embodiments, a complex composed of, but not limited to, a type VI-E or VI-F CRISPR-Cas effector protein or its derivative or its functional fragment and a crRNA does not exhibit collateral RNase activity after target recognition. This "collateral-free" embodiment may include wild-type, engineered / derivative effector proteins, or functional fragments thereof.

[0165] In some embodiments, the type VI-E or VI-F CRISPR-Cas effector protein or its derivative or its functional fragment recognizes and cleaves target RNA without additional requirements (i.e., protospacer adjacent motif "PAM" or protospacer flanking sequence "PFS" requirements) adjacent to or proximal to the protospacer.

[0166] The present disclosure also provides split forms of the CRISPR-related proteins described herein (e.g., type VI-E or VI-F CRISPR-Cas effector proteins). Split forms of CRISPR-related proteins can be advantageous for delivery. In some embodiments, the CRISPR-related protein is split into two parts of the enzyme, which together substantially comprise a functional CRISPR-related protein.

[0167] The splitting can be done in a way that does not affect the catalytic domain. The CRISPR-related protein can function as a nuclease or be an inactivated enzyme, which is an RNA-binding protein that has little or no catalytic activity, substantially (e.g., due to mutations in its catalytic domain). The degrading enzyme is described, for example, in Wright et al., “Rational design of a split-Cas9 enzyme complex”, Proc. Nat’l. Acad. Sci. 112(10):2984-2989, 2015, which is hereby incorporated by reference in its entirety.

[0168] For example, in some embodiments, the nuclease lobe and the α-helical lobe are expressed as separate polypeptides. The lobes do not interact with each other by themselves, but the crRNA recruits them into a ternary complex, which recapitulates the activity of the full-length CRISPR-related protein and catalyzes site-specific DNA cleavage. The use of modified crRNA enables the generation of an inducible dimerization system by preventing dimerization and eliminating degrading enzyme activity.

[0169] In some embodiments, the split CRISPR-related protein can be fused to a dimerization partner, for example, by using a rapamycin-sensitive dimerization domain. This enables the generation of a chemically inducible CRISPR-related protein for temporal control of protein activity. Thus, the CRISPR-related protein is split into two fragments It can be chemically induced, and the rapamycin-sensitive dimerization domain can be used for the controlled reorganization of proteins.

[0170] The cleavage points are typically designed in silico and cloned into constructs. During this process, mutations can be introduced into the split CRISPR-associated proteins, and non-functional domains can be removed.

[0171] In some embodiments, the two parts or fragments of the split CRISPR-associated protein (i.e., the N-terminal and C-terminal fragments) can form a complete CRISPR-associated protein that contains, for example, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% of the sequence of the wild-type CRISPR-associated protein.

[0172] The CRISPR-associated proteins described herein (e.g., type VI-E or VI-F CRISPR-Cas effector proteins) can be designed to be self-activating or self-inactivating. For example, a target sequence can be introduced into the coding construct of the CRISPR-associated protein. Thus, the CRISPR-associated protein can cleave the target sequence, as well as the construct encoding the protein, thereby self-inactivating their expression. Methods for constructing self-inactivating CRISPR systems are described, for example, in Epstein and Schaffer, Mol. Ther. 24:S50, 2016, which is incorporated herein by reference.

[0173] In some other embodiments, additional crRNAs expressed under the control of a weak promoter (e.g., the 7SK promoter) can target nucleic acid sequences encoding CRISPR-associated proteins and prevent and / or block their expression (e.g., by preventing transcription and / or translation of the nucleic acid). Transfection of cells with a vector expressing a CRISPR-associated protein, a crRNA, and a crRNA targeting the nucleic acid encoding the CRISPR-associated protein can result in efficient disruption of the nucleic acid encoding the CRISPR-associated protein, reducing the level of the CRISPR-associated protein, thereby restricting genome editing activity.

[0174] In some embodiments, the genome editing activity of a CRISPR-associated protein can be regulated via an endogenous RNA signature (e.g., miRNA) within mammalian cells. A CRISPR-associated protein switch can be made by using a miRNA complementary sequence in the 5'-UTR of the mRNA encoding the CRISPR-associated protein. The switch selectively and efficiently responds to miRNAs within the target cell. Thus, the switch can differentially control genome editing by sensing endogenous miRNA activity within a heterogeneous cell population. Accordingly, the switch system can provide a framework for cell type-selective genome editing and cell engineering based on intracellular miRNA information (see, e.g., Hirosawa et al., Nucl. Acids Res. 45(13):e118, 2017).

[0175] CRISPR-related proteins (e.g., type VI-E and VI-F CRISPR-Cas effector proteins) can be expressed inducibly, e.g., their expression can be light-inducible or chemically inducible. This mechanism enables the activation of functional domains in the CRISPR-related proteins. Light-inducibility can be achieved by various methods known in the art, e.g., by designing a fusion complex in which the CRY2 PHR / CIBN pairing is used to split the CRISPR-related protein (see, e.g., Konermann et al., “Optical control of mammalian endogenous transcription and epigenetic states”, Nature 500:7463, 2013).

[0176] Chemical inducibility can be achieved, e.g., by designing a fusion complex in which the FKBP / FRB (FK506-binding protein / FKBP rapamycin-binding domain) pairing is used to split the CRISPR-related protein. Rapamycin is required to form the fusion complex, thereby activating the CRISPR-related protein (see, e.g., Zetsche et al., “A split-Cas9 architecture for inducible genome editing and transcription modulation”, Nature Biotech. 33:2:139-42, 2015).

[0177] Furthermore, the expression of CRISPR-related proteins can be regulated by inducible promoters, such as transcription activation regulated by tetracycline or doxycycline (Tet-On and Tet-Off expression systems), hormone-inducible gene expression systems (e.g., ecdysone-inducible gene expression systems), and arabinose-inducible gene expression systems. When delivered as RNA, the expression of RNA-targeting effector proteins can be regulated via riboswitches that can sense small molecules such as tetracycline (see, e.g., Goldfless et al., “Direct and specific chemical control of eukaryotic translation with a synthetic RNA-protein interaction”, Nucl. Acids Res. 40:9:e64-e64, 2012).

[0178] Various embodiments of inducible CRISPR-related proteins and inducible CRISPR systems are described, for example, in U.S. Patent No. 8,871,445, U.S. Patent Application Publication No. 2016 / 0208243, and International Publication No. 2016 / 205764, each of which is hereby incorporated by reference in its entirety.

[0179] In some embodiments, the CRISPR-related protein comprises at least one (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) nuclear localization signal (NLS) that is bound to the N-terminus or C-terminus of the protein. Non-limiting examples of NLSs include NLS sequences derived from the following: the NLS of SV40 virus large T-antigen having the amino acid sequence PKKKRKV (Array No. 77) ; the NLS derived from nucleoplasmin (e.g., the nucleoplasmin bipartite NLS having the sequence KRPAATKKAGQAKKKK (Array No. 78) ); the amino acid sequence PAAKRVKLD (Array No. 79) or RQRRNELKRSP (Array No. 80)The c-myc NLS having; the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (Array No. 81) The hRNPA1 M9 NLS having; the sequence of the IBB domain from importin-α, RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (Array No. 82) ; the sequence of the muscle tumor T protein, VSRKRPRP (Array No. 83) and PPKKARED (Array No. 84) ; the sequence of human p53, PQPKKKPL (Array No. 85) ; the sequence of mouse c-abl IV, SALIKKKKKMAP (Array No. 86) ; the sequence of influenza virus NS1, DRLRR (Array No. 87) and PKQKKRK (Array No. 88) ; the sequence of the hepatitis delta virus antigen, RKLKKKIKKL (Array No. 89) ; the sequence of the mouse Mx1 protein, REKKKFLKRR (Array No. 90) ; the sequence of human poly(ADP-ribose) polymerase, KRKGDEVDGVDEVAKKKSKK (Array No. 91) ; and the sequence of the human glucocorticoid receptor, RKCLQAGMNLEARKTKK (Array No. 92) . In some embodiments, the CRISPR-related protein comprises at least one (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) nuclear export signal (NES) that is attached to the N-terminus or C-terminus of the protein. In preferred embodiments, the C-terminus and / or N-terminus NLS or NES is attached for optimal expression and nuclear targeting within eukaryotic cells, such as human cells.

[0180] In some embodiments, the CRISPR-related proteins described herein are mutated at one or more amino acid residues to modify one or more functional activities.

[0181] For example, in some embodiments, the CRISPR-related protein is mutated at one or more amino acid residues to modify its helicase activity.

[0182] In some embodiments, the CRISPR - associated protein is mutated at one or more amino acid residues to modify its nuclease activity (e.g., endonuclease activity or exonuclease activity).

[0183] In some embodiments, the CRISPR - associated protein is mutated at one or more amino acid residues to modify its ability to functionally associate with a guide RNA.

[0184] In some embodiments, the CRISPR - associated protein is mutated at one or more amino acid residues to modify its ability to functionally associate with a target nucleic acid.

[0185] In some embodiments, the CRISPR - associated proteins described herein are capable of cleaving a target RNA molecule.

[0186] In some embodiments, the CRISPR - associated protein is mutated at one or more amino acid residues to modify its cleavage activity. For example, in some embodiments, the CRISPR - associated protein can comprise one or more mutations that render the enzyme unable to cleave the target nucleic acid.

[0187] In some embodiments, the CRISPR - associated protein is capable of cleaving a strand of a target nucleic acid that is complementary to the strand to which the guide RNA hybridizes.

[0188] In some embodiments, the CRISPR - associated proteins described herein can be engineered to have deletions in one or more amino acid residues to reduce the size of the enzyme while retaining one or more desired functional activities (e.g., nuclease activity and the ability to functionally interact with a guide RNA). The truncated CRISPR - associated protein can be advantageously used in combination with a delivery system having a load limitation.

[0189] In some embodiments, the CRISPR-related proteins described herein can be fused to one or more peptide tags including His-tag, GST-tag, V5-tag, FLAG-tag, HA-tag, VSV-G-tag, Trx-tag, or myc-tag.

[0190] In some embodiments, the CRISPR-related proteins described herein can be fused to a detectable moiety such as GST, a fluorescent protein (e.g., GFP, HcRed, DsRed, CFP, YFP, or BFP), or an enzyme (such as HRP or CAT).

[0191] In some embodiments, the CRISPR-related proteins described herein can be fused to MBP, the LexA DNA binding domain, or the Gal4 DNA binding domain.

[0192] In some embodiments, the CRISPR-related proteins described herein can be linked or conjugated to a detectable label such as a fluorescent dye, including FITC and DAPI.

[0193] In any of the embodiments herein, the bond between the CRISPR-related protein described herein and another moiety can be at the N-terminus or C-terminus of the CRISPR-related protein and can also be internal via a covalent chemical bond. The bond can be made by a peptide bond, a bond via a side chain of an amino acid such as D, E, S, T, or an amino acid derivative (Ahx, β-Ala, GABA or Ava), or any chemical bond known in the art such as a PEG bond.

[0194] 3. Polynucleotide The present invention also provides a nucleic acid encoding a protein (e.g., a CRISPR-related protein or an accessory protein) and a guide RNA (e.g., crRNA) described herein.

[0195] ​In some embodiments, the nucleic acid is a synthetic nucleic acid. In some embodiments, the nucleic acid is a DNA molecule. In some embodiments, the nucleic acid is an RNA molecule (e.g., an mRNA molecule encoding Cas, its derivatives, or functional fragments). In some embodiments, the mRNA is capped, polyadenylated, substituted with 5-methylcytidine, substituted with pseudouridine, or a combination thereof.

[0196] In some embodiments, the nucleic acid (e.g., DNA) is operably linked to a regulatory element (e.g., a promoter) for controlling the expression of the nucleic acid. In some embodiments, the promoter is a constitutive promoter. In some embodiments, the promoter is an inducible promoter. In some embodiments, the promoter is a cell-specific promoter. In some embodiments, the promoter is an organism-specific promoter.

[0197] Suitable promoters are known in the art and include, for example, pol I promoter, pol II promoter, pol III promoter, T7 promoter, U6 promoter, H1 promoter, Rous sarcoma virus LTR promoter of retrovirus, cytomegalovirus (CMV) promoter, SV40 promoter, dihydrofolate reductase promoter, and β-actin promoter. For example, the U6 promoter can be used to regulate the expression of the guide RNA molecules described herein.

[0198] In some embodiments, the nucleic acid is present within a vector (e.g., a viral vector or phage). The vector can be a cloning vector or an expression vector. The vector can be a plasmid, phagemid, cosmid, or the like. The vector may include one or more regulatory elements that enable the growth of the vector within a cell of interest (e.g., a bacterial cell or a mammalian cell). In some embodiments, the vector includes a nucleic acid encoding a single component of the CRISPR-associated (Cas) system described herein. In some embodiments, the vector includes a plurality of nucleic acids, each encoding a component of the CRISPR-associated (Cas) system described herein.

[0199] In one aspect, the disclosure provides a nucleic acid sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the nucleic acid sequences described herein, i.e., the DR sequences of SEQ ID NOs: 8 - 14, which encode a Cas protein, derivative, functional fragment, or guide / crRNA.

[0200] In another aspect, the disclosure also provides a nucleic acid sequence encoding an amino acid sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequences described herein, e.g., SEQ ID NOs: 1 - 7.

[0201] In some embodiments, the nucleic acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, e.g., consecutive nucleotides or non - consecutive nucleotides) that is the same as the sequences described herein. In some embodiments, the nucleic acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, e.g., consecutive nucleotides or non - consecutive nucleotides) that is different from the sequences described herein.

[0202] In related embodiments, the invention provides an amino acid sequence having at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid residues, e.g., consecutive amino acid residues or non - consecutive amino acid residues) that is the same as the sequences described herein. In some embodiments, the amino acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid residues, e.g., consecutive amino acid residues or non - consecutive amino acid residues) that is different from the sequences described herein.

[0203] To determine the percent identity of two amino acid sequences or two nucleic acid sequences, the sequences are aligned for optimal comparison purposes (e.g., gaps can be introduced into one or both of the first and second amino acids or one or both of the first and second nucleic acid sequences for optimal alignment, and non-homologous sequences can be ignored for comparison purposes). Generally, the length of the reference sequence aligned for comparison purposes should be at least 80% of the length of the reference sequence, and in some embodiments, at least 90%, 95%, or 100% of the length of the reference sequence. Subsequently, the amino acid residues or nucleotides at corresponding amino acid positions or nucleotide positions are compared. If a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, the molecules are identical at that position. The percent identity between two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps that need to be introduced for optimal alignment of the two sequences and the length of each gap. For the purposes of the present disclosure, comparison of sequences and determination of percent identity between two sequences can be achieved using the Blossum 62 scoring matrix with a gap penalty of 12, a gap extension penalty of 4, and a frameshift gap penalty of 5.

[0204] The proteins described herein (e.g., CRISPR-related proteins or accessory proteins) can be delivered and used as nucleic acid molecules or polypeptides.

[0205] In certain embodiments, the nucleic acid molecule, derivative, or functional fragment encoding the CRISPR-related protein is codon-optimized for expression in a host cell or organism. The host cell can be an established cell line (e.g., 293T cells) or an isolated primary cell. The nucleic acid can be codon-optimized for use in any organism of interest, particularly human cells or bacteria. For example, the nucleic acid can be for any prokaryote (such as E. coli), or any eukaryote, such as a human, as well as yeast, worms, insects, plants, and algae (including food crops, rice, corn, vegetables, fruits, trees, grasses), vertebrates, fish, non-human mammals (such as mice, rats, rabbits, dogs, birds (such as chickens), livestock (such as cows, pigs, horses, sheep, goats, and others), or non-human primates), or other non-human eukaryotes. Codon usage tables are readily available, for example, in the "Codon Usage Database" available at www.kazusa.orjp / codon / , and the table can be constructed in several ways. N See akamura et al., Nucl. Acids Res. 28:292, 2000 (which is incorporated herein by reference in its entirety). Also, computer algorithms, such as Gene Forge (Aptagen; Jacobus, Pa.), are available for codon-optimizing specific sequences for expression in a particular host cell.

[0206] Examples of codon-optimized sequences are, in this case, sequences optimized for expression in eukaryotes, such as humans (i.e., optimized for expression in humans), or for another eukaryote, animal, or mammal, as discussed herein; see, for example, the SaCas9 human codon-optimized sequence in WO 2014 / 093622 (PCT / US2013 / 074667). While this is preferred, it is of course possible that other examples exist, and codon optimization for host species other than humans, or for specific organs, is known. In general, codon optimization refers to the process of modifying a nucleic acid sequence by replacing at least one codon of the original sequence (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) with codons that are used more frequently, or most frequently, in the genes of the host cell, while maintaining the original amino acid sequence, for enhanced expression in the host cell of interest. Different species exhibit a particular bias for particular codons of a particular amino acid. Codon bias (differences in codon usage frequency between organisms) often correlates with the efficiency of messenger RNA (mRNA) translation. This is thought to depend, in turn, at least in part, on the properties of the codons to be translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs within a cell is usually a reflection of the codons that are most frequently used in peptide synthesis. Thus, genes can be adjusted for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the "Codon Usage Database" available at http: / / www.kazusa.orjp / codon / , and the table can be constructed in several ways. See Nakamura, Y., et al. "Codon usage tabulated from the international DNA sequence databases: status for the year 2000" Nucl. Acids Res. 28:292 (2000).Also, computer algorithms for codon-optimizing specific sequences for expression in a particular host cell, such as Gene Forge (Aptagen; Jacobus, PA), are available. In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more, or all codons) within the sequence encoding Cas correspond to the codons most frequently used for a particular amino acid.

[0207] 4. RNA guide or crRNA In some embodiments, the CRISPR systems described herein include at least an RNA guide (e.g., gRNA or crRNA).

[0208] The architectures of multiple RNA guides are known in the art (see, e.g., WO 2014 / 093622 and WO 2015 / 070083, the entire contents of each of which are incorporated herein by reference).

[0209] In some embodiments, the CRISPR systems described herein include multiple RNA guides (e.g., 1, 2, 3, 4, 5, 6, 7, 8, or more RNA guides).

[0210] In some embodiments, the RNA guide includes crRNA. In some embodiments , the RNA guide includes crRNA rather than tracrRNA.

[0211] The sequences of guide RNAs derived from multiple CRISPR systems are generally known in the art. See, for example, Grissa et al. (Nucleic Acids Res. 35 (Web Server issue): W52-7, 2007; Grissa et al., BMC Bioinformatics 8:172, 2007; Grissa et al., Nucleic Acids Res. 36 (Web Server issue): W145-8, 2008; and Moller and Liang, PeerJ 5:e3788, 2017; CRISPR database at crispr.i2bc.paris-saclay.fr / crispr / BLAST / CRISPRsBlast.php; and MetaCRAST available at github.com / molleraj / MetaCRAST). All are incorporated herein by reference.

[0212] In some embodiments, the crRNA comprises a direct repeat (DR) sequence and a spacer sequence. In certain embodiments, the crRNA preferably comprises, consists essentially of, or consists of a guide sequence or a direct repeat sequence linked to the spacer sequence at the 3' end of the spacer sequence.

[0213] Generally, the Cas protein forms a complex with the mature crRNA, and the spacer sequence directs the complex to sequence-specific binding with a target RNA that is complementary to the spacer sequence and / or hybridizes to the spacer sequence. The resulting complex comprises the Cas protein bound to the target RNA and the mature crRNA.

[0214] The direct repeat sequences for the Cas13e system and the Cas13f system are generally well conserved, especially at the ends. At the 5' end, they have GCTG for Cas13e and GCTGT for Cas13f, and at the 3' end, they are reverse complementary to CAGC for Cas13e and ACAGC for Cas13f. This conservation suggests strong base pairing for the RNA stem-loop structures that potentially interact with the protein within the locus.

[0215] In some embodiments, the direct repeat sequence, when in RNA, includes a general secondary structure of 5'-S1a-Ba-S2a-L-S2b-Bb-S1b-3'. Segments S1a and S1b are reverse complementary sequences that form a first stem (S1) having 4 nucleotides in Cas13e and 5 nucleotides in Cas13f; segments Ba and Bb form a symmetric or nearly symmetric bulge (B) without base pairing with each other, having 5 nucleotides each in Cas13e, and 5 (Ba) and 4 (Bb), or 6 (Ba) and 5 (Bb) nucleotides in Cas13f respectively; segments S2a and S2b are reverse complementary sequences that form a second stem (S2) having 5 base pairs in Cas13e and 6 or 5 base pairs in Cas13f; L is an 8-nucleotide loop in Cas13e and a 5-nucleotide loop in Cas13f. See Figure 2.

[0216] In certain embodiments, S1a has a series of GCUG in Cas13f and a series of GCUGU in Cas13e.

[0217] In certain embodiments, S2a has a series of GCCCC in Cas13f and a series of A / G CCUC G / A in Cas13e (where the first A or G may be absent).

[0218] In some embodiments, the direct repeat array comprises or consists of the nucleic acid sequences of SEQ ID NOs: 8-14.

[0219] As used herein, "direct repeat array" can refer to the DNA coding sequence within the CRISPR locus or the RNA encoded by the same within the crRNA. Thus, when any of SEQ ID NOs: 8-14 are referred to in the context of an RNA molecule, such as a crRNA, it is understood that each T represents U.

[0220] In some embodiments, the direct repeat array comprises or consists of a nucleic acid sequence having a deletion, insertion, or substitution of up to 1, 2, 3, 4, 5, 6, 7, or 8 nucleotides of SEQ ID NOs: 8-14. In some embodiments, the direct repeat array comprises or consists of a nucleic acid sequence having at least 80%, 85%, 90%, 95%, or 97% sequence identity to SEQ ID NOs: 8-14 (e.g., due to nucleotide deletions, insertions, or substitutions in SEQ ID NOs: 8-14). In some embodiments, the direct repeat array comprises or consists of a nucleic acid sequence that is not identical to any one of SEQ ID NOs: 8-14 but can hybridize under stringent hybridization conditions to a complement of any one of SEQ ID NOs: 8-14 or can bind to a complement of any one of SEQ ID NOs: 8-14 under physiological conditions.

[0221] In certain embodiments, the deletion, insertion, or substitution does not change the overall secondary structure of SEQ ID NOs: 8-14 (e.g., the relative positions and / or sizes of the stems, bulges, and loops do not deviate significantly from the original stems, bulges, and loops). For example, the deletion, insertion, or substitution can be within a bulge or loop region such that the overall symmetry of the bulge remains mostly the same. The deletion, insertion, or substitution can be within a stem such that the overall length of the stem does not deviate significantly from the original stem (e.g., adding or deleting one base pair in each of two stems corresponds to a total of four base changes).

[0222] In certain embodiments, the deletion, insertion, or substitution may have ±1 or 2 base pairs in one or both stems (see Figure 2), may have ±1, 2, or 3 bases in one or both stems of the single strand within the bulge, and / or may have ±1, 2, 3, or 4 bases in the loop region, resulting in a derived DR sequence.

[0223] In certain embodiments, any subsequent directory repeat sequence that is different from any one of SEQ ID NOs: 8-14 retains the ability to function as a DR sequence of SEQ ID NOs: 8-14 and as a directory repeat sequence within the Cas13e protein or Cas13f protein.

[0224] In some embodiments, the directory repeat sequence comprises or consists of a nucleic acid having a nucleic acid sequence of any one of SEQ ID NOs: 8-14, with a truncation of the first 3, 4, 5, 6, 7, or 8 3' nucleotides.

[0225] In some embodiments, the Cas protein comprises the amino acid sequence of SEQ ID NO: 1, and the crRNA comprises a directory repeat sequence that comprises or consists of the nucleic acid sequence of SEQ ID NO: 8.

[0226] In some embodiments, the Cas protein comprises the amino acid sequence of SEQ ID NO: 2, and the crRNA comprises a directory repeat sequence that comprises or consists of the nucleic acid sequence of SEQ ID NO: 9.

[0227] In some embodiments, the Cas protein comprises the amino acid sequence of SEQ ID NO: 3, and the crRNA comprises a directory repeat sequence that comprises or consists of the nucleic acid sequence of SEQ ID NO: 10.

[0228] In some embodiments, the Cas protein comprises the amino acid sequence of SEQ ID NO: 4, and the crRNA comprises a directory repeat sequence, the directory repeat sequence comprising or consisting of the nucleic acid sequence of SEQ ID NO: 11.

[0229] In some embodiments, the Cas protein comprises the amino acid sequence of SEQ ID NO: 5, and the crRNA comprises a directory repeat sequence, the directory repeat sequence comprising or consisting of the nucleic acid sequence of SEQ ID NO: 12.

[0230] In some embodiments, the Cas protein comprises the amino acid sequence of SEQ ID NO: 6, and the crRNA comprises a directory repeat sequence, the directory repeat sequence comprising or consisting of the nucleic acid sequence of SEQ ID NO: 13.

[0231] In some embodiments, the Cas protein comprises the amino acid sequence of SEQ ID NO: 7, and the crRNA comprises a directory repeat sequence, the directory repeat sequence comprising or consisting of the nucleic acid sequence of SEQ ID NO: 14.

[0232] In classical CRISPR systems, the degree of complementarity between the guide sequence (e.g., crRNA) and its corresponding target sequence can be about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or 100%. In some embodiments, the degree of complementarity is 90 - 100%.

[0233] The guide RNA can be about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, 100, 125, 150, 175, 200, or more nucleotides in length. For example, the spacer used with a functional Cas13e or Cas13f effector protein, or a homolog, ortholog, derivative, fusion, conjugate, or functional fragment thereof, can be 10 - 60 nucleotides, 20 - 50 nucleotides, 25 - 45 nucleotides, 25 - 35 nucleotides, or about 27, 28, 29, 30, 31, 32, or 33 nucleotides. However, the spacer used with any of the previous dCas versions can be 10 - 200 nucleotides, 20 - 150 nucleotides, 25 - 100 nucleotides, 25 - 85 nucleotides, 35 - 75 nucleotides, 45 - 60 nucleotides, or about 46, 47, 48, 49, 50, 51, 52, 53, 54, or 55 nucleotides.

[0234] To reduce off - target interactions, for example, to reduce the guide interaction with off - target sequences with low complementarity, mutations can be introduced into the CRISPR system so that the CRISPR system can distinguish between a target and an off - target sequence having 80%, 85%, 90%, or more than 95% complementarity. In some embodiments, the degree of complementarity is 80% - 95%, for example, about 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, or 95% (e.g., distinguishing an 18 - nucleotide target from an 18 - nucleotide off - target having 1, 2, or 3 mismatches). Thus, in some embodiments, the degree of complementarity between the guide sequence and its corresponding target sequence is greater than 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 99.9%. In some embodiments, the degree of complementarity is 100%.

[0235] In this field, if there is sufficient complementarity such that it is functional, perfect complementarity is not required It is known that this is not the case. The regulation of cleavage efficiency can be developed by introducing one or more mismatches, such as one or two mismatches, etc. (including the position of the mismatch along the spacer / target), between the spacer sequence and the target sequence. A mismatch, such as a double mismatch, is positioned more centrally (i.e., not at the 3' or 5' end); the cleavage efficiency is more affected. Therefore, the cleavage efficiency can be regulated by selecting the mismatch position along the spacer sequence. For example, if less than 100% cleavage of the target is desired (e.g., in a cell population), one or two mismatches between the spacer and the target sequence can be introduced into the spacer sequence.

[0236] Type VI CRISPR-Cas effectors have been demonstrated to enable the ability of such effectors, as well as systems and complexes containing the same, to target multiple nucleic acids by using multiple RNA guides. In some embodiments, the CRISPR systems described herein include multiple (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, or more) RNA guides. In some embodiments, the CRISPR systems described herein include single-stranded RNAs, or nucleic acids encoding single-stranded RNAs, and the RNA guides are arranged in tandem. The single-stranded RNA can include multiple copies of the same RNA guide, multiple copies of different RNA guides, or combinations thereof. The processing ability of the type VI-E and type VI-F CRISPR-Cas effector proteins described herein enables these effectors to target multiple target nucleic acids (e.g., target RNAs) without loss of activity. In some embodiments, the type VI-E and type VI-F CRISPR-Cas effector proteins can be delivered complexed with multiple RNA guides directed to different target RNAs. In some embodiments, the type VI-E and type VI-F CRISPR-Cas effector proteins can be delivered with multiple RNA guides, each specific for a different target nucleic acid. Methods of multiplexing using CRISPR-associated proteins are described, for example, in U.S. Patent No. 9,790,490 B2 and European Patent No. 3,009,511 B1, the entire contents of each of these patents being expressly incorporated herein by reference.

[0237] The spacer length of the crRNA can range from about 10 to 60 nucleotides, for example, 15 to 50 nucleotides, 20 to 50 nucleotides, 25 to 50 nucleotides, or 19 to 50 nucleotides. In some embodiments, the spacer length of the guide RNA is at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, or at least 22 nucleotides. In some embodiments, the spacer length is 15 to 17 nucleotides (e.g., 15, 16, or 17 nucleotides), 17 to 20 nucleotides (e.g., 17, 18, 19, or 20 nucleotides), 20 to 24 nucleotides (e.g., 20, 21, 22, 23, or 24 nucleotides), 23 to 25 nucleotides (e.g., 23, 24, or 25 nucleotides), 24 to 27 nucleotides, 27 to 30 nucleotides, 30 to 45 nucleotides (e.g., 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 nucleotides), 30 or 35 to 40 nucleotides, 41 to 45 nucleotides, 45 to 50 nucleotides (e.g., 45, 46, 47, 48, 49, or 50 nucleotides), or longer. In some embodiments, the spacer length is about 15 to about 42 nucleotides.

[0238] In some embodiments, the length of the direct repeat of the guide RNA is 15 to 36 nucleotides, at least 16 nucleotides, 16 to 20 nucleotides (e.g., 16, 17, 18, 19, or 20 nucleotides), 20 to 30 nucleotides (e.g., 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides), 30 to 40 nucleotides (e.g., 30, 31, 32, 33, 34, 3 5, 36, 37, 38, 39, or 40 nucleotides), or about 36 nucleotides (e.g., 33, 34, 35, 36, 37, 38, or 39 nucleotides). In some embodiments, the length of the direct repeat of the guide RNA is 36 nucleotides.

[0239] In some embodiments, the full length of the crRNA / guide RNA is about 36 nucleotides longer than any one of the spacer sequence lengths described previously herein. For example, the full length of the crRNA / guide RNA can be 45-86 nucleotides, 60-86 nucleotides, 62-86 nucleotides, or 63-86 nucleotides.

[0240] The crRNA sequence can be modified to enable the formation of a complex between the crRNA and the CRISPR-associated protein and successful binding to the target, while at the same time not enabling successful nuclease activity (i.e., no nuclease activity / does not cause indels). These modified guide sequences are referred to as "dead crRNA", "dead guide", or "dead guide sequence". These dead guides or dead guide sequences can be catalytically inactive or inactive in terms of higher-order structure with respect to nuclease activity. Dead guide sequences are typically shorter than each guide sequence that results in active RNA cleavage. In some embodiments, the dead guide is 5%, 10%, 20%, 30%, 40%, or 50% shorter than each guide RNA having nuclease activity. The dead guide sequence of the guide RNA can be 13-15 nucleotides in length (e.g., 13, 14, or 15 nucleotides in length), 15-19 nucleotides in length, or 17-18 nucleotides in length (e.g., 17 nucleotides in length).

[0241] Thus, in one aspect, the present disclosure provides a non-naturally occurring or engineered CRISPR system comprising a functional CRISPR-associated protein and a crRNA described herein, wherein the crRNA, by including a dead crRNA sequence, can hybridize to a target sequence such that the CRISPR system is directed to a genomic locus of interest within a cell without detectable nuclease activity (e.g., RNase activity).

[0242] A detailed description of the dead guide is described, for example, in Pamphlet of International Publication No. 2016 / 094872, and this application is hereby incorporated by reference in its entirety into this specification.

[0243] Guide RNAs (e.g., crRNAs) can be generated as components of an inducible system. The inducible nature of the system enables spatio-temporal control of gene editing or gene expression. In some embodiments, examples of stimuli for the inducible system include electromagnetic radiation, acoustic energy, chemical energy, and / or thermal energy.

[0244] In some embodiments, the transcription of guide RNAs (e.g., crRNAs) can be regulated by inducible promoters, such as tetracycline or doxycycline-controlled transcriptional activation (Tet-On and Tet-Off expression systems), hormone-inducible gene expression systems (e.g., ecdysone-inducible gene expression systems), and arabinose-inducible gene expression systems. Other examples of inducible systems include small molecule 2-hybrid transcriptional activation systems (FKBP, ABA, etc.), light-inducible systems (phytochrome, LOV domain, or cryptochrome), or light-inducible transcriptional effectors (LITE). These inducible systems are described, for example, in Pamphlet of International Publication No. 2016205764 and U.S. Patent No. 8795965, both of which are hereby incorporated by reference in their entirety into this specification.

[0245] Chemical modifications can be made to the phosphate backbone, sugar, and / or base of the crRNA. Pho Skeletal modifications such as phosphorothioates modify the charge on the phosphate backbone and assist in the delivery of oligonucleotides and nuclease resistance (see, e.g., Eckstein, “Phosphorothioates, essential components of therapeutic oligonucleotides”, Nucl. Acid Ther., 24, pp. 374 - 387, 2014); modifications of sugars such as 2'-O-methyl (2'-OMe), 2'-F, and locked nucleic acid (LNA) enhance both base pairing and nuclease resistance (see, e.g., Allerson et al., “Fully 2'-modified oligonucleotide duplexes with improved in vitro potency and stability compared to unmodified small interfering RNA”, J. Med. Chem. 48.4:901 - 904, 2005). Chemically modified bases such as 2-thiouridine or N6-methyladenosine can enable, among other things, stronger or weaker base pairing (see, e.g., Bramsen et al., “Development of therapeutic-grade small interfering RNAs by chemical engineering”, Front. Genet., August 20, 2012; 3:154). Additionally, RNA is applicable to 5' and 3' end conjugations with various functional moieties including fluorescent dyes, polyethylene glycol, or proteins.

[0246] A wide variety of modifications can be applied to chemically synthesized crRNA molecules. For example, modification of oligonucleotides with 2'-OMe to improve nuclease resistance can change the binding energy of Watson-Crick base pairing. Furthermore, 2'-OMe modification can affect the way oligonucleotides interact with transfection reagents, proteins, or any other molecule within a cell. The impact of such modifications can be determined by experimental-based tests.

[0247] In some embodiments, the crRNA comprises one or more phosphorothioate modifications. In some embodiments, the crRNA comprises one or more locked nucleic acids for the purpose of enhancing base pairing and / or increasing nuclease resistance.

[0248] A summary of these chemical modifications can be found, for example, in Kelley et al., “Versatility of chemically synthesized guide RNAs for CRISPR-Cas9 genome editing”, J. Biotechnol. 233:74-83, 2016; WO 2016 / 205764; and U.S. Pat. No. 8,795,965 B2, each of which is incorporated by reference in its entirety.

[0249] The sequences and full lengths of the RNA guides (e.g., crRNA) described herein can be optimized. In some embodiments, the optimized full length of the RNA guide can be determined by identifying the processed form of the crRNA (i.e., mature crRNA) or by full-length studies based on experiments on the crRNA tetraloop.

[0250] In addition, the crRNA may include one or more aptamer sequences. An aptamer is an oligonucleotide molecule or a peptide molecule, has a specific three-dimensional structure, and can bind to a specific target molecule. The aptamer may be specific for a gene effector, a gene activator, or a gene repressor. In some embodiments, the aptamer may be specific for a specific gene effector, gene activator, or gene repressor and specific for a protein that recruits and / or binds to it. The effector, activator, or repressor may exist in the form of a fusion protein. In some embodiments, the guide RNA has two or more aptamer sequences that are specific for the same adapter protein. In some embodiments, the two or more aptamer sequences are specific for different adapter proteins. Examples of adapter proteins may include MS2, PP7, Qβ, F2, GA, fr, JP501, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φkCb5, φkCb8r, φkCb12r, φkCb23r, 7s, and PRR1. Thus, in some embodiments, the aptamer is selected from binding proteins that specifically bind to any one of the adapter proteins described herein. In some embodiments, the aptamer sequence is the MS2 binding loop (5’-ggcccAACAUGAGGAUCACCCAUGUCUGCAGgggcc-3’ (Array No. 93) ). In some embodiments, the aptamer sequence is the QBeta binding loop (5’-ggcccAUGCUGUCUAAGACAGCAUgggcc-3’ (Array No. 94) ). In some embodiments, the aptamer sequence is the PP7 binding loop (5’-ggcccUAAGGGUUUAUAUGGAAACCCUUAgggcc-3’ (Array No. 95)) is. A detailed description of the aptamer can be found, for example, in Nowak et al., “Guide RNA engineering for versatile Cas9 functionality”, Nucl. Acid. Res., 44(20):9555-9564, 2016; and International Publication No. 2016 / 205764 pamphlet, which are hereby incorporated by reference in their entirety.

[0251] In certain embodiments, the method uses a chemically modified guide RNA. Examples of guide RNA chemical modifications include, but are not limited to, incorporation of 2'-O-methyl (M), 2'-O-methyl 3'-phosphorothioate (MS), or 2'-O-methyl 3'-thioPACE (MSP) at one or more terminal nucleotides. Such chemically modified guide RNAs may include increased stability and increased activity compared to unmodified guide RNAs, but on-target versus off-target specificity is not predictable. See Hendel, Nat Biotechnol. 33(9):985-9, 2015 (incorporated by reference). Additionally, chemically modified guide RNAs may include, but are not limited to, RNAs having phosphorothioate linkages and locked nucleic acid (LNA) nucleotides containing a methylene bridge between the 2' and 4' carbons of the ribose ring.

[0252] In addition, the present invention encompasses methods of delivering a plurality of nucleic acid components, each nucleic acid component modifying a plurality of target loci of interest by being specific for various target loci of interest. The nucleic acid components of the complex may include one or more protein-binding RNA aptamers. The one or more aptamers can bind to a bacteriophage coat protein. The bacteriophage coat protein can be selected from the group consisting of Qβ, F2, GA, fr, JP501, MS2, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s, and PRR1. In certain embodiments, the bacteriophage coat protein is MS2.

[0253] 5. Target RNA The target RNA can be any RNA molecule of interest, including naturally occurring RNA molecules and engineered RNA molecules. The target RNA can be mRNA, tRNA, ribosomal RNA (rRNA), microRNA (miRNA), interfering RNA (siRNA), ribozyme, riboswitch, satellite RNA, microswitch, microzyme, or viral RNA.

[0254] In some embodiments, the target nucleic acid is associated with a receptor or a disease (e.g., an infectious disease or cancer). with.

[0255] Thus, in some embodiments, the systems described herein can be used to treat a receptor or a disease by targeting these nucleic acids. For example, as a target nucleic acid associated with a receptor or a disease, there can be an RNA molecule that is overexpressed in diseased cells (e.g., cancer or tumor cells). Also, as a target nucleic acid, there can be toxic RNA and / or mutant RNA (e.g., an mRNA molecule having splicing defects or mutations). Also, as a target nucleic acid, there can be RNA specific to a particular microorganism (e.g., a pathogenic bacterium).

[0256] 6. Complexes and Cells One aspect of the invention provides a CRISPR / Cas13e complex or a CRISPR / Cas13f complex comprising (1) any Cas13e / Cas13f effector protein, homolog, ortholog, fusion, derivative, conjugate, or functional fragment described herein, and (2) any guide RNA described herein (each comprising a spacer sequence designed to be at least partially complementary to a target RNA), and a DR sequence compatible with the Cas13e / Cas13f effector protein, homolog, ortholog, fusion, derivative, conjugate, or functional fragment.

[0257] In certain embodiments, the complex further comprises a target RNA bound by the guide RNA.

[0258] In certain embodiments, the complex is non-naturally occurring. For example, at least one of the components of the complex is non-naturally occurring. In certain embodiments, the Cas13e / Cas13f effector protein, homolog, ortholog, fusion, derivative, conjugate, or functional fragment is non-naturally occurring, for example, due to the presence of at least one amino acid mutation (deletion, insertion, and / or substitution) compared to the wild-type protein. In certain embodiments, the DR sequence is non-naturally occurring, i.e., not any of SEQ ID NOs: 8-14, for example, due to the addition, deletion, and / or substitution of at least one nucleotide base within the wild-type sequence. In certain embodiments, the spacer sequence is non-naturally occurring in that it is either absent or not encoded by any spacer sequence present within the wild-type CRISPR locus of the prokaryote in which the subject Cas13e or Cas13f is present. The spacer sequence may not be naturally occurring if it is not 100% complementary to a naturally occurring bacteriophage nucleic acid.

[0259] In a related aspect, the invention also provides a cell comprising any of the complexes of the invention.

[0260] In certain embodiments, the cell is a prokaryote.

[0261] In certain embodiments, the cell is a eukaryote. When the cell is a eukaryote, the complex within the eukaryotic cell can be the naturally occurring Cas13e / Cas13f complex in the prokaryote from which the Cas13e / Cas13f is isolated.

[0262] 7. Methods of using CRISPR systems The CRISPR systems described herein have a wide variety of utilities, including modifying (e.g., deleting, inserting, translocating, inactivating, or activating) a target polynucleotide or nucleic acid upon multiplexing of cell types. CRISPR systems are useful, for example, for DNA / RNA detection (e.g., Specific High Sensitivity Enzyme Reporter UnLOCK (SHERLOCK)), nucleic acid tracking and labeling, enrichment assays (extracting a desired sequence from a background), control of interfering RNAs or miRNAs, detection of circulating tumor DNA, preparation of next-generation libraries, drug screening, disease diagnosis and prediction, and a wide range of applications in the treatment of various genetic disorders.

[0263] DNA / RNA detection In one aspect, the CRISPR systems described herein can be used for DNA or RNA detection. As shown in the examples, the Cas13e and Cas13f proteins of the present invention exhibit non-specific / collateral RNase activity upon activation of their guide RNA-dependent specific RNase activity when the spacer sequence is about 30 nucleotides. Thus, the CRISPR-associated proteins of the present invention can be reprogrammed by CRISPR RNAs (crRNAs) to provide a platform for specific RNA detection. By selecting a specific spacer sequence length and by recognition of its RNA target, the activated CRISPR-associated protein participates in the “collateral” cleavage of nearby non-target RNA. This crRNA-programmed collateral cleavage activity allows the CRISPR system to detect the presence of a specific RNA by triggering programmed cell death or by non-specific degradation of the labeled RNA.

[0264] The SHERLOCK method (Specific High Sensitivity Enzymatic Reporter UnLOCKing) provides an attomolar sensitivity in vitro nucleic acid detection platform based on nucleic acid amplification and collateral cleavage of reporter RNA, enabling real-time detection of targets. To achieve signal detection, the detection can be combined with various isothermal amplification steps. For example, recombinase polymerase amplification (RPA) can be coupled with T7 transcription to convert the amplified DNA into RNA for subsequent detection. The combination of amplification by RPA, T7 RNA polymerase transcription of the amplified DNA into RNA, and detection of target RNA by collateral RNA cleavage-mediated release of the reporter signal is referred to as SHERLOCK. A method using CRISPR in SHERLOCK is described in detail, for example, in Gootenberg, et al. “Nucleic acid detection with CRISPR-Cas13a / C2c2”, Science, April 28, 2017; 356(6336):438-442, which is hereby incorporated by reference in its entirety.

[0265] CRISPR-related proteins can be used in Northern blot assays, which separate RNA samples by size using electrophoresis. CRISPR-related proteins can be specifically bound to target RNA sequences and detected. In addition, CRISPR-related proteins can be fused to fluorescent proteins (e.g., GFP) and used to track RNA localization in living cells. More specifically, CRISPR-related proteins can be inactivated in that they no longer cleave RNA as described above. Therefore, CRISPR-related proteins can be used to determine the localization of RNA or specific splice variants, the levels of mRNA transcripts, the upregulation or downregulation of transcripts, and disease-specific diagnoses. CRISPR-related proteins can be used, for example, for visualization of RNA in (living) cells using fluorescence microscopy or flow cytometry, such as fluorescence-activated cell sorting (FACS), which enables high-throughput screening of cells and recovery of living cells after cell sorting. A detailed description of how to detect DNA and RNA can be found, for example, in the pamphlet of International Publication No. WO 2017 / 070605, which is hereby incorporated by reference in its entirety.

[0266] In some embodiments, the CRISPR systems described herein can be used in multiplex error-robust fluorescence in situ hybridization (MERFISH). The method is described, for example, in Chen et al., “Spatially resolved, highly multiplexed RNA profiling in single cells”, Science, April 24, 2015; 348(6233):aaa6090, which is hereby incorporated by reference in its entirety.

[0267] In some embodiments, the CRISPR systems described herein can be used to detect target RNAs in a sample (e.g., a clinical sample, a cell, or a cell lysate). The collateral RNase activity of the type VI-E and / or type VI-F CRISPR-Cas effector proteins described herein is activated when the spacer sequence is of a particular selected length (e.g., about 30 nucleotides) and the effector protein binds to the target nucleic acid. By binding to the target RNA of interest, the effector protein cleaves a labeled detector RNA to generate a signal (e.g., an increased signal or a decreased signal), enabling qualitative and quantitative detection of the target RNA in the sample. Specific detection and quantification of RNA in a sample enables a number of applications, including diagnostic methods. In some embodiments, the method comprises contacting the sample with: i) an RNA guide (e.g., crRNA), and / or a nucleic acid encoding the RNA guide; the RNA guide consisting of a direct repeat sequence and a spacer sequence that can hybridize to the target RNA; (ii) a type VI-E or type VI-F CRISPR-Cas effector protein (Cas13e or Cas13f), and / or a nucleic acid encoding the effector protein; and (iii) a labeled detector RNA; the effector protein binds to the RNA guide to form a complex; the RNA guide hybridizes to the target RNA; upon binding of the complex to the target RNA, the effector protein exhibits collateral RNase activity and cleaves the labeled detector RNA; and b) measuring a detection signal generated by cleavage of the labeled detector RNA, wherein said measuring achieves detection of single-stranded target RNA in the sample. In some embodiments, the method further comprises comparing the detection signal to a reference signal and determining the amount of target RNA in the sample. In some embodiments, the measuring is performed using gold nanoparticle detection, fluorescence polarization, colloid phase transition / dispersion, electrochemical detection, and semiconductor-based sensing.In some embodiments, the labeled detector RNA comprises a pair of fluorescent dyes, a fluorescence resonance energy transfer (FRET) pair, or a quencher / fluorophore pair. In some embodiments, cleavage of the labeled detector RNA by an effector protein decreases or increases the amount of the detection signal generated by the labeled detector RNA. In some embodiments, the labeled detector RNA generates a first detection signal before cleavage by the effector protein and a second detection signal after cleavage by the effector protein. In some embodiments, a detection signal is generated when the labeled detector RNA is cleaved by an effector protein. In some embodiments, the labeled detector RNA comprises modified nucleobases, modified sugar moieties, modified nucleic acid linkages, or combinations thereof. In some embodiments, the method comprises multi-channel detection of multiple independent target RNAs (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, or more target RNAs) in a sample using multiple VIE-type and / or VI-F type CRISPR-Cas (Cas13e and / or Cas13f) systems, each comprising a different orthologous effector protein and corresponding RNA guide, enabling differentiation of the multiple target RNAs in the sample. In some embodiments, the method comprises multi-channel detection of multiple independent target RNAs in a sample using multiple examples of VI-E type and / or VI-F type CRISPR-Cas systems, each containing an orthologous effector protein with distinguishable collateral RNase substrates. Methods for detecting RNA in a sample using CRISPR-related proteins are, for example. For example, it is described in US Patent Application Publication No. 2017 / 0362644, the entire content of which is incorporated herein by reference.

[0268] Tracking and Labeling of Nucleic Acids Cell processes are determined by a network of molecular interactions among proteins, RNAs, and DNA. Accurate detection of protein-DNA and protein-RNA interactions is key to understanding such processes. In vitro proximity labeling techniques use an affinity tag combined with a reporter group, e.g., a photoactivatable group, to label polypeptides and RNAs in vitro near a protein or RNA of interest. After UV irradiation, the photoactivatable group reacts with proteins and other molecules proximal to the tagged molecule to label the tagged molecule. Subsequently, the labeled interacting molecules can be recovered and identified. CRISPR-related proteins can be used, for example, to direct a probe to a selected RNA sequence. Also, the applications can be applied to animal models for in vivo imaging of diseases or cell types that are difficult to culture. Methods for tracking and labeling nucleic acids are described, for example, in U.S. Patent No. 8,795,965, International Publication No. 2016 / 205764, and International Publication No. 2017 / 070605, each of which is incorporated herein by reference in its entirety.

[0269] Isolation, purification, enrichment, and / or depletion of RNA RNAs can be isolated and / or purified using the CRISPR systems (e.g., CRISPR-related proteins) described herein. The CRISPR-related proteins can be fused to an affinity tag that can be used to isolate and / or purify the RNA-CRISPR-related protein complex. Such applications are useful, for example, for analyzing gene expression profiles within cells.

[0270] In some embodiments, the activity of a specific non-coding RNA (ncRNA) can be blocked by targeting it with a CRISPR-associated protein. In some embodiments, a CRISPR-associated protein can be used to specifically enrich a specific RNA (including, but not limited to, increasing its stability), or alternatively, a specific RNA (e.g., a specific splice variant, isoform, etc.) can be specifically depleted.

[0271] These methods are described, for example, in U.S. Patent No. 8,795,965, International Publication No. 2016 / 205764, and International Publication No. 2017 / 070605; each of which is hereby incorporated by reference in its entirety.

[0272] High-throughput screening The CRISPR systems described herein can be used to prepare next-generation sequencing (NGS) libraries. For example, to generate a cost-effective NGS library, the CRISPR system can be used to disrupt the coding sequence of a target gene, and clones transfected with the CRISPR-associated protein can be simultaneously screened by next-generation sequencing (e.g., by an Ion Torrent PGM system). A detailed description of how to prepare an NGS library can be found, for example, in Bell et al., “A high-throughput screening strategy for detecting CRISPR-Cas9 induced mutations using next-generation sequencing”, BMC Genomics, 15.1 (2014): 1002, which is hereby incorporated by reference in its entirety.

[0273] Microorganisms to be engineered Microorganisms (e.g., Escherichia coli (E. coli), yeast, and microalgae) are widely used in synthetic biology. The development of synthetic biology has broad utility with various clinical applications. For example, using a programmable CRISPR system, a protein of a toxic domain for target cell death can be split using, for example, cancer-related RNA as a target transcript. Further, a pathway involved in protein-protein interaction can be affected in a synthetic biological system by a fusion complex with an appropriate effector such as a kinase or an enzyme.

[0274] In some embodiments, a crRNA targeting a phage sequence can be introduced into a microorganism. Thus, the present disclosure also provides a method of inoculating a microorganism (e.g., a production strain) against phage infection.

[0275] In some embodiments, the CRISPR systems provided herein can be used to engineer microorganisms, e.g., to improve yield or fermentation efficiency. For example, the CRISPR systems described herein can be used to engineer microorganisms, such as yeast, to produce biofuels or biopolymers from fermentable sugars, or to break down plant-derived lignocellulose from agricultural waste as a source of fermentable sugars. More specifically, the methods described herein can be used to modify the expression of endogenous genes required for biofuel production and / or to modify endogenous genes that may interfere with biofuel synthesis. These methods of engineering microorganisms are described, for example, in Verwaal et al., “CRISPR / Cpf1 enables fast and simple genome editing of Saccharomyces cerevisiae”, Yeast doi:10.1002 / yea.3278, 2017; and Hlavova et al., “Improving microalgae for biotechnology-from genetics to synthetic biology”, Biotechnol. Adv., 33:1194-203, 2015, both of which are hereby incorporated by reference in their entirety.

[0276] In some embodiments, the CRISPR systems provided herein can be used to induce cell death or a quiescent state in cells (e.g., microorganisms, e.g., engineered microorganisms). Using this method, among others, mammalian cells (e.g., cancer cells or tissue culture cells), protozoa, fungal cells, virus-infected cells, intracellular bacteria-infected cells, intracellular protozoa-infected cells, prion-infected cells, bacteria (e.g., pathogenic and non-pathogenic bacteria), protozoa, and prokaryotic and eukaryotic cells including single-celled and multi-celled parasites can be induced to enter a quiescent state or die. For example, in the field of synthetic biology, it is highly desirable to have a mechanism to control engineered microorganisms (e.g., bacteria) that are engineered to prevent growth or spread. The systems described herein can be used as a "kill-switch" to regulate and / or prevent the growth or spread of engineered microorganisms. Additionally, alternatives to current antibiotic treatments are needed in the art. Also, the systems described herein can be used in applications where it is desired to kill or control a specific microbial population (e.g., a bacterial population). For example, the systems described herein may include an RNA guide (e.g., crRNA) that targets a genus-, species-, or strain-specific nucleic acid (e.g., RNA) and can be delivered to a cell. Upon complex formation and binding to the target nucleic acid, the collateral RNase activity of type VI-E and / or type VI-F CRISPR-Cas effector proteins is activated, leading to cleavage of non-target RNA within the microorganism, ultimately resulting in a quiescent state or death. In some embodiments, the method comprises introducing into the cell a type VI-E and / or type VI-F CRISPR-Cas effector protein, or a nucleic acid encoding the effector protein, and RN Contacting the systems described herein with an A guide (e.g., crRNA), or a nucleic acid encoding an RNA guide, wherein the spacer sequence is complementary to at least 15 nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or more nucleotides) of a target nucleic acid (e.g., a genus-, strain-, or species-specific RNA guide). Without wishing to be bound by any particular theory, cleavage of non-target RNA by type VI-E and / or type VI-F CRISPR-Cas effector proteins can induce programmed cell death, cytotoxicity, apoptosis, necrosis, necroptosis, cell death, cell cycle arrest, cell anergy, reduction of cell growth, or reduction of cell proliferation. For example, in bacteria, cleavage of non-target RNA by type VI-E and / or type VI-F CRISPR-Cas effector proteins can be bacteriostatic or bactericidal.

[0277] Use in plants The CRISPR systems described herein have a variety of utilities in plants. In some embodiments, the CRISPR system can be used to engineer the genome of a plant (e.g., to improve production, produce a product with a desired post-translational modification, or introduce a gene for the production of an industrial product). In some embodiments, the CRISPR system can be used to introduce a desired trait into a plant (e.g., with or without heritable modification of the genome), or to regulate the expression of an endogenous gene in a plant cell or plant.

[0278] In some embodiments, the CRISPR system can be used to identify, edit, and / or silence genes encoding specific proteins, such as allergen proteins (e.g., allergen proteins in peanuts, soybeans, lentils, peas, green pod beans, and fava beans). Details regarding how to identify, edit, and / or silence genes encoding proteins are described, for example, in Nicolaou et al., “Molecular diagnosis of peanut and legume allergy”, Curr. Opin. Allergy Clin. Immunol. 11(3):222 - 8, 2011, and International Publication No. WO 2016 / 205764 A1, both of which are hereby incorporated by reference in their entirety.

[0279] Gene drive A gene drive is a phenomenon in which the inheritance of a particular gene or set of genes is advantageously biased. The CRISPR system described herein can be used to construct gene drives. For example, the CRISPR system can be designed to target and disrupt a particular allele of a gene, such that the cell can copy the second allele and fix the sequence. Thanks to the copy, the first allele is converted to the second allele, increasing the likelihood that the second allele will be passed on to the offspring. Details regarding how to use the CRISPR system described herein to construct gene drives are described, for example, in Hammond et al., “A CRISPR - Cas9 gene drive system targeting female reproduction in the malaria mosquito vector Anopheles gambiae”, Nat. Biotechnol. 34(1):78 - 83, 2016, which is hereby incorporated by reference in its entirety.

[0280] Pooled screening The pooled CRISPR screening described herein is a powerful tool for identifying genes involved in biological mechanisms such as cell proliferation, drug resistance, and viral infection. Cells are transduced in large numbers with a library of guide RNA (gRNA) - encoding vectors described herein, and the distribution of gRNAs is measured before and after a selective challenge. The pooled CRISPR screen functions well for mechanisms affecting cell survival and proliferation and can be extended to measure the activity of individual genes (e.g., using engineered reporter cell lines). RNAseq can be used as a read - out with arrayed CRISPR screens, where only one gene is targeted at a time. In some embodiments, the CRISPR systems described herein can be used for single - cell CRISPR screens. A detailed description of pooled CRISPR screening can be found, for example, in Datlinger et al., “Pooled CRISPR screening with single - cell transcriptome read - out”, Nat.Methods.14(3):297 - 301,2017, which is hereby incorporated by reference in its entirety.

[0281] Saturation mutagenesis (bassing) The CRISPR systems described herein can be used for in - situ saturation mutagenesis. In some embodiments, in - situ saturation mutagenesis can be performed for a particular gene or regulatory element using a pooled guide RNA library. Such methods can reveal the important minimal features and distinct vulnerabilities of these genes or regulatory elements (e.g., enhancers). The method is, for example, Canver et al., “BCL11A enhancer dissection by Cas9-mediated in situ saturating mutagenesis”, Nature 527(7577):192-7, 2015, which is hereby incorporated by reference in its entirety.

[0282] RNA-related uses The CRISPR systems described herein can have various RNA-related uses, such as regulating gene expression, degrading RNA molecules, inhibiting RNA expression, screening for RNA or RNA products, determining the function of lincRNA or non-coding RNA, inducing a cell quiescent state, inducing cell cycle arrest, reducing cell growth and / or cell proliferation, inducing cell anergy, inducing apoptosis, inducing necrosis, inducing cell death, and / or inducing programmed cell death. A detailed description of these uses can be found, for example, in WO 2016 / 205764 A1 pamphlet, which application is hereby incorporated by reference in its entirety. In various embodiments, the methods described herein can be performed in vitro, in vivo, or ex vivo.

[0283] For example, the CRISPR systems described herein can be administered to a subject having a disease or disorder to target cells in a diseased state (e.g., cancer cells, or cells infected with an infectious agent) and induce cell death. For example, in some embodiments, the CRISPR systems described herein can be used to target cancer cells and induce cell death, and the cancer cells are derived from a subject having Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, or bladder cancer.

[0284] Regulating gene expression Gene expression can be regulated using the CRISPR systems described herein. The CRISPR system can be used with an appropriate guide RNA to target gene expression through control of RNA processing. Examples of control of RNA processing can include RNA processing reactions such as RNA splicing (e.g., alternative splicing), viral replication, and tRNA biosynthesis. In addition, RNA targeting proteins can be used in combination with an appropriate guide RNA to control RNA activation (RNAa). RNA activation is a small RNA guide and Argonaute (Ago)-dependent gene regulatory phenomenon in which promoter-targeted short double-stranded RNA (dsRNA) induces target gene expression at the transcriptional / epigenetic level. Since RNAa leads to the promotion of gene expression, regulation of gene expression can be achieved by disruption or knockdown of RNAa. In some embodiments, the method includes using RNA-targeted CRISPR as an alternative to, for example, interfering ribonucleic acids (e.g., siRNA, shRNA, or dsRNA). Methods for regulating gene expression are described, for example, in WO 2016 / 205764, which is incorporated herein by reference in its entirety.

[0285] Controlling RNA interference Controlling interfering RNAs or microRNAs (miRNAs) can help reduce off-target effects by shortening the lifespan of interfering RNAs or miRNAs in vivo or in vitro. In some embodiments, interfering RNAs, i.e., RNAs involved in the RNA interference pathway, such as small hairpin RNAs (shRNAs), small interfering (siRNAs), and others, can be used as target RNAs. In some embodiments, examples of target RNAs include miRNAs or double-stranded RNAs (dsRNAs).

[0286] In some embodiments, an RNA-targeting protein and an appropriate guide RNA can be used to protect a cell or system from RNA interference (RNAi) in the cell (in vivo or in vitro) if they are selectively expressed (e.g., spatially or temporally under the control of a regulatory promoter, such as a tissue- or cell cycle-specific promoter and / or enhancer). This can be useful for comparing adjacent tissues or cells where RNAi is not required, or for cells or tissues in which CRISPR-related proteins and appropriate crRNAs are expressed and not expressed (i.e., where RNAi is not controlled and is controlled, respectively). An RNA-targeting protein can be used to control or bind to a molecule that contains or consists of RNA, such as a ribozyme, ribosome, or riboswitch. In some embodiments, the guide RNA can recruit the RNA-targeting protein to the molecule so that the RNA-targeting protein can bind to the molecule. The methods are described, for example, in International Publication No. WO 2016 / 205764 and International Publication No. WO 2017 / 070605, both of which are hereby incorporated by reference in their entirety.

[0287] Modifying riboswitches to control metabolic regulation A riboswitch is a regulatory segment of messenger RNA that binds to a small molecule and then regulates gene expression. By this mechanism, a cell can detect the intracellular concentration of the small molecule. Certain riboswitches typically regulate a gene by altering the transcription, translation, or splicing of an adjacent gene. Thus, in some embodiments, riboswitch activity can be controlled by using an RNA-targeting protein in combination with a guide RNA suitable for targeting the riboswitch. This can be achieved through cleavage of the riboswitch or binding to the riboswitch. Methods of controlling riboswitches using the CRISPR system are described, for example, in International Publication No. WO 2016 / 205764 and International Publication No. WO 2017 / 070605, both of which are hereby incorporated by reference in their entirety.

[0288] RNA Modification In some embodiments, the CRISPR-related proteins described herein can be fused to a base-editing domain, such as ADAR1, ADAR2, APOBEC, or activation-induced cytidine deaminase (AID), and used to modify an RNA sequence (e.g., mRNA). In some embodiments, the CRISPR-related protein contains one or more mutations (e.g., within the catalytic domain), whereby the CRISPR-related protein is unable to cleave RNA.

[0289] In some embodiments, the CRISPR-related protein can be used in an RNA-binding fusion polypeptide that includes a base-editing domain (e.g., ADAR1, ADAR2, APOBEC, or AID) fused to an RNA-binding domain, such as MS2 (also known as the MS2 coat protein), Qbeta (also known as the Qbeta coat protein), or PP7 (also known as the PP7 coat protein). The amino acid sequences of the RNA-binding domains MS2, Qbeta, and PP7 are described below: MS2 (MS2 coat protein) MASNFTQFVLVDNGGTGDVTVAPSNFANGVAEWISSNSRSQAYKVTCSVRQSSAQKRKYTIKVEVPKVATQTVGGVELPVAAWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIY (Array No. 96) Qbeta (Qbeta coat protein) MAKLETVTLGNIGKDGKQTLVLNPRGVNPTNGVASLSQAGAVPALEKRVTVSVSQPSRNRKNYKVQVKIQNPTACTANGSCDPSVTRQAYADVTFSFTQYSTDEERAFVRTELAALLASPLLIDAIDQLNPAY (Array No. 97) PP7 (PP7 coat protein) MSKTIVLSVGEATRTLTEIQSTADRQIFEEKVGPLVGRLRLTASLRQNGAKTAYRVNLKLDQADVVDCSTSVCGELPKVRYTQVWSHDVTIVANSTEASRKSLYDLTKSLVVQATSEDLVVNLVPLGR (Array No. 98)

[0290] In some embodiments, the RNA binding domain recruits an RNA binding fusion polypeptide (which has a base - editing domain) to the effector complex by binding to a specific sequence (e.g., an aptamer sequence) or a secondary structure motif on the crRNA of the systems described herein (e.g., when the crRNA is within the effector - crRNA complex). For example, in some embodiments, the CRISPR system includes a CRISPR - associated protein, a crRNA having an aptamer sequence (e.g., an MS2 binding loop, a QBeta binding loop, or a PP7 binding loop), and an RNA binding fusion polypeptide having a base - editing domain fused to an RNA binding domain that specifically binds to the aptamer sequence. In this system, the CRISPR - associated protein forms a complex with the crRNA having the aptamer sequence. Further, the RNA binding fusion polypeptide binds to the crRNA (via the aptamer sequence) to form a tripartite complex that can modify the target RNA.

[0291] Methods of using the CRISPR system for base editing are described, for example, in WO 2017 / 219027, which is incorporated herein by reference in its entirety, particularly with respect to the discussion of RNA modification.

[0292] RNA splicing In some embodiments, the inactivated CRISPR - associated proteins described herein (For example, a CRISPR-related protein having one or more mutations within a catalytic domain) can be used to target and bind to a specific splicing site on an RNA transcript. Binding of an inactivated CRISPR-related protein to RNA sterically inhibits the interaction of the spliceosome with the transcript, enabling modification of the frequency of generation of a specific transcript isoform. Using such methods, diseases can be treated by exon skipping such that an exon having a mutation can be skipped within the mature protein. Methods of modifying splicing using the CRISPR system are described, for example, in WO 2017 / 219027, which application is hereby incorporated by reference in its entirety, particularly with respect to the discussion of RNA splicing.

[0293] Therapeutic uses The CRISPR systems described herein can have various therapeutic uses. Such uses can be based in both in vitro and in vivo on one or more of the following capabilities of a subject CRISPR / Cas13e or Cas13f system: the ability to induce cell senescence, to induce cell cycle arrest, to inhibit cell growth and / or proliferation, to induce apoptosis, to induce necrosis, and the like.

[0294] In some embodiments, a new CRISPR system can be used to treat various diseases and disorders, such as genetic disorders (e.g., single gene diseases), diseases treatable by nuclease activity (e.g., Pcsk9 targeting, Duchenne muscular dystrophy (DMD), BCL11a targeting), and various cancers and the like.

[0295] In some embodiments, the CRISPR systems described herein can be used to edit a target nucleic acid (e.g., by inserting, deleting, or mutating one or more nucleic acid residues) to modify the target nucleic acid. For example, in some embodiments, the CRISPR systems described herein include a foreign donor template nucleic acid (e.g., a DNA molecule or an RNA molecule) that contains a desired nucleic acid sequence. Through the resolution of cleavage events induced by the CRISPR systems described herein, the cellular machinery will utilize the foreign donor template nucleic acid to repair and / or resolve the cleavage event. Alternatively, the cellular machinery may utilize an endogenous template to repair and / or resolve the cleavage event. In some embodiments, the CRISPR systems described herein are used to modify a target nucleic acid such that insertions, deletions, and / or point mutations can occur. In some embodiments, the insertion is a scarless insertion (i.e., the insertion of an intended nucleic acid sequence into the target nucleic acid such that no additional unintended nucleic acid sequences are generated upon resolution of the cleavage event). The donor template nucleic acid can be a double-stranded or single-stranded nucleic acid molecule (e.g., DNA or RNA). Methods for designing foreign donor template nucleic acids are described, for example, in WO 2016 / 094874 A1, the entire content of which is hereby expressly incorporated by reference.

[0296] In one aspect, the CRISPR systems described herein can be used to treat diseases caused by the overexpression of RNA, toxic RNA, and / or mutant RNA (e.g., splicing defects or truncations). For example, the expression of toxic RNA can be associated with the formation of nuclear inclusions and delayed degenerative changes in the brain, heart, or skeletal muscle. In some embodiments, the disorder is myotonic dystrophy. In myotonic dystrophy, the major pathogenic effect of the toxic RNA is to sequester binding proteins and impair the regulation of alternative splicing (see, e.g., Osborne et al., “RNA-dominant diseases”, Hum. Mol. Genet., 2009 Apr. 15; 18(8):1471-81). Myotonic dystrophy (DM) causes a very broad range of clinical features and is thus of particular interest to geneticists. The classical form of DM, now called DM type 1 (DM1), is caused by an expansion of CTG repeats within the 3’ untranslated region (UTR) of the gene DMPK, which encodes a cytosolic protein kinase. The CRISPR systems described herein can target overexpressed RNA or toxic RNA in DM1 skeletal muscle, heart, or brain, such as the DMPK gene, or any aberrantly regulated alternative splicing.

[0297] Also, the CRISPR systems described herein can target trans-acting mutations that affect RNA-dependent functions that cause various diseases, such as Prader-Willi syndrome, spinal muscular atrophy (SMA), and dyskeratosis congenita. Lists of diseases that can be treated using the CRISPR systems described herein are summarized in Cooper et al., “RNA and disease”, Cell, 136.4 (2009):777-793, and International Publication No. WO 2016 / 205764 A1, both of which are incorporated herein by reference in their entirety. One of ordinary skill in the art will understand how to use the novel CRISPR systems to treat these diseases.

[0298] In addition, the CRISPR systems described herein can be used, for example, in the treatment of various tauopathies, such as primary tauopathy and secondary tauopathy, for example, primary age-related tauopathy (PART) / neurofibrillary tangle (NFT)-predominant Alzheimer's disease (AD) (similar to that seen in AD but without plaques, NFT), boxer dementia (chronic traumatic encephalopathy), and progressive supranuclear palsy. Useful lists of tauopathies and methods of treating these diseases are described, for example, in WO 2016 / 205764, which application is hereby incorporated by reference in its entirety.

[0299] In addition, using the CRISPR systems described herein, mutations that disrupt cis-acting splicing codes that can cause splicing defects and diseases can be targeted. Examples of these diseases include motor neuron degenerative diseases (e.g., spinal muscular atrophy) resulting from deletions in the SMN1 gene, Duchenne muscular dystrophy (DMD), frontotemporal dementia, Parkinson's syndrome associated with chromosome 17 (FTDP-17), and cystic fibrosis.

[0300] The CRISPR systems described herein can further be used, in particular, for antiviral activity against RNA viruses. The CRISPR-associated proteins can target viral RNA using an appropriate guide RNA selected to target the viral RNA sequence.

[0301] In addition, the CRISPR systems described herein can be used to treat cancer in a subject (e.g., a human subject). For example, the CRISPR-associated proteins described herein can be programmed with a crRNA that targets an RNA molecule that is abnormal (e.g., contains a point mutation or is alternatively spliced) and found within cancer cells to induce cell death in the cancer cells (e.g., via apoptosis).

[0302] In addition, the CRISPR systems described herein can be used to treat autoimmune diseases or disorders in a subject (e.g., a human subject). For example, the CRISPR-associated proteins described herein can be programmed with a crRNA that targets an RNA molecule that is abnormal (e.g., contains a point mutation or is alternatively spliced) and is found within a cell responsible for causing an autoimmune disease or disorder.

[0303] Furthermore, the CRISPR systems described herein can be used to treat infectious diseases in a subject. For example, the CRISPR-associated proteins described herein can be programmed with a crRNA that targets an RNA molecule expressed by an infectious agent (e.g., a bacterium, virus, parasite, or protozoan) to target the infectious agent cell and induce cell death in the infectious agent cell. In addition, the CRISPR system can be used to treat diseases in which an intracellular infectious agent has infected a host subject's cells. By programming the CRISPR-associated protein to target an RNA molecule encoded by the infectious agent gene, cells infected with the infectious agent can be targeted and cell death can be induced.

[0304] In addition, an in vitro RNA detection assay can be used to detect a specific RNA substrate. CRISPR-associated proteins can be used for RNA-based detection in living cells. Examples of uses include, for example, diagnostic methods by detecting disease-specific RNAs.

[0305] A detailed description of the therapeutic uses of the CRISPR systems described herein can be found, for example, in U.S. Patent No. 8,795,965, EP 3009511, International Publication No. 2016 / 205764, and International Publication No. 2017 / 070605, each of which is hereby incorporated by reference in its entirety.

[0306] Cells and their progeny In certain embodiments, the methods of the invention are used to introduce the CRISPR systems described herein into cells to modify the production of one or more cell products, such as antibodies, starch, ethanol, or any other desired product, in the cells and / or their progeny. Such cells and their progeny are within the scope of the invention.

[0307] In certain embodiments, the methods and / or CRISPR systems described herein result in the modification of the translation and / or transcription of one or more RNA products of a cell. For example, the modification can result in an increase in the transcription / translation / expression of the RNA product. In other embodiments, the modification can result in a decrease in the transcription / translation / expression of the RNA product.

[0308] In certain embodiments, the cell is a prokaryotic cell.

[0309] In certain embodiments, the cells are eukaryotic cells, such as mammalian cells including human cells (primary human cells or established human cell lines). In certain embodiments, the cells are non-human mammalian cells, such as cells derived from non-human primates (e.g., monkeys), cows, sheep, goats, pigs, horses, dogs, cats, rodents (e.g., rabbits, mice, rats, hamsters, etc.). In certain embodiments, the cells are derived from fish (e.g., salmon), birds (e.g., domestic birds such as chicks, ducks, geese), reptiles, crustaceans (e.g., oysters, clams, lobsters, shrimp), insects, worms, yeast, and others. In certain embodiments, the cells are derived from plants, such as monocotyledonous or dicotyledonous plants. In certain embodiments, the plants are food crops, such as wheat, cassava, cotton, soybean, or peanut, corn, millet, oil palm fruit, potato, legumes, rapeseed or canola, rice, rye, sorghum, soybean, sugarcane, sugar beet, sunflower, and wheat. In certain embodiments, the plants are cereals (wheat, corn, millet, rice, rye, sorghum, and wheat). In certain embodiments, the plants are tubers (cassava and potato). In certain embodiments, the plants are sugar crops (sugar beet and sugarcane). In certain embodiments, the plants are oil-containing crops (soybean, soybean, or peanut, rapeseed or canola, sunflower, and oil palm fruit). In certain embodiments, the plants are fiber crops (cotton). In certain embodiments, the plants are, It is a tree (peach or nectarine tree, apple or pear tree, nut tree, such as almond, walnut, or pistachio tree, or citrus tree, such as orange, grapefruit, or lemon tree, etc.), gramineous plant, vegetable, fruit, or alga. In certain embodiments, the plant is a solanaceous plant; a Brassica genus plant; a Lactuca genus plant; a Spinacia genus plant; a Capsicum genus plant; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, and others.

[0310] Related aspects provide a cell or its progeny modified by the method of the invention using the CRISPR system described herein.

[0311] In certain embodiments, the cell is modified in vitro, in vivo, or ex vivo.

[0312] In certain embodiments, the cell is a stem cell.

[0313] 7. Delivery Based on the present disclosure and the knowledge in the art, the CRISPR system described herein, or any component thereof described herein (Cas protein, its derivative, functional fragment, or various fusions or adducts, and guide RNA / crRNA), its nucleic acid molecule, and / or the nucleic acid molecule encoding or providing its components can be delivered by various delivery systems, such as vectors, e.g., plasmid vectors and viral delivery vectors, using any suitable means in the art. Such methods include, but are not limited to, electroporation, lipofection, microinjection, transfection, sonication, gene gun, and others.

[0314] In certain embodiments, the CRISPR-related protein and / or any RNA (e.g., guide RNA or crRNA) and / or accessory proteins can be delivered using a suitable vector, such as a plasmid or viral vector, such as an adeno-associated virus (AAV), lentivirus, adenovirus, retroviral vector, and other viral vectors, or combinations thereof. The protein and one or more crRNAs can be packaged into one or more vectors, such as a plasmid or viral vector. For bacterial use, a nucleic acid encoding any of the components of the CRISPR system described herein can be delivered to bacteria using a phage. Exemplary phages include, but are not limited to, T4 phage, Mu, λ phage, T5 phage, T7 phage, T3 phage, Φ29, M13, MS2, Qβ, and ΦX174.

[0315] In some embodiments, a vector, such as a plasmid or viral vector, is delivered to the tissue of interest, for example, by intramuscular injection, intravenous administration, transdermal administration, intranasal administration, oral administration, or mucosal administration. Such delivery may be via a single dose or multiple doses. One of ordinary skill in the art will understand that the actual dosage to be delivered herein can vary widely depending on various factors, such as the choice of vector, target cells, organism, tissue, the general condition of the subject to be treated, the degree of transformation / modification required, the route of administration, the mode of administration, the type of transformation / modification required, and others.

[0316] In certain embodiments, the delivery is via an adenovirus, which is made in a single dose containing at least 1×10 5 adenovirus particles (also referred to as particle units pu) In some embodiments, the dose is preferably at least about 1×10 6 particles, at least about 1×10 7 particles, at least about 1×10 8 particles, at least about 1×10 9The particle is an adenovirus. Delivery methods and dosages are described, for example, in International Publication No. 2016 / 205764 Pamphlet A1 and U.S. Patent No. 8,454,972 B2, both of which are hereby incorporated by reference in their entirety.

[0317] In some embodiments, delivery is via a plasmid. The dosage may be the number of plasmids sufficient to induce a response. In some cases, an appropriate amount of plasmid DNA in the plasmid composition can be from about 0.1 to about 2 mg. A plasmid typically includes (i) a promoter; (ii) sequences encoding a nucleic acid-targeting CRISPR-associated protein and / or accessory protein, each operably linked to a promoter (e.g., the same promoter or different promoters); (iii) a selectable marker; (iv) an origin of replication; and (v) a transcription terminator operably linked to (ii) downstream of (ii). Also, a plasmid may encode the RNA components of the CRISPR complex, although one or more of these may instead be encoded on a different vector. The frequency of administration is within the purview of a medical or veterinary practitioner (e.g., a physician, veterinarian), or one of ordinary skill in the art.

[0318] In another embodiment, delivery can be via liposomes or lipofection formulations, etc., prepared by methods known to those of skill in the art. Such methods are described, for example, in International Publication No. 2016 / 205764 Pamphlet, as well as U.S. Patent No. 5,593,972; U.S. Patent No. 5,589,466; and U.S. Patent No. 5,580,859, each of which is hereby incorporated by reference in its entirety.

[0319] In some embodiments, delivery is via nanoparticles or exosomes. For example, exosomes have been shown to be particularly useful for delivering RNA.

[0320] Additional means of introducing one or more components of the new CRISPR system into cells is by using cell-penetrating peptides (CPPs). In some embodiments, the cell-penetrating peptide is linked to a CRISPR-associated protein. In some embodiments, the CRISPR-associated protein and / or guide RNA are linked to one or more CPPs to be effectively transported inside a cell (e.g., a plant protoplast). In some embodiments, the CRISPR-associated protein and / or guide RNA are encoded by one or more circular or linear DNA molecules linked to one or more CPPs for cell delivery.

[0321] A CPP is a short peptide of less than 35 amino acids derived from either a protein or a chimeric sequence that can transport biomolecules across the cell membrane independently of receptors. A CPP can be a cationic peptide, a peptide having a hydrophobic sequence, an amphiphilic peptide, a peptide having a proline-rich sequence and an antimicrobial sequence, and a chimeric peptide or a bipartite peptide. Examples of CPPs include Tat (a transcriptional activator protein required for viral replication by human immunodeficiency virus type 1), penetratin, the Kaposi fibroblast growth factor (FGF) signal peptide sequence, the integrin β3 signal peptide sequence, the polyarginine peptide Arg sequence, the guanine-rich-molecule transporter, and the sweet arrow peptide. CPPs and methods of using them are described, for example, in Haellbrink et al., “Prediction of cell-penetrating peptides”, Methods Mol. Biol., 2015;1324:39-58; Ramakrishna et al., “Gene disruption by cell-penetrating peptide-mediated delivery of Cas9 protein and guide RNA”, Ge Nome Res., 2014 June; 24(6): 1020 - 7; and International Publication No. WO 2016 / 205764 A1 pamphlet, which are each incorporated herein by reference in their entirety.

[0322] Also, various delivery methods of the CRISPR systems described herein are described, for example, in U.S. Patent No. 8,795,965, EP 3009511, International Publication No. WO 2016 / 205764 pamphlet, and International Publication No. WO 2017 / 070605 pamphlet, which are each incorporated herein by reference in their entirety.

[0323] 8. Kit Another aspect of the present invention provides a kit comprising any two or more components of the subject CRISPR / Cas systems described herein, such as Cas13e and Cas13f proteins, their derivatives, functional fragments, or various fusions or adducts, guide RNA / crRNA, their complexes, vectors containing them, or hosts containing them.

[0324] In certain embodiments, the kit further includes instructions for use of the components contained therein and / or instructions for combination with additional components available elsewhere.

[0325] In certain embodiments, the kit further includes one or more nucleotides, for example, nucleotides useful for inserting a guide RNA coding sequence into a vector and operably linking the coding sequence to one or more control elements of the vector.

[0326] In certain embodiments, the kit further includes one or more buffers that can be used to lyse any of the components and / or provide reaction conditions suitable for one or more of the components. Such buffers may include one or more of PBS, HEPES, Tris, MOP, Na2CO3, NaHCO3, NaB, or combinations thereof. In certain embodiments, the reaction conditions include an appropriate pH, such as a basic pH. In certain embodiments, the pH is from 7 to 10.

[0327] In certain embodiments, any one or more of the kit components may be stored in a suitable container. Non-limitingly, the present invention includes the following aspects. [Aspect 1] (1) An RNA guide sequence comprising a spacer sequence capable of hybridizing to a target RNA and a direct repeat (DR) sequence on the 3' side with respect to the spacer sequence; and (2) A clustered regularly interspaced short palindromic repeat (CRISPR)-Cas complex comprising a CRISPR-associated protein (Cas) having an amino acid sequence of any one of SEQ ID NOs: 1 to 7, or a derivative or functional fragment of the Cas wherein; the Cas, the derivative of the Cas, and the functional fragment are capable of (i) binding to the RNA guide sequence and (ii) targeting the target RNA, wherein the spacer sequence is a CRISPR-Cas complex on the condition that it is not 100% complementary to a naturally occurring bacteriophage nucleic acid when the complex comprises any one of the Cas of SEQ ID NOs: 1 to 7. [Aspect 2] The CRISPR-Cas complex according to Aspect 1, wherein the DR sequence has a secondary structure substantially the same as the secondary structure of any one of SEQ ID NOs: 8 to 14. [Aspect 3] The CRISPR-Cas complex according to Aspect 1, wherein the DR sequence is encoded by any one of SEQ ID NOs: 8 to 14. [Aspect 4] The target RNA is the CRISPR-Cas complex according to any one of aspects 1 to 3, which is encoded by eukaryotic DNA. [Aspect 5] The eukaryotic DNA is non-human mammalian DNA, non-human primate DNA, human DNA, plant DNA, insect DNA, avian DNA, reptilian DNA, rodent DNA, fish DNA, worm / nematode DNA, yeast DNA, and is the CRISPR-Cas complex according to aspect 4. [Aspect 6] The target RNA is mRNA, and is the CRISPR-Cas complex according to any one of aspects 1 to 5. [Aspect 7] The spacer sequence is 15 to 60 nucleotides, 25 to 50 nucleotides, or about 30 nucleotides, and is the CRISPR-Cas complex according to any one of aspects 1 to 6. [Aspect 8] The spacer sequence is 90 to 100% complementary to the target RNA, and is the CRISPR-Cas complex according to any one of aspects 1 to 7. [Aspect 9] The derivative includes conservative amino acid substitutions of one or more residues of any one of SEQ ID NOs: 1 to 7, and is the CRISPR-Cas complex according to any one of aspects 1 to 8. [Aspect 10] The derivative includes only conservative amino acid substitutions, and is the CRISPR-Cas complex according to aspect 9. [Aspect 11] The derivative has the same sequence as any one of the wild-type Cas of SEQ ID NOs: 1 to 7 within the HEPN domain or the RXXXXH motif, and is the CRISPR-Cas complex according to any one of aspects 1 to 10. [Aspect 12] The derivative can bind to the RNA guide sequence that hybridizes to the target RNA, but does not have RNase catalytic activity due to mutations within the RNase catalytic site of the Cas, and is the CRISPR-Cas complex according to any one of aspects 1 to 9. [Aspect 13] The derivative is the CRISPR-Cas complex according to aspect 12, having an N-terminal deletion of 210 residues or less and / or a C-terminal deletion of 180 residues or less. [Aspect 14] The derivative is the CRISPR-Cas complex according to aspect 13, having an N-terminal deletion of about 180 residues and / or a C-terminal deletion of about 150 residues. [Aspect 15] The derivative is further the CRISPR-Cas complex according to any one of aspects 12 to 14, comprising an RNA base-editing domain. [Aspect 16] The RNA base-editing domain is adenosine deaminase, such as double-stranded RNA-specific adenosine deaminase (e.g., ADAR1 or ADAR2); apolipoprotein B mRNA editing enzyme; catalytic polypeptide-like (APOBEC); or activation-induced cytidine deaminase (AID), the CRISPR-Cas complex according to aspect 15. [Aspect 17] The ADAR2 has an E488Q / T375G double mutation or is ADAR2DD, the CRISPR-Cas complex according to aspect 16. [Aspect 18] The base-editing domain is further fused to an RNA binding domain, such as MS2, the CRISPR-Cas complex according to any one of aspects 15 to 17. [Aspect 19] The derivative is further the CRISPR-Cas complex according to any one of aspects 12 to 14, comprising an RNA methyltransferase, an RNA demethylase, an RNA splicing modifier, a localization factor, or a translation modifier. [Aspect 20] The Cas, the derivative, or the functional fragment comprises a nuclear localization signal (NLS) sequence or a nuclear export signal (NES), the CRISPR-Cas complex according to any one of aspects 1 to 19. [Aspect 21] The targeting of the target RNA results in the modification of the target RNA, the CRISPR-Cas complex according to any one of aspects 1 to 20. [Aspect 22] The CRISPR-Cas complex according to aspect 21, wherein the modification of the target RNA is cleavage of the target RNA. [Aspect 23] The CRISPR-Cas complex according to aspect 21, wherein the modification of the target RNA is deamination of adenosine (A) to inosine (I). [Aspect 24] The CRISPR-Cas complex according to any one of aspects 1 to 23, further comprising a target RNA comprising a sequence capable of hybridizing to the spacer sequence. [Aspect 25] (1) Cas, a derivative thereof, or a functional fragment thereof according to any one of aspects 1 to 24, and (2) a fusion protein comprising a heterologous functional domain. [Aspect 26] The heterologous functional domain is: a nuclear localization signal (NLS), a reporter protein or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP), a localization signal, a protein targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD), an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc.), a transcriptional activation domain (e.g., VP64 or VPR), a transcriptional repression domain (e.g., KRAB moiety or SID moiety), a nuclease (e.g., FokI), a deamination domain (e.g., ADAR1, ADAR2, APOBEC, AID, or TAD), a methylase, a demethylase, a transcriptional release factor, HDAC, a polypeptide having ssRNA cleavage activity, a polypeptide having dsRNA cleavage activity, a polypeptide having ssDNA cleavage activity, a polypeptide having dsDNA cleavage activity, a DNA or RNA ligase, or any combination thereof. The fusion protein according to aspect 25. [Aspect 27] The heterologous functional domain is fused to the N-terminus, C-terminus, or internally within the fusion protein. The fusion protein according to aspect 25 or 26. [Aspect 28] A conjugate comprising a Cas, a derivative thereof, or a functional fragment thereof according to any one of aspects 1 to 24, conjugated to a heterologous functional moiety. [Aspect 29] The heterologous functional moiety is: a nuclear localization signal (NLS), a reporter protein or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP), a localization signal, a protein targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD), an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx and others), a transcriptional activation domain (e.g., VP64 or VPR), a transcriptional repression domain (e.g., the KRAB moiety or the SID moiety), a nuclease (e.g., FokI), a deamination domain (e.g., ADAR1, ADAR2, APOBEC, AID, or TAD), a methylase, a demethylase, a transcriptional release factor, HDAC, a polypeptide having ssRNA cleavage activity, a polypeptide having dsRNA cleavage activity, a polypeptide having ssDNA cleavage activity, a polypeptide having dsDNA cleavage activity, a DNA or RNA ligase, or any combination thereof, the conjugate according to aspect 28. [Aspect 30] The heterologous functional moiety is conjugated to the Cas, the derivative thereof, or the functional fragment thereof at the N-terminus, C-terminus, or internally, the conjugate according to aspect 28 or 29. [Aspect 31] A polynucleotide encoding any one of SEQ ID NOs: 1 to 7, or a derivative thereof, or a functional fragment thereof, or a fusion protein thereof, provided that it is not any of SEQ ID NOs: 15 to 21. [Aspect 32] The polynucleotide according to aspect 31, which is codon-optimized for expression in cells. [Aspect 33] The polynucleotide according to aspect 32, wherein the cell is a eukaryotic cell. [Aspect 34] A non-naturally occurring polynucleotide comprising any one derivative of SEQ ID NOs: 8 to 14, wherein the derivative has (i) an addition, deletion, or substitution of one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) nucleotides as compared to any one of SEQ ID NOs: 8 to 14; (ii) at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 97% sequence identity to any one of SEQ ID NOs: 8 to 14; (iii) hybridizes with any one of SEQ ID NOs: 8 to 14, or any one of (i) and (ii) under stringent conditions; or (iv) is a complement of any one of (i) to (iii), provided that the derivative is not any of SEQ ID NOs: 8 to 14, and the derivative encodes (or is) an RNA that maintains a substantially the same secondary structure as any of the RNAs encoded by SEQ ID NOs: 8 to 14, a non-naturally occurring polynucleotide. [Aspect 35] The derivative functions as a DR sequence of any one of the Cas, its derivative, or its functional fragment described in any one of Aspects 1 to 24, the non-naturally occurring polynucleotide described in Aspect 34. [Aspect 36] A vector comprising the polynucleotide described in any one of Aspects 31 to 35. [Aspect 37] The polynucleotide is operably linked to a promoter and optionally an enhancer, the vector described in Aspect 36. [Aspect 38] The promoter is a constitutive promoter, an inducible promoter, a ubiquitin promoter, or a tissue-specific promoter, the vector described in Aspect 37. [Aspect 39] A plasmid, the vector described in any one of Aspects 36 to 38. [Aspect 40] The vector according to any one of aspects 36 to 38, which is a retroviral vector, phage vector, adenoviral vector, herpes simplex virus (HSV) vector, AAV vector, or lentiviral vector. [Aspect 41] The vector according to aspect 40, wherein the AAV vector is a recombinant AAV vector of serotype AAV1, AAV2, AAV4, AAV5, AAV6, AAV7, AAVrh74, AAV8, AAV9, AAV10, AAV11, AAV12, or AAV13. [Aspect 42] (1) A delivery vehicle, and (2) a CRISPR-Cas complex according to any one of aspects 1 to 24, a fusion protein according to any one of aspects 25 to 27, a conjugate according to any one of aspects 28 to 30, a polynucleotide according to any one of aspects 31 to 33, or a vector according to any one of aspects 36 to 41. [Aspect 43] The delivery system according to aspect 42, wherein the delivery vehicle is a nanoparticle, liposome, exosome, microvesicle, or gene gun. [Aspect 44] A cell or its progeny comprising a CRISPR-Cas complex according to any one of aspects 1 to 24, a fusion protein according to any one of aspects 25 to 27, a conjugate according to any one of aspects 28 to 30, a polynucleotide according to any one of aspects 31 to 33, or a vector according to any one of aspects 36 to 41. [Aspect 45] The cell or its progeny according to aspect 44, which is a eukaryotic cell (e.g., a non-human mammalian cell, a human cell, or a plant cell), or a prokaryotic cell (e.g., a bacterial cell). [Aspect 46] A non-human multicellular eukaryote comprising the cell according to aspect 44 or 45. [Aspect 47] The non-human multicellular eukaryote according to aspect 46, which is an animal (e.g., a rodent or a primate) model for a human genetic disorder. [Aspect 48] A method of modifying a target RNA, comprising contacting the target RNA with a CRISPR-Cas complex according to any one of aspects 1 to 24, wherein the spacer sequence is complementary to at least 15 nucleotides of the target RNA; the Cas, the derivative, or the functional fragment binds to the RNA guide sequence to form the complex; the complex binds to the target RNA; and binding of the complex to the target RNA causes the Cas, the derivative, or the functional fragment to modify the target RNA. [Aspect 49] The method according to aspect 48, wherein the target RNA is modified by cleavage by the Cas. [Aspect 50] The method according to aspect 48, wherein the target RNA is modified by deamination by a derivative comprising a double-stranded RNA-specific adenosine deaminase. [Aspect 51] The method according to any one of aspects 48 to 50, wherein the target RNA is mRNA, tRNA, rRNA, non-coding RNA, lncRNA, or nuclear RNA. [Aspect 52] The method according to any one of aspects 48 to 51, wherein binding of the complex to the target RNA causes the Cas, the derivative, and the functional fragment to exhibit no substantial (or detectable) collateral RNase activity. [Aspect 53] The method according to any one of aspects 48 to 52, wherein the target RNA is intracellular. [Aspect 54] The method according to aspect 53, wherein the cell is a cancer cell. [Aspect 55] The method according to aspect 53, wherein the cell is infected with an infectious agent. [Aspect 56] The method according to aspect 55, wherein the infectious agent is a virus, prion, protozoan, fungus, or parasite. [Aspect 57] The CRISPR-Cas complex is encoded by a first polynucleotide encoding any one of SEQ ID NOs: 1 to 7, or a derivative or functional fragment thereof, and a second polynucleotide comprising any one of SEQ ID NOs: 8 to 14 and a sequence encoding a spacer RNA capable of binding to the target RNA, wherein the first polynucleotide and the second polynucleotide are introduced into the cell by the method according to any one of aspects 53 to 56. [Aspect 58] The method according to aspect 57, wherein the first polynucleotide and the second polynucleotide are introduced into the cell by the same vector. [Aspect 59] (i) in vitro or in vivo induction of cell senescence; (ii) in vitro or in vivo cell cycle arrest; (iii) in vitro or in vivo inhibition of cell proliferation and / or cell growth inhibition; (iv) in vitro or in vitro induction of anergy; (v) in vitro or in vitro induction of apoptosis; and (vi) the method according to any one of aspects 53 to 58, which causes one or more of in vitro or in vitro induction of necrosis. [Aspect 60] A method for treating a receptor or a disease in a subject in need of treatment of the receptor or the disease, comprising administering to the subject a composition comprising the CRISPR-Cas complex according to any one of aspects 1 to 24, or a polynucleotide encoding the same; the spacer sequence is complementary to at least 15 nucleotides of the target RNA associated with the receptor or the disease; the Cas, the derivative, or the functional fragment binds to the RNA guide sequence to form the complex; the complex binds to the target RNA; and binding of the complex to the target RNA causes the Cas, the derivative, or the functional fragment to cleave the target RNA to treat the receptor or the disease in the subject. [Aspect 61] The method according to aspect 60, wherein the receptor or the disease is cancer or an infectious disease. [Aspect 62] The cancer is Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, or bladder cancer, the method according to aspect 61. [Aspect 63] The method according to any one of aspects 60 to 62, which is an in vitro method, an in vivo method, or an ex vivo method. [Aspect 64] Cells or their progeny obtained by the method according to any one of aspects 48 to 59, which cells and their progeny contain a modification that does not occur naturally (for example, a modification that does not occur naturally in the transcribed RNA of the cells / progeny). [Aspect 65] A method for detecting the presence of a target RNA, comprising contacting the target RNA with a composition comprising a fusion protein according to any one of aspects 25 to 27, or a conjugate according to any one of aspects 28 to 30, or a polynucleotide encoding the fusion protein, wherein the fusion protein or the conjugate comprises a detectable label (for example, one that can be detected by fluorescence, Northern blot, or FISH) and a synthetic spacer sequence capable of binding to the target RNA. [Aspect 66] A eukaryotic cell comprising a clustered regularly interspaced short palindromic repeat (CRISPR)-Cas complex, wherein the CRISPR-Cas complex: (1) an RNA guide sequence comprising a spacer sequence capable of hybridizing to a target RNA and a direct repeat (DR) sequence on the 3' side of the spacer sequence; and (2) a CRISPR-associated protein (Cas) having an amino acid sequence of any one of SEQ ID NOs: 1 to 7, or a derivative or functional fragment of the Cas. The Cas, the derivative of the Cas, and the functional fragment are eukaryotic cells that can (i) bind to the RNA guide sequence and (ii) target the target RNA.

Example

[0328] Example 1 Identification of novel Cas13e and Cas13f systems Using a computer pipeline, an extended database of class 2 CRISPR-Cas systems was generated from genomic and metagenomic sources. Genomic and metagenomic sequences were downloaded from NCBI (Benson et al., 2013; Pruitt et al., 2012), NCBI whole genome sequencing (WGS), and DOE JGI Integrated Microbial Genomes (Markowitz et al., 2012). Proteins were predicted in all contigs at least 5 kb in length (Prodigal in anon mode (Hyatt et al., 2010)), and duplicates were removed (i.e., identical protein sequences were removed) to construct a complete protein database. Proteins larger than 600 residues were considered large proteins (LPs). Since most currently identified Cas13 proteins are larger than 900 residue size, only large proteins were further considered to reduce the computational complexity.

[0329] CRISPR arrays were identified using Piler-CR (Ed gar, PILER-CR: Fast and accurate identification of CRISPR repeats. BMC Bioinformatics 8:18, 2007) with all default parameters. Non-redundant large protein sequence-encoding ORFs located within ±10 kb from the CRISPR arrays were classified into CRISPR-proximal large protein-encoding clusters, and the encoded LPs were defined as Cas-LPs.

[0330] First, pairwise alignment between Cas-LPs was performed using BLASP to obtain BLASTP alignment results with an Evalue < 1E-10. Subsequently, MCL was used to further cluster Cas-LPs based on the BLASTP results, generating families of Cas proteins.

[0331] Next, BLASTP was used to align Cas-LPs against all LPs to obtain BLASTP alignment results with an Evalue < 1E-10. The Cas-LP families were further expanded according to the BLASTP alignment results. After expansion, Cas-LP families with a doubling or less were obtained for further analysis.

[0332] For the functional characterization of candidate Cas proteins, candidate Cas proteins were annotated using the protein family database Pfam (Finn et al., 2014), the NR database, and Cas proteins in NCBI. Subsequently, multiple sequence alignment was performed for each candidate Cas effector protein using MAFFT (Katoh and Standley, 2013). Subsequently, JPred and HHpred were used to analyze conserved regions within these proteins to identify candidate Cas proteins / families with two conserved RXXXXH motifs.

[0333] This analysis led to the identification of seven new Cas13 effector proteins belonging to two new Cas13 families that are different from all previously identified class 2 CRISPR-Cas systems. These include Cas13e.1 (SEQ ID NO: 1) and Cas13e.2 (SEQ ID NO: 2) of the new Cas13e family, and Cas13f.1 (SEQ ID NO: 3), Cas13f.2 (SEQ ID NO: 4), Cas13f.3 (SEQ ID NO: 5), Cas13f.4 (SEQ ID NO: 6), and Cas13f.5 (SEQ ID NO: 7) of the new Cas13f family.

Chemical formula

Chem.

Chem.

[0334] The DNAs encoding the corresponding direct repeat (DR) sequences within each pre-crRNA sequence are SEQ ID NOs: 8 to 14, respectively. GCTGGAGCAGCCCCCGATTTGTGGGGTGATTACAGC (SEQ ID NO: 8) GCTGAAGAAGCCTCCGATTTGAGAGGTGATTACAGC (SEQ ID NO: 9) GCTGTGATAGACCTCGATTTGTGGGGTAGTAACAGC (SEQ ID NO: 10) GCTGTGATAGACCTCGATTTGTGGGGTAGTAACAGC (SEQ ID NO: 11) GCTGTGATAGACCTCGATTTGTGGGGTAGTAACAGC (SEQ ID NO: 12) GCTGTGATGGGCCTCAATTTGTGGGGAAGTAACAGC (SEQ ID NO: 13) GCTGTGATAGGCCTCGATTTGTGGGGTAGTAACAGC (SEQ ID NO: 14)

[0335] The native (wild-type) DNA coding sequences of Cas13e.1, Cas13e.2, Cas13f.1, Cas13f.2, Cas13f.3, Cas13f.4, and Cas13f.5 proteins are SEQ ID NOs: 15 to 21, respectively.

Chem.

Chem.

Chem.

Chem.

Chem.

[0336] The human codon-optimized coding sequences of seven Cas13e and Cas13f proteins (i.e., Cas13e.1, Cas13e.2, Cas13f.1, Cas13f.2, Cas13f.3, Cas13f.4, and Cas13f.5) generated for further functional experiments are SEQ ID NOs: 22 to 28, respectively.

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

[0337] Seven CRISPR / Cas13e and Cas13f locus structures are shown in FIG. 1.

[0338] Further analysis of the RNA secondary structures of the seven DR sequences within the pre-crRNA was performed using RNAfold. The results are shown in FIG. 2. It is clear that all shared a highly conserved secondary structure.

[0339] For example, in the Cas13e family, each DR sequence forms a secondary structure consisting of a symmetric bulge of 5 + 5 nucleotides (excluding 4 stem nucleotides) following a 4-base pair stem (5'-GCUG-3'), followed by a 5-base pair stem (5'-GCC C / U C-3'), and a terminal 8-base loop (5'-CGAUUUGU-3', excluding 2 stem nucleotides).

[0340] Similarly, in the Cas13f family (with one exception (Cas13f.4)) , each DR array forms a secondary structure consisting of a nearly symmetric bulge of 5 + 4 nucleotides (excluding 4 stem nucleotides) following a 5-base pair stem (5’-GCUGU3’), followed by a 6-base pair stem (5’A / G CCUCG3’), and a terminal 5-base loop (5’AUUUG3’, excluding 2 stem nucleotides). The only exception is the DR of Cas13f.4, where the second step is 1 base pair shorter and an additional 2 bases are added to the first bulge, forming a mostly symmetric 6 + 5 bulge.

[0341] Multi-sequence alignment of Cas13e and Cas13f proteins, and previously identified Cas13a, Cas13b, Cas13c, and Cas13d family proteins using MAFFT revealed that Cas13e and Cas13f proteins are relatively closest to Cas13b protein on the phylogenetic tree (Figure 3).

[0342] Furthermore, regarding the positions of the RXXXXH motif with respect to the N-terminus and C-terminus of the Cas proteins, Cas13e and Cas13f proteins, and to a lesser extent Cas13b protein, the RXXXXH motif is closer to the N-terminus and C-terminus compared to Cas13a, Cas13c, and Cas13d (see Figure 4).

[0343] Subsequently, the 3D structure of the Cas13e protein was predicted using I-TASSER, and then the predicted structure was visualized using PyMOL. The two RXXXXH motifs are positioned very close to the N-terminus and C-terminus of Cas13e.1 and are very close in the 3D structure (Figure 5).

[0344] Example 2 Cas13e is an effector RNase To confirm that the newly identified Cas13e protein is a functional RNase in the CRISPR / Cas system, the Cas13e.1 coding sequence was codon-optimized for human expression (SEQ ID NO: 22) and cloned into a first plasmid having a GFP gene. On the other hand, the coding sequence of a guide RNA (gRNA) targeting the reporter gene (mCherry) mRNA was cloned into a second plasmid having a GFP gene. The gRNA consists of a spacer coding region flanked by two direct repeat sequences for Cas13e.1 (SEQ ID NO: 29). The sequences of the GFP and mCherry reporter genes are SEQ ID NOs: 30-31, respectively.

Chem.

[0345] HEK293T cells were cultured in a 24-well tissue culture plate according to a standard protocol for triple plasmid transfection using LIPOFECTAMINE® 3000 and P3000™ reagents to introduce three plasmids encoding Cas13e.1 protein, mCherry-targeting gRNA, and the mCherry coding sequence, respectively. In the negative control experiment, a control plasmid encoding a non-target gRNA was used instead of the plasmid encoding mCherry-targeting gRNA. Since the GFP coding sequence is present in the Cas13e.1 and gRNA plasmids, the expression of GFP can be used as an internal control for the success / efficiency of transfection. See the schematic diagram in Figure 6. Subsequently, the transfected HEK293T cells were incubated at 37 °C under 5% CO2 for about 24 hours and then the cells were tested under a fluorescence microscope.

[0346] As shown in Fig. 7, cells transfected with mCherry-targeted gRNA and cells transfected with control non-targeted (NT) gRNA had equal growth and morphology under bright-field microscopy, and GFP expression in both was mostly equal. However, the RFP signal derived from mCherry expression decreased by up to 75% based on flow cytometry analysis (Fig. 8). This suggests that Cas13e can efficiently knockdown mCherry mRNA levels and consequently mCherry protein expression using mCherry-targeted gRNA.

[0347] Example 3 Effective direction of sgRNA for Cas13e Since the Cas13e system can theoretically utilize either the DR+spacer (5’DR) or spacer+DR (3’DR) orientation, this experiment was designed to determine which is the exact orientation utilized by Cas13e.

[0348] Using a triple-transfection experimental setup similar to Example 2, it was found that only the 3’DR orientation (spacer+DR) supported significant mCherry knockdown. This demonstrated that Cas13e utilizes crRNAs with a DR sequence at the 3’ end of the spacer. See Fig. 9.

[0349] The SgRNAs for DR+spacer (5’DR) and spacer+DR (3’DR) are SEQ ID NO: 32 and SEQ ID NO: 33, respectively. GCTGGAGCAGCCCCCGATTTGTGGGGTGATTACAGCGGTCTTCGATATTCAAGCGTCGGAAGACCT (SEQ ID NO: 32) GGTCTTCGATATTCAAGCGTCGGAAGACCTGCTGGAGCAGCCCCCGATTTGTGGGGTGATTACAGC (SEQ ID NO: 33)

[0350] Example 4 Influence of spacer sequence length on the specific and collateral activities of Cas13e.1 To study the effect of spacer sequence length on the specific and collateral activities of Cas13e.1, a set of sgRNAs targeting the mCherry reporter gene with spacer sequence lengths of 20 nt, 25 nt, 30 nt, 35 nt, 40 nt, 45 nt, or 50 nt were designed (SEQ ID NOs: 34 - 40). TTGGTGCCGCGCAGCTTCAC (SEQ ID NO: 34) TTGGTGCCGCGCAGCTTCACCTTGT (SEQ ID NO: 35) TTGGTGCCGCGCAGCTTCACCTTGTAGATG (SEQ ID NO: 36) TTGGTGCCGCGCAGCTTCACCTTGTAGATGAACTC (SEQ ID NO: 37) TTGGTGCCGCGCAGCTTCACCTTGTAGATGAACTCGCCGT (SEQ ID NO: 38) TTGGTGCCGCGCAGCTTCACCTTGTAGATGAACTCGCCGTCCTGC (SEQ ID NO: 39) TTGGTGCCGCGCAGCTTCACCTTGTAGATGAACTCGCCGTCCTGCAGGGA (SEQ ID NO: 40)

[0351] Using a triple transfection experimental setup similar to Example 2, the knockdown efficiencies of the mCherry and GFP genes were analyzed by flow cytometry.

[0352] The results of the mCherry and GFP knockdown experiments showed the specific and non-specific (collateral) activities of Cas13e.1, respectively. Cas13e.1 was found to have high specific activity with spacer lengths of approximately 30 nt to approximately 50 nt. See Figure 10. On the other hand, Cas13e.1 had the highest non-specific activity when the spacer length was approximately 30 nt. See Figure 11.

[0353] Single-base RNA editing using the dCas13e.1-ADAR2DD fusion of Example 5 To test whether Cas13e can be used for single-base RNA editing, dCas13e.1 was generated by mutating two RXXXXH motifs to eliminate RNase activity. Subsequently, a high-fidelity ADAR2DD mutant with double mutations of E488Q and T375G was fused to the C-terminus of dCas13e.1 to generate a putative single-base RNA editor from A to G, named dCas13e.1-ADAR2DD. See the coding sequence within SEQ ID NO: 41.

Chemical formula

[0354] To be used as a target for the putative RNA base-editor, the wild-type mCherry coding sequence was mutated to generate an early termination codon TAG (see the sequence underlined in bold in SEQ ID NO: 42) so that a functional mCherry protein is not generated without being corrected from A to G by the RNA base editor. See FIGS. 12 and 14. Subsequently, gRNAs were designed to generate the desired A-to-G edits (FIGS. 12 and 14), and the CX530 plasmid encoding the dCas13e.1-ADAR2DD base editor, the CX537 / Cx538 plasmids encoding the sgRNAs, and the CX337 plasmid encoding the mutated mCherry gene were triple-transfected into HEK293T cells using standard protocols. The transfected HEK293T cells were incubated at 37 °C under 5% CO2 for 24 hours, and then the cells were subjected to flow cytometry to isolate cells that have corrected mCherry mRNA and express the mCherry protein. See the explanatory drawing of FIG. 12. The results of the flow cytometry analysis are shown in FIG. 13. isolated. See the explanatory drawing of FIG. 12. The results of the flow cytometry analysis are shown in FIG. 13.

[0355] It is clear that both gRNA-1 (SEQ ID NO: 43) and gRNA-2 (SEQ ID NO: 44) successfully correct the TAG early termination codon to generate a functional mCherry protein.

Chemical formula

[0356] Example 6 Single-base RNA editing using short dCas13e.1-ADAR2DD fusions To determine the minimum size of dCas13e.1 that can be used for RNA single-base editing, a series of five constructs were generated that expressed progressively larger C-terminal deletions of dCas13e.1, each having less than 30 residues from the C-terminus (i.e., 30-, 60-, 90-, 120-, and 150-residue deletions). The resulting constructs were used to generate coding sequences of dCas13e.1 that fuse to high-fidelity adar2 (ADAR2DD) at each C-terminus. These constructs were cloned into the Vysz15 (“V15”) to Vysz-19 (“V19”) plasmids (Figure 15) used in experiments similar to Example 4. In all of these constructs, the fusion protein was expressed from the CMV promoter (pCMV) and an enhancer (eCMV) that further enhanced protein expression, immediately downstream of the intron. Two nuclear localization signals (NLS) were placed at the N-terminus and C-terminus of the dCas13e.1 portion of the fusion, fusing the ADAR2 domain (e.g., ADAR2DD) to the C-terminal NLS via an NLS linker and tagging the C-terminus with an HA-tag. The EGFP coding sequence under the independent control of the EFS promoter (pEFS) was present downstream of the polyA addition sequence of all plasmids.

[0357] Interestingly, progressive C-terminal deletions were found to continuously increase the RNA base editing activity within the fusion editor, such that the editor with a 150 C-terminal residue deletion (within V19) exhibited the highest base editing activity. See Figure 16. However, a 180-residue deletion from the C-terminus appeared to abolish the base editing activity. This suggests that the maximum / optimal deletion from the C-terminus of Cas13e.1 is likely to be between 150 and 180 residues.

[0358] Based on this discovery, a series of N-terminal deletion mutants were generated for dCas13e.1 with a 150 C-terminal residue deletion. Seven such N-terminal deletion mutants with 30-, 60-, 90-, 120-, 150-, 180-, and 210-residue deletions were generated, respectively. See Figure 17. The results in Figure 18 showed that the best RNA editing activity was observed in mutants with a 180 N-terminal residue deletion and a 150 C-terminal residue deletion, i.e., a total 330-residue deletion from the 775-residue Cas13e.1 protein (generating the 445-residue dCas13e.1 optimal for generating the ADAR2DD fusion). showed that it was observed in mutants having

[0359] Example 7 Comparison of Mammalian Endogenous mRNA Knockdown Efficiency Using Various Cas13 Proteins This experiment demonstrated that Cas13e and Cas13f proteins, especially Cas13f.1, are very effective in knocking down mammalian endogenous target mRNAs much more than previously identified Cas13 proteins.

[0360] Specifically, five plasmids were constructed, each expressing one of the Cas13 proteins, namely, Cas13e.1 (SEQ ID NO: 22), Cas13f.1( Array No. 24 ), LwaCas13a( Array No. 51) ), PspCas13b( Array No. 45 ), and RxCas13d( Array No. 46 ). Also, each plasmid encoded an mCherry reporter gene and the sgRNA / crRNA coding sequence of each Cas13 protein flanked by two native DR sequences. These sgRNAs were designed to have spacer sequences targeting the ANXA4 mRNA. See SEQ ID NOs: 48 - 50. As negative controls, five additional plasmids were constructed, each encoding a non-targeting sgRNA / crRNA instead of the ANXA4-targeting sgRNA / crRNA (the "control NT constructs"). See Figure 19.

Chemical formula

[0361] Five Cas13 / sgRNA-encoding plasmids were transfected into HEK293 cells as in Example 4. After culturing for 24 hours, cells expressing mCherry were isolated by flow cytometry, and the expression of ANXA4 mRNA was determined using RT-PCR to evaluate the knockdown efficiency compared to control cells transfected with the Cas13 / NT-encoding plasmid.

[0362] Figure 20 showed that only Cas13b had marginal ANXA4 mRNA knockdown, while Cas13e.1, Cas13f.1, and Cas13d each had knockdown exceeding 80% of the target ANXA4 mRNA. Notably, Cas13e.1 appeared to have the most robust knockdown efficiency.

Claims

**Claim 1** A clustered regularly interspaced short palindromic repeat (CRISPR)-associated protein (Cas) comprising an amino acid sequence having at least about 90% sequence identity to the amino acid sequence of SEQ ID NO:

3. **Claim 2** The Cas according to claim 1, having the same sequence as the amino acid sequence of SEQ ID NO: 3 within the HEPN domain or the RXXXXH motif. **Claim 3** The Cas according to claim 1, having a mutation within the RNase catalytic site of the Cas and not having RNase catalytic activity. **Claim 4** The Cas according to claim 1, which does not exhibit substantial (or detectable) collateral RNase activity. **Claim 5** The Cas according to any one of claims 1 to 4, fused to a nuclear localization signal (NLS) sequence or a nuclear export signal (NES). **Claim 6** (1) A guide RNA, or a polynucleotide encoding the guide RNA, the guide RNA comprising: (i) a spacer sequence capable of hybridizing to a target RNA, and (ii) a direct repeat (DR) sequence on the 3' side of the spacer sequence; And (2) The Cas according to any one of claims 1 to 5, or a polynucleotide encoding the Cas, A CRISPR-Cas system comprising: Provided that when the Cas has the amino acid sequence of SEQ ID NO: 3, the spacer sequence is not 100% complementary to a naturally occurring bacteriophage nucleic acid. The CRISPR-Cas system. **Claim 7** The DR sequence has substantially the same secondary structure as the secondary structure of the DR sequence encoded by the complementary sequence of SEQ ID NO: 10; or the DR sequence is encoded by the complementary sequence of SEQ ID NO:

10. The CRISPR-Cas system according to claim 6. **Claim 8** The target RNA is eukaryotic RNA or is encoded by eukaryotic DNA. The CRISPR-Cas system according to claim 6 or 7. **Claim 9** The eukaryotic DNA is non-human mammalian DNA, non-human primate DNA, human DNA, plant DNA, insect DNA, avian DNA, reptilian DNA, rodent DNA, fish DNA, worm / nematode DNA, or yeast DNA. The CRISPR-Cas system according to claim 8. **Claim 10** The CRISPR-Cas system according to claim 6 or 7, wherein the spacer array is 15 to 60 nucleotides, 25 to 50 nucleotides, or about 30 nucleotides.

11. The CRISPR-Cas system according to claim 6 or 7, wherein the spacer array is 90 to 100% complementary to the target RNA.

12. A polynucleotide encoding Cas according to any one of claims 1 to 5, wherein the polynucleotide is not SEQ ID NO:

17.

13. A vector comprising the polynucleotide according to claim 12.

14. A cell or its progeny comprising Cas according to any one of claims 1 to 5, a CRISPR-Cas system according to any one of claims 6 to 11, a polynucleotide according to claim 12, or a vector according to claim 13.

15. The cell or its progeny according to claim 14, which is a eukaryotic cell or a prokaryotic cell.

16. An in vitro method for modifying a target RNA, the method comprising contacting the target RNA with a CRISPR-Cas system according to any one of claims 6 to 11, wherein the spacer array is complementary to at least 15 nucleotides of the target RNA; the Cas binds to the RNA guide sequence to form a complex; the complex binds to the target RNA; and binding of the complex to the target RNA causes the Cas to modify the target RNA. The method.

Citation Information

Patent Citations

  • Novel CAS13b orthologues crispr enzymes and systems

    WO2018170333A1

  • Nucleic acids encoding crispr-associated proteins and uses thereof

    WO2018172556A1

  • Systems, methods, and compositions for targeted nucleic acid editing

    WO2019084063A1

  • Novel crispr enzymes and systems

    WO2020028555A2