Novel CRISPR DNA Targeting Enzymes and Systems
Engineered CRISPR-Cas systems, particularly type V-I systems, address the limitations of existing CRISPR-Cas systems by offering novel effector proteins and RNA guides for advanced nucleic acid editing, enhancing genome and epigenome manipulation and therapeutic applications.
Patent Information
- Application Number
- JP2024000695
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-12-05
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2039-03-14
AI Technical Summary
Current CRISPR-Cas systems, particularly Class 2 systems like CRISPR-Cas9, lack additional programmable effectors that can modify nucleic acids beyond their current capabilities, limiting their applications in genome and epigenome engineering.
Development of engineered, non-naturally occurring CRISPR-Cas systems, specifically type V-I systems, which include novel CRISPR-Cas effector proteins and RNA guides, offering unique properties such as smaller size, versatile delivery strategies, and capabilities for DNA/RNA editing, insertion, excision, and mobilization.
The novel CRISPR-Cas systems enhance genome and epigenome manipulation techniques by providing additional programmable tools for targeted nucleic acid modifications, enabling a wide range of applications including cancer and infectious disease treatment.
Smart Images

Figure 0007706581000024 
Figure 0007706581000025 
Figure 0007706581000026
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the benefit of priority of U.S. Patent Application No. 62 / 642,919, filed Mar. 14, 2018; U.S. Patent Application No. 62 / 666,397, filed May 3, 2018; U.S. Patent Application No. 62 / 672,489, filed May 16, 2018; U.S. Patent Application No. 62 / 679,628, filed Jun. 1, 2018; U.S. Patent Application No. 62 / 703,857, filed Jul. 26, 2018; U.S. Patent Application No. 62 / 740,856, filed Oct. 3, 2018; U.S. Patent Application No. 62 / 746,528, filed Oct. 16, 2018; U.S. Patent Application No. 62 / 772,038, filed Nov. 27, 2018; and U.S. Patent Application No. 62 / 775,885, filed Dec. 5, 2018. The entire contents of each of the foregoing applications are hereby incorporated by reference in their entirety.
[0002] The present disclosure relates to systems, methods, and compositions for gene expression control involving sequence targeting and nucleic acid editing that use a vector system related to Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) and its components.
Background Art
[0003] In recent years, with the application of the progress of genome sequencing technology and analysis, important insights have been obtained regarding the genetic basis of biological activities in a wide variety of natural areas ranging from the biosynthetic pathways of prokaryotes to human pathology. To fully understand and evaluate the vast amount of information generated by gene sequencing technology, improvements in the scale, effectiveness, and ease of corresponding genome and epigenome manipulation techniques are required. Such novel genome and epigenome engineering technologies will accelerate the development of new applications in numerous areas, including biotechnology, agriculture, and human therapeutics.
[0004] Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) and CRISPR-associated (Cas) genes are collectively known as the CRISPR-Cas or CRISPR / Cas system and are currently understood to confer immunity against phage infection in bacteria and archaea. The CRISPR-Cas system of prokaryotic adaptive immunity is a highly diverse group of protein effectors, non-coding elements, and locus architectures, and several examples of which have been engineered and adapted to generate important biotechnologies.
[0005] The components of this system involved in host defense include one or more effector proteins with the ability to modify DNA or RNA, and an RNA guide element responsible for targeting these protein activities to specific sequences on phage DNA or RNA. The RNA guide is composed of CRISPR RNA (crRNA) and may require an additional trans-activating RNA (tracrRNA) to enable the manipulation of target nucleic acids by one or more effector proteins. The crRNA consists of a direct repeat responsible for protein binding to the crRNA and a spacer sequence complementary to the desired nucleic acid target sequence. The CRISPR system can be reprogrammed to target another DNA or RNA target by modifying the spacer sequence of the crRNA.
[0006] The CRISPR-Cas system can be broadly divided into two classes: Class 1 systems are composed of multiple effector proteins that together form a complex around the crRNA, and Class 2 systems consist of a single effector protein that complexes with an RNA guide to target a DNA or RNA substrate. The single-subunit effector composition of Class 2 systems provides a more straightforward set of components for engineering and translational applications, and thus has been an important source of programmable effectors to date. Therefore, the discovery, engineering, and optimization of novel Class 2 systems could lead to broad and powerful programmable technologies for genome engineering and beyond.
Summary of the Invention
Problems to be Solved by the Invention
[0007] The CRISPR-Cas system is an adaptive immune system that defends species from foreign genetic elements in archaea and bacteria. Class 2, exemplified by CRISPR-Cas9 Characterization and engineering of the CRISPR-Cas system have opened the way to a wide range of biotechnology applications in genome editing and beyond. Nevertheless, there remains a need for additional programmable effectors and systems that go beyond current CRISPR-Cas systems to realize new applications by virtue of their unique properties for modifying nucleic acids and polynucleotides (i.e., DNA, RNA, or any hybrid, derivative, or modification thereof).
[0008] The citation or identification of any document in this application does not admit that such document is available as prior art for the present invention.
Means for Solving the Problems
[0009] The present disclosure provides engineered systems and compositions of novel single-effector Class 2 CRISPR-Cas systems that do not occur in nature, along with methods for computational identification from genomic databases, development from native loci into engineered systems, and experimental validation and translational application. This novel effector has a sequence different from orthologs and homologs of existing Class 2 CRISPR effectors and also has a unique domain composition. This provides, but is not limited to, 1) novel DNA / RNA editing properties and control mechanisms, 2) a smaller size compared to higher versatility in delivery strategies, 3) cellular processes such as cell death induced by genotype, and 4) additional features including programmable RNA-guided DNA insertion, excision, and mobilization. The novel DNA targeting systems described herein add to the toolbox of genomic and epigenomic manipulation techniques, enabling a wide range of applications to specific and programmed perturbations.
[0010] Generally, the present disclosure relates to novel CRISPR-Cas systems that include newly discovered enzymes and other components used to create a minimal system that can be used in a non-native environment, such as bacteria other than the bacteria in which the system was originally discovered.
[0011] In one aspect, the present disclosure provides an engineered, non-naturally occurring CRISPR-Cas system comprising: i) one or more type V-I (CLUST.029130) RNA guides or one or more nucleic acids encoding one or more type V-I RNA guides, wherein the type V-I RNA guide comprises or consists of a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid; and ii) a type V-I (CLUST.029130) CRISPR-Cas effector protein or a nucleic acid encoding a type V-I CRISPR-Cas effector protein, wherein the type V-I CRISPR-Cas effector protein has the ability to bind to the type V-I RNA guide and target a target nucleic acid sequence complementary to the spacer sequence, wherein the target nucleic acid is DNA. As used herein, the type V-I (CLUST.029130) CRISPR-Cas effector protein is also referred to as a Cas12i effector protein, and these two terms are used interchangeably in the present disclosure.
[0012] In some embodiments of any part of the systems described herein, the type V-I CRISPR-Cas effector protein is less than or equal to about 1100 amino acids in length (excluding any amino acid signal sequence or peptide tag fused thereto) and comprises at least one RuvC domain. In some embodiments, none, one, or more of the RuvC domains are catalytically inactive. In some embodiments, the type V-I CRISPR-Cas effector protein comprises or consists of the amino acid sequence X1SHX4DX6X7 (SEQ ID NO: 200), wherein X1 is S or T, X4 is Q or L, X6 is P or S, and X7 is F or L.
[0013] In some embodiments, the type V-I CRISPR-Cas effector protein has the amino acid sequence X1XDXNX6X7XXXX 11 (SEQ ID NO: 201), wherein X1 is A or G or S, X is any amino acid, X6 is Q or I, X7 is T or S or V, and X 10comprises or consists of T or A). In some embodiments, the type V-I CRISPR-Cas effector protein comprises or consists of the amino acid sequence X1X2X3E (SEQ ID NO: 210), where X1 is C or F or I or L or M or P or V or W or Y, X2 is C or F or I or L or M or P or R or V or W or Y, and X3 is C or F or G or I or L or M or P or V or W or Y).
[0014] In some embodiments, the type V-I CRISPR-Cas effector protein comprises two or more sequences from the set of SEQ ID NO: 200, SEQ ID NO: 201, and SEQ ID NO: 210. In some embodiments, the type V-I CRISPR-Cas effector protein comprises or consists of an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequences provided in Table 4 (e.g., SEQ ID NOs: 1-5, and 11-18).
[0015] In some embodiments of any part of the systems described herein, the type V-I CRISPR-Cas effector protein comprises or consists of an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) identical to the amino acid sequence of Cas12i1 (SEQ ID NO: 3) or Cas12i2 (SEQ ID NO: 5). In some embodiments, the type V-I CRISPR-Cas effector protein is Cas12i1 (SEQ ID NO: 3) or Cas12i2 (SEQ ID NO: 5).
[0016] In some embodiments, the type V-I CRISPR-Cas effector protein has the ability to recognize a protospacer adjacent motif (PAM), and the target nucleic acid comprises or consists of a PAM that includes or consists of the nucleic acid sequence 5'-TTN-3' or 5'-TTH-3' or 5'-TTY-3' or 5'-TTC-3'.
[0017] In some embodiments of any part of the systems described herein, the type V-I CRISPR-Cas effector protein comprises one or more amino acid substitutions within the scope of at least one of the RuvC domains. In some embodiments, the one or more amino acid substitutions include substitutions at amino residues corresponding to D647 or E894 or D948 of SEQ ID NO: 3, such as alanine substitutions. In some embodiments, the one or more amino acid substitutions include alanine substitutions at amino residues corresponding to D599 or E833 or D886 of SEQ ID NO: 5. In some embodiments, the one or more amino acid substitutions result in a decrease in the nuclease activity of the type V-I CRISPR-Cas effector protein as compared to the nuclease activity of a type V-I CRISPR-Cas effector protein without the one or more amino acid substitutions.
[0018] In some embodiments of any part of the systems described herein, the type V-I RNA guide comprises a direct repeat sequence that includes a stem-loop structure proximal to the 3' end (adjacent directly to the spacer sequence). In some embodiments, the type V-I RNA guide direct repeat includes a stem-loop proximal to the 3' end, where the stem is 5 nucleotides in length. In some embodiments, the type V-I RNA guide direct repeat includes a stem-loop proximal to the 3' end, where the stem is 5 nucleotides in length and the loop is 7 nucleotides in length. In some embodiments, the type V-I RNA guide direct repeat includes a stem-loop proximal to the 3' end, where the stem is 5 nucleotides in length and the loop is 6, 7, or 8 nucleotides in length.
[0019] In some embodiments, the V-I type RNA guide directory repeat contains the sequence 5'-CCGUCNNNNNNUGACGG-3' (SEQ ID NO: 202) proximal to the 3' end, where N refers to any nucleobase. In some embodiments, the V-I type RNA guide directory repeat contains the sequence 5'-GUGCCNNNNNNUGGCAC-3' (SEQ ID NO: 203) proximal to the 3' end, where N refers to any nucleobase.
[0020] In some embodiments, the V-I type RNA guide directory repeat has the sequence 5'-GUGUCN 5-6 UGACAX1-3' (SEQ ID NO: 204), where N 5-6 refers to any consecutive sequence of 5 or 6 nucleobases, and X1 refers to C or T or U. In some embodiments, the V-I type RNA guide directory repeat contains the sequence 5'-UCX3UX5X6X7UUGACGG-3' (SEQ ID NO: 205), where X3 refers to C or T or U, X5 refers to A or T or U, X6 refers to A or C or G, and X7 refers to A or G. In some embodiments, the V-I type RNA guide directory repeat contains the sequence 5'-CCX3X4X5CX7UUGGCAC-3' (SEQ ID NO: 206), where X3 refers to C or T or U, X4 refers to A or T or U, X5 refers to C or T or U, and X7 refers to A or G.
[0021] In some embodiments, the V-I type RNA guide comprises or consists of a directory repeat sequence that is at least 80% identical, e.g., 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the nucleotide sequences provided in Table 5A (e.g., SEQ ID NOs: 6-19 and 19-24).
[0022] In some embodiments, the type V-I RNA guide comprises or consists of the nucleotide sequences provided in Table 5B or a subsequence thereof (e.g., SEQ ID NOs: 150-163). In some embodiments, the type V-I RNA guide comprises or consists of a nucleotide sequence constructed by concatenation of a directory repeat, a spacer, and a directory repeat sequence, where the directory repeat sequence is provided in Table 5A and the length of the spacer is provided in the Spacer Length 1 column of Table 5B. In some embodiments, the type V-I RNA guide comprises or consists of a nucleotide sequence constructed by concatenation of a directory repeat, a spacer, and a directory repeat sequence, where the directory repeat sequence is provided in Table 5A and the length of the spacer is provided in the Spacer Length 2 column of Table 5B. In some embodiments, the type V-I RNA guide comprises or consists of a nucleotide sequence constructed by concatenation of a directory repeat, a spacer, and a directory repeat sequence, where the directory repeat sequence is provided in Table 5A and the length of the spacer is provided in the Spacer Length 3 column of Table 5B.
[0023] In some embodiments of any part of the systems described herein, the spacer sequence of the type V-I RNA guide comprises or consists of from about 15 to about 34 nucleotides (e.g., 16, 17, 18, 19, 20, 21, or 22 nucleotides). In some embodiments of any part of the systems described herein, the spacer is from 17 nucleotides to 31 nucleotides in length.
[0024] In some embodiments of any part of the systems provided herein, the target nucleic acid is DNA. In some embodiments of any part of the systems described herein, the target nucleic acid comprises or consists of a protospacer adjacent motif (PAM), e.g., a PAM comprising the nucleic acid sequence 5'-TTN-3' or 5'-TTH-3' or 5'-TTY-3' or 5'-TTC-3'.
[0025] In certain embodiments of any of the systems provided herein, targeting of a target nucleic acid by a type V-I CRISPR-Cas effector protein and an RNA guide results in a modification (e.g., a single-stranded or double-stranded cleavage event) of the target nucleic acid. In some embodiments, the modification is a deletion event. In some embodiments, the modification is an insertion event. In some embodiments, the modification results in cytotoxicity and / or cell death.
[0026] In some embodiments, the type V-I CRISPR-Cas effector protein has non-specific (i.e., “collateral”) nuclease (e.g., DNase) activity. In certain embodiments of any of the systems provided herein, the system further comprises a donor template nucleic acid (e.g., DNA or RNA).
[0027] In some embodiments of any of the systems provided herein, the system is within a cell (e.g., a eukaryotic cell (e.g., a mammalian cell) or a prokaryotic cell (e.g., a bacterial cell)).
[0028] In another aspect, the present disclosure provides methods for targeting and editing a target nucleic acid, where the method comprises contacting the target nucleic acid with any of the systems described herein. This can be done ex vivo or in vitro. In some embodiments, the methods described herein do not modify the identity of human germline genes.
[0029] In other aspects, the present disclosure provides methods for targeting the insertion of a payload nucleic acid at a site in a target nucleic acid, where the method comprises contacting the target nucleic acid with any of the systems described herein.
[0030] In yet another aspect, the present disclosure provides methods for targeting the excision of a payload nucleic acid from a site in a target nucleic acid, where the method comprises contacting the target nucleic acid with any of the systems described herein.
[0031] In another aspect, the present disclosure provides a method for targeting and nicking a non-target strand (non-spacer complementary strand) of double-stranded target DNA in response to recognition of the target strand (spacer complementary strand) of double-stranded target DNA. This method includes contacting any of the systems described herein with double-stranded target DNA.
[0032] In yet another aspect, the present disclosure provides a method for targeting and cleaving double-stranded target DNA, the method including contacting any of the systems described herein with double-stranded target DNA.
[0033] In some embodiments of the method for targeting and cleaving double-stranded target DNA, the non-target strand (non-spacer complementary strand) of double-stranded target DNA is nicked before nicking of the target strand (spacer complementary strand) of the double-stranded target nucleic acid.
[0034] In yet another aspect, the present disclosure provides a method for specific editing of double-stranded nucleic acids, the method including: (a) a type V-I effector protein and one other enzyme having sequence-specific nicking activity; (b) a type V-I RNA guide that induces the type V-I effector protein to nick the opposite strand compared to the activity of the other sequence-specific nickase; and (c) contacting the double-stranded nucleic acid, wherein this method results in a reduced likelihood of off-target modification.
[0035] In some embodiments, the type V-I effector protein further includes a linker sequence. In some embodiments, the type V-I effector protein includes one or more mutations or amino acid substitutions that result in a CRISPR-related protein that is unable to cleave DNA.
[0036] In yet another aspect, the present disclosure provides a method for base editing of double-stranded nucleic acids, the method comprising: (a) a fusion protein comprising a type V-I effector protein and a protein domain having DNA modification activity (e.g., cytidine deamination); (b) a type V-I RNA guide that targets the double-stranded nucleic acid; and (c) contacting the double-stranded nucleic acid. The type V-I effector of the fusion protein may be modified to nick the non-target strand of the double-stranded nucleic acid. In some embodiments, the type V-I effector of the fusion protein may be modified to be nuclease-deficient. zzz
[0037] In another aspect, the present disclosure provides a method for modifying a DNA molecule, the method comprising contacting the system described herein with the DNA molecule.
[0038] In some embodiments of any of the methods described herein (and compositions for use in such methods), the cell is a eukaryotic cell. In some embodiments, the cell is an animal cell. In some embodiments, the cell is a cancer cell (e.g., a tumor cell). In some embodiments, the cell is an infectious pathogen cell or a cell infected with an infectious pathogen. In some embodiments, the cell is a bacterial cell, a cell infected with a virus, a cell infected with a prion, a fungal cell, a protozoan, or a parasitic cell.
[0039] In another aspect, the present disclosure provides a method for treating a condition or disease in a subject in need thereof and a composition for use in such method. The method comprises administering to the subject the system described herein, wherein the spacer sequence is complementary to at least 15 nucleotides of a target nucleic acid associated with the condition or disease, wherein the type V-I CRISPR-Cas effector protein associates with the RNA guide to form a complex, wherein the complex binds to a target nucleic acid sequence complementary to at least 15 nucleotides of the spacer sequence, and wherein when the complex binds to the target nucleic acid sequence, the type V-I CRISPR-Cas effector protein cleaves or silences the target nucleic acid, thereby treating the condition or disease of the subject.
[0040] In some embodiments of the methods described herein (and compositions for use in such methods), the condition or disease is cancer or an infectious disease. In some embodiments, the condition or disease is cancer, where the cancer is selected from the group consisting of Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myelogenous leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, and bladder cancer.
[0041] In some embodiments, the type V-I effector protein comprises or consists of at least one (e.g., 2, 3, 4, 5, 6, or more) nuclear localization signal (NLS). In some embodiments, the type V-I effector protein comprises or consists of at least one (e.g., 2, 3, 4, 5, 6, or more) nuclear export signal (NES). In some embodiments, the type V-I effector protein comprises at least one (e.g., 2, 3, 4, 5, 6, or more) NLS and at least one (e.g., 2, 3, 4, 5, 6, or more) NES.
[0042] In some embodiments, the systems described herein comprise a nucleic acid encoding one or more RNA guides. In some embodiments, the nucleic acid encoding one or more RNA guides is operably linked to a promoter (e.g., a constitutive promoter or an inducible promoter).
[0043] In some embodiments, the systems described herein comprise a nucleic acid encoding a target nucleic acid (e.g., a target DNA). In some embodiments, the nucleic acid encoding the target nucleic acid is operably linked to a promoter (e.g., a constitutive promoter or an inducible promoter).
[0044] In some embodiments, the system described herein comprises a nucleic acid encoding a V-I type CRISPR-Cas effector protein in a vector. In some embodiments, the system further comprises one or more nucleic acids encoding an RNA guide present in the vector.
[0045] In some embodiments, the vector included in the system is a viral vector (e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector). In some embodiments, the vector included in the system is a phage vector.
[0046] In some embodiments, the system provided herein is in a delivery system. In some embodiments, the delivery system is nanoparticles, liposomes, exosomes, microvesicles, and gene guns.
[0047] The present disclosure also provides a cell (e.g., a eukaryotic cell or a prokaryotic cell (e.g., a bacterial cell)) comprising the system described herein. In some embodiments, the eukaryotic cell is a mammalian cell (e.g., a human cell) or a plant cell. The present disclosure also provides an animal model (e.g., a rodent, rabbit, dog, monkey, or ape model) and a plant model comprising the cell. In some embodiments, the method is used for the treatment of a subject, such as a mammal, e.g., a human patient. The mammalian subject may also be a domesticated mammal such as a dog, cat, horse, monkey, rabbit, rat, mouse, female cow, goat, or sheep.
[0048] In yet another aspect, the present disclosure provides a method for detecting a target nucleic acid (e.g., DNA or RNA) in a sample, the method comprising: (a) contacting the system provided herein and a labeled reporter nucleic acid with the sample, wherein cleavage of the labeled reporter nucleic acid occurs when the crRNA hybridizes to the target nucleic acid; and (b) measuring a detectable signal generated by cleavage of the labeled reporter nucleic acid, thereby detecting the presence of the target nucleic acid in the sample.
[0049] In some embodiments, the method for detecting a target nucleic acid can also include comparing the level of the detectable signal to a reference signal level and determining the amount of the target nucleic acid in the sample based on the level of the detectable signal.
[0050] In some embodiments, the measurement is performed using gold nanoparticle detection, fluorescence polarization, colloid phase transition / dispersion, electrochemical detection, or semiconductor-based sensing.
[0051] In some embodiments, the labeled reporter nucleic acid can include a pair of fluorescent emission dyes, a fluorescence resonance energy transfer (FRET) pair, or a quencher / fluorophore pair, wherein when the labeled reporter nucleic acid is cleaved by an effector protein, an increase or decrease in the amount of the signal generated by the labeled reporter nucleic acid occurs.
[0052] In another aspect, the present disclosure includes a method of modifying a target DNA, the method comprising contacting a complex comprising a Cas12i effector protein and an engineered type V-I RNA guide designed to hybridize to a target sequence of the target DNA (e.g., at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% complementary thereto) with the target DNA, and the system is distinguished by (a) the absence of tracrRNA in the system, and (b) the Cas12i effector protein and the type V-I RNA guide forming a complex that associates with the target DNA, thereby modifying the target DNA.
[0053] In certain embodiments, modifying the target DNA comprises cleaving at least one strand of the target DNA (e.g., creating a single-strand break or "nick", or creating a double-strand break). Alternatively, or in addition, modifying the target DNA comprises either (i) binding to the target DNA, thereby preventing the target DNA from associating with another biomolecule or complex, or (ii) unwinding a portion of the target DNA. In some examples, the target DNA comprises a protospacer adjacent motif (PAM) sequence recognized by the Cas12i effector protein, such as 5'-TTN-3' or 5'-TTH-3' or 5'-TTY-3' or 5'-TTC-3'. The Cas12 effector protein is, in certain embodiments, a Cas12i1 effector protein or a Cas12i2 effector protein.
[0054] Continuing with this aspect of the present disclosure, in certain embodiments, contacting the complex with the target DNA comprises, for example, (a) contacting the complex with a cell, where the complex is formed in vitro, or (b) contacting the cell with one or more nucleic acids encoding the Cas12i effector protein and the type V-I RNA guide, followed by their expression by the cell and formation of the complex intracellularly. In some cases, the cell is a prokaryotic cell; in other cases, it is a eukaryotic cell.
[0055] In another aspect, the present disclosure relates to a method of modifying a target DNA, the method comprising contacting a genomic editing system comprising a Cas12i protein and a type V-I RNA guide (e.g., crRNA, guide RNA or similar structure, optionally comprising one or more nucleotide, nucleobase or backbone modifications) comprising a 15-24 nucleotide spacer sequence having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% complementarity to a sequence in the target DNA, but not comprising a tracrRNA, to the target DNA within a cell. In various embodiments, the Cas12i protein comprises or consists of an amino acid sequence having at least 95%, such as 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 3, and the type V-I RNA guide comprises a direct repeat sequence having at least 95%, such as 96%, 97%, 98%, 99%, or 100% sequence identity to one of SEQ ID NO: 7 or 24; or the Cas12i protein comprises or consists of an amino acid sequence having at least 95%, such as 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 5, and the type V-I RNA guide comprises a direct repeat sequence having at least 95%, such as 96%, 97%, 98%, 99%, or 100% sequence identity to one of SEQ ID NO: 9 or 10. The target DNA is optionally cellular DNA, and the contacting is optionally performed within a cell such as a prokaryotic cell or a eukaryotic cell (e.g., a mammalian cell, a plant cell, or a human cell).
[0056] In some embodiments, the type V-I CRISPR-Cas effector protein comprises an amino acid sequence having at least 90%, or at least 95% sequence identity with one of SEQ ID NOs: 1-5 or 11-18. According to certain embodiments, the type V-I CRISPR-Cas effector protein comprises the amino acid sequence provided by SEQ ID NO: 3, or the amino acid sequence provided by SEQ ID NO: 5. The full length of the CRISPR-Cas effector protein according to certain embodiments is less than 1100 amino acids excluding any amino acid signal sequence or peptide tag fused thereto. In some cases, the CRISPR-Cas effector protein comprises an amino acid substitution, such as a substitution at the amino acid residue corresponding to D647, E894, or D948 of SEQ ID NO: 3, or a substitution at the amino acid residue corresponding to D599, E833, or D886 of SEQ ID NO: 5. The substitution is optionally alanine.
[0057] In yet another aspect, the present disclosure relates to an engineered, non-naturally occurring CRISPR-Cas system comprising or consisting of a Cas12i effector protein and an engineered type V-I RNA guide having a 15-34 nucleotide spacer sequence that is at least 80%, such as 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% complementary to a target sequence (e.g., crRNA, guide RNA, or similar construct, optionally including one or more nucleotide, nucleobase, or backbone modifications). The system does not include a tracrRNA, and the Cas12i effector protein and the type V-I RNA guide form a complex that associates with the target sequence. In some examples, the complex of the Cas12i effector protein and the type V-I RNA guide results in cleavage of at least one strand of the DNA containing the target sequence. The target sequence can include a protospacer adjacent motif (PAM) sequence recognized by the Cas12i effector protein, which PAM sequence is optionally 5'-TTN-3', 5'-TTY-3', or 5'-TTH-3' or 5'-TTC-3'. The type V-I RNA guide can include a direct repeat sequence having at least 95%, such as 96%, 97%, 98%, 99%, or 100% sequence identity to one of SEQ ID NOs: 7, 9, 10, 24, 100, or 101.
[0058] In certain embodiments, the Cas12i effector protein comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 3, and the direct repeat sequence has at least 95% sequence identity with SEQ ID NO: 100, or the Cas12i effector protein comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 5, and the direct repeat sequence has at least 95% sequence identity with SEQ ID NO: 101. Alternatively, or in addition, the Cas12i effector protein comprises an amino acid substitution (optionally, an alanine substitution) selected from the group consisting of (a) a substitution at an amino acid residue corresponding to D647, E894, or D948 of SEQ ID NO: 3; and (b) a substitution at an amino acid residue corresponding to D599, E833, or D886 of SEQ ID NO: 5.
[0059] In yet another aspect, the disclosure relates to a composition comprising one or more nucleic acids encoding a CRISPR-Cas system (or genome editing system) according to one of the aspects of the disclosure. And in another aspect, the disclosure relates to a viral vector encoding a CRISPR-Cas system (or genome editing system) according to one of the aspects of the disclosure.
[0060] The disclosure also includes a method of targeting and nicking the non-spacer complementary strand of a double-stranded target DNA in response to recognition of the spacer complementary strand of the double-stranded target DNA, the method comprising contacting the double-stranded target DNA with any of the systems described herein.
[0061] In another aspect, the disclosure includes a method of targeting and cleaving a double-stranded target DNA, the method comprising contacting the double-stranded target DNA with a system as described herein. In these methods, the non-spacer complementary strand of the double-stranded target DNA is nicked prior to nicking of the spacer complementary strand of the double-stranded target nucleic acid.
[0062] In other embodiments, the present disclosure includes a method for detecting a target nucleic acid in a sample, the method comprising: (a) contacting a system as described herein and a labeled reporter nucleic acid with the sample, wherein cleavage of the labeled reporter nucleic acid occurs when the crRNA hybridizes to the target nucleic acid; and (b) measuring a detectable signal generated by cleavage of the labeled reporter nucleic acid, thereby detecting the presence of the target nucleic acid in the sample. These methods may further comprise comparing the level of the detectable signal to a reference signal level and determining the amount of the target nucleic acid in the sample based on the level of the detectable signal. In some embodiments, the measurement is performed using gold nanoparticle detection, fluorescence polarization, colloid phase transition / dispersion, electrochemical detection, or semiconductor-based sensing. In some embodiments, the labeled reporter nucleic acid comprises a pair of fluorescent emission dyes, a fluorescence resonance energy transfer (FRET) pair, or a quencher / fluorophore pair, wherein cleavage of the labeled reporter nucleic acid by an effector protein results in an increase or decrease in the amount of signal generated by the labeled reporter nucleic acid.
[0063] In another aspect, the methods herein include specifically editing a double-stranded nucleic acid, the method comprising: (a) a type V-I CRISPR-Cas effector and one other enzyme having sequence-specific nicking activity, and a crRNA that induces the type V-I CRISPR-Cas effector to nick the opposite strand compared to the activity of the other sequence-specific nickase; and (b) contacting the double-stranded nucleic acid under sufficient conditions for a sufficient length of time, wherein formation of a double-strand break occurs by this method.
[0064] Another aspect includes a method for editing a double-stranded nucleic acid, the method comprising: (a) a fusion protein comprising a type V-I CRISPR-Cas effector, a protein domain having DNA modification activity, and an RNA guide that targets the double-stranded nucleic acid; and (b) contacting the double-stranded nucleic acid under sufficient conditions for a sufficient length of time; wherein the type V-I CRISPR-Cas effector of the fusion protein is modified to nick the non-target strand of the double-stranded nucleic acid.
[0065] Another aspect includes a method for inducing genotype-specific or transcription state-specific cell death or dormancy in a cell, the method comprising contacting a cell, such as a prokaryotic or eukaryotic cell, with any of the systems disclosed herein, wherein when the RNA guide hybridizes to the target DNA, collateral DNase activity-mediated cell death or dormancy occurs. For example, the cell may be a mammalian cell, such as a cancer cell. The cell may be an infectious cell or a cell infected with an infectious pathogen, such as a cell infected with a virus, a prion, a fungal cell, a protozoan, or a parasite cell.
[0066] In another aspect, the present disclosure provides a method for treating a condition or disease in a subject in need thereof, the method comprising administering any of the systems described herein to the subject, wherein the spacer sequence is complementary to at least 15 nucleotides of a target nucleic acid associated with the condition or disease; wherein the type V-I CRISPR-Cas effector protein associates with the RNA guide to form a complex; Here, the complex binds to a target nucleic acid sequence complementary to at least 15 nucleotides of the spacer array; and here, in response to the complex binding to the target nucleic acid sequence, a type V-I CRISPR-Cas effector protein cleaves the target nucleic acid, thereby treating a condition or disease of interest. For example, the condition or disease may be cancer or an infectious disease. For example, the condition or disease may be cancer, where the cancer is selected from the group consisting of Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, bile duct cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myelogenous leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, and bladder cancer.
[0067] The present disclosure also includes a system or cell as described herein for use as a medicament or for use in the treatment or prevention of cancer or an infectious disease, for example, where the cancer is selected from the group consisting of Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, bile duct cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myelogenous leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, and bladder cancer.
[0068] The present disclosure also includes in vitro or ex vivo a) methods for targeting and editing a target nucleic acid; b) methods for non-specific degradation of single-stranded DNA in response to recognition of a DNA target nucleic acid; c) methods for targeting and nicking the non-spacer complementary strand of a double-stranded target DNA in response to recognition of the spacer complementary strand of the double-stranded target DNA; d) methods for targeting and cleaving a double-stranded target DNA; e) methods for detecting a target nucleic acid in a sample; f) Method for specifically editing double-stranded nucleic acid; g) Method for base editing of double-stranded nucleic acid; h) Method for inducing genotype-specific or transcription state-specific cell death or dormancy in cells, i) Method for creating indels in double-stranded target DNA; j) Method for inserting a sequence into double-stranded target DNA, or k) Method for deleting or forming an inversion of a sequence in double-stranded target DNA also provides the use of the system or cell as described in this specification in the above.
[0069] In another aspect, the present disclosure provides a) Method for targeting and editing a target nucleic acid; b) Method for non-specific degradation of single-stranded DNA in response to recognition of a DNA target nucleic acid; c) Method for targeting and nicking the non-spacer complementary strand of double-stranded target DNA in response to recognition of the spacer complementary strand of double-stranded target DNA; d) Method for targeting and cleaving double-stranded target DNA; e) Method for detecting a target nucleic acid in a sample; f) Method for specifically editing double-stranded nucleic acid; g) Method for base editing of double-stranded nucleic acid; h) Method for inducing genotype-specific or transcription state-specific cell death or dormancy in cells; i) Method for creating indels in double-stranded target DNA; j) Method for inserting a sequence into double-stranded target DNA, or k) Method for deleting or forming an inversion of a sequence in double-stranded target DNA provides the use of the system or cell described in this specification in the above, wherein this method does not include the process of modifying the identity of human germline genes and does not include a method for treating the human or animal body.
[0070] In the methods described herein, cleavage of the target DNA or target nucleic acid results in the formation of indels, or cleavage of the target DNA or target nucleic acid results in the insertion of a nucleic acid sequence, or cleavage of the target DNA or target nucleic acid involves cleavage of the target DNA or target nucleic acid at two sites, resulting in a deletion or inversion of the sequence between those two sites.
[0071] In some embodiments, the various systems described herein may lack a tracrRNA. In some embodiments, a type V-I CRISPR-Cas effector protein and a type V-I RNA guide form a complex that modifies a target nucleic acid by associating with the target nucleic acid.
[0072] In some embodiments of the systems described herein, the spacer sequence is 15 to 47 nucleotides in length, for example, 20 to 40 nucleotides in length, or 24 to 38 nucleotides in length.
[0073] In another aspect, the disclosure provides a eukaryotic cell, such as a mammalian cell, such as a human cell, comprising a modified target locus, wherein the target locus has been modified by the method of any one of the preceding claims or by use of a composition thereof. For example, modification of the target locus may (i) a eukaryotic cell comprising a change in the expression of at least one gene product; (ii) a eukaryotic cell comprising a change in the expression of at least one gene product, wherein the expression of at least one gene product is increased; (iii) a eukaryotic cell comprising a change in the expression of at least one gene product, wherein the expression of at least one gene product is decreased; or (iv) a eukaryotic cell comprising an edited genome can be generated.
[0074] In another aspect, the disclosure provides a eukaryotic cell described herein, or a eukaryotic cell line comprising the same, or progeny thereof, or a multicellular organism comprising one or more of the eukaryotic cells described herein.
[0075] The present disclosure also provides a plant or animal model comprising one or more cells as described herein.
[0076] In another aspect, the present disclosure provides a method for producing a plant in which a trait of interest encoded by a gene of interest is modified, the method comprising contacting a plant cell with any of the systems described herein, thereby effecting either modification or introduction of the gene of interest, and regenerating a plant from the plant cell.
[0077] The present disclosure also provides a method for identifying a trait of interest in a plant, where the trait of interest is encoded by a gene of interest, the method comprising contacting a plant cell with any of the systems described herein, thereby identifying the gene of interest. For example, the method may further comprise introducing the identified gene of interest into a plant cell or a plant cell line or a plant germplasm and generating a plant therefrom, whereby the plant contains the gene of interest. The method may comprise causing the plant to exhibit the trait of interest.
[0078] The present disclosure also includes a method for targeting and cleaving single-stranded target DNA, the method comprising contacting a target nucleic acid with any of the systems described herein. The method may comprise the disease or disorder being infectious, where the infectious agent is selected from the group consisting of human immunodeficiency virus (HIV), herpes simplex virus type 1 (HSV1), and herpes simplex virus type 2 (HSV2).
[0079] In some of the methods described herein, both strands of the target DNA may be cleaved at different sites, resulting in a sticky-end type of cleavage site. In other embodiments, both strands of the target DNA may be cleaved at the same site, resulting in a blunt-end type of double-stranded break (DSB).
[0080] In some of the therapies described herein, the condition or disease is selected from the group consisting of cystic fibrosis, Duchenne muscular dystrophy, Becker muscular dystrophy, α1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, hypercholesterolemia, Leber congenital amaurosis, sickle cell disease, and β-thalassemia.
[0081] As used herein, the term "cleavage event" refers to DNA cleavage in a target nucleic acid created by a nuclease of a CRISPR system described herein. In some embodiments, the cleavage event is a double-stranded DNA cleavage. In some embodiments, the cleavage event is a single-stranded DNA cleavage.
[0082] As used herein, the terms "CRISPR-Cas system", "type V-I CRISPR-Cas system", or "type V-I system" refer to a type V-I CRISPR-Cas effector protein (i.e., a Cas12i effector protein) and one or more type V-I RNA guides, and / or a nucleic acid encoding a type V-I CRISPR-Cas effector protein or one or more type V-I RNA guides, and optionally, a promoter operably linked to the expression of the CRISPR effector or the RNA guide or both.
[0083] As used herein, the term "CRISPR array" refers to a nucleic acid (e.g., DNA) segment that includes CRISPR repeats and spacers, starting from the first nucleotide of the first CRISPR repeat and ending at the last nucleotide of the last (terminal) CRISPR repeat. Typically, each spacer in a CRISPR array is located between two repeats. The term "CRISPR repeat", or "CRISPR direct repeat", or "direct repeat", as used herein, refers to a plurality of short, unidirectionally repeating sequences that show very little or no sequence variation within the CRISPR array. Preferably, type V-I direct repeats can form a stem-loop structure.
[0084] "Stem-loop structure" refers to a nucleic acid having a secondary structure that includes a nucleotide region (stem portion) that is known or predicted to form a double-stranded structure that is linked on one side mainly by a single-stranded nucleotide region (loop portion). The terms "hairpin" and "foldback" structures are also used herein to refer to stem-loop structures. Such structures are well known in the art, and these terms are used in accordance with their known meaning in the art. As is known in the art, a stem-loop structure does not require exact base pairing. Thus, the stem may contain one or more base mismatches. Alternatively, the base pairing may be exact, i.e., it may contain no mismatches. The predicted stem-loop structures of some type V-I direct repeats are shown in FIG. 3. The stems of type V-I direct repeats contained in an RNA guide are composed of complementary 5-nucleotide bases that hybridize to each other, and the loops are 6, 7, or 9 nucleotides in length.
[0085] As used herein, the term "CRISPR RNA" or "crRNA" refers to an RNA molecule that contains a guide sequence used for targeting a specific nucleic acid sequence by a CRISPR effector. Typically, a crRNA includes a spacer sequence that mediates target recognition and a direct repeat sequence (referred to herein as the direct repeat or "DR" sequence) that forms a complex with a CRISPR-Cas effector protein.
[0086] As used herein, the term "donor template nucleic acid" refers to a nucleic acid molecule that can be used by one or more cellular proteins to modify the structure of a target nucleic acid after the CRISPR enzyme described herein has modified the target nucleic acid. In some embodiments, the donor template nucleic acid is a double-stranded nucleic acid. In some embodiments, the donor template nucleic acid is a single-stranded nucleic acid. In some embodiments, the donor template nucleic acid is linear. In some embodiments, the donor template nucleic acid is circular (e.g., a plasmid). In some embodiments, the donor template nucleic acid is an exogenous nucleic acid molecule. In some embodiments, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome).
[0087] As used herein, the terms "CRISPR-Cas effector", "CRISPR effector", "effector", "CRISPR-associated protein", or "CRISPR enzyme", "type V-I CRISPR-Cas effector protein", "type V-I CRISPR-Cas effector", "type V-I effector", or "Cas12i effector protein" refer to a protein that binds to a target site on a nucleic acid specified by a protein or RNA guide that performs enzymatic activity. The related CRISPR-Cas type V-I effector protein within the type V-I CRISPR-Cas system can also be referred to herein as "Cas12i" or "Cas12i enzyme". The Cas12i enzyme can recognize a related short motif near the target DNA called the protospacer adjacent motif (PAM). Preferably, the Cas12i enzyme of the present disclosure can recognize a PAM that includes or consists of TTN (wherein N means any nucleotide). For example, the PAM may be TTN, TTH, TTY, or TTC.
[0088] In some embodiments, the type V-I CRISPR-Cas effector protein has endonuclease activity, nickase activity, and / or exonuclease activity.
[0089] As used herein, the terms "CRISPR effector complex", "effector complex", "binary complex", or "surveillance complex" refer to a complex comprising a type V-I CRISPR-Cas effector protein and a type V-I RNA guide.
[0090] As used herein, the term "RNA guide" refers to any RNA molecule that facilitates targeting of the proteins described herein to a target nucleic acid. Exemplary "RNA guides" include, but are not limited to, crRNA, pre-crRNA (e.g., DR-spacer-DR), and mature crRNA (e.g., mature_DR-spacer, mature DR-spacer-mature_DR).
[0091] As used herein, the term "targeting" refers to the ability of a complex comprising a CRISPR-associated protein and an RNA guide, such as a crRNA, to preferentially or specifically bind, e.g., hybridize, to a particular target nucleic acid as compared to other nucleic acids that do not have the same or a similar sequence as the target nucleic acid.
[0092] As used herein, the term "target nucleic acid" refers to a particular nucleic acid substrate that contains a nucleic acid sequence complementary to all or part of the spacer in the RNA guide. In some embodiments, the target nucleic acid comprises a gene or a sequence within a gene. In some embodiments, the target nucleic acid comprises a non-coding region (e.g., a promoter). In some embodiments, the target nucleic acid is single-stranded. In some embodiments, the target nucleic acid is double-stranded.
[0093] The terms "activated CRISPR complex", "activated complex", or "ternary complex", as used herein, refer to a CRISPR effector complex after binding to or modifying a target nucleic acid.
[0094] The term "collateral RNA" or "collateral DNA" as used herein refers to a nucleic acid substrate that is non-specifically cleaved by an activated CRISPR complex.
[0095] The term "collateral DNase activity", as used herein in reference to a CRISPR enzyme, refers to the non-specific DNase activity of an activated CRISPR complex.
[0096] Unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. In the practice or testing of the present invention, methods and materials similar or equivalent to those described herein may be used, but suitable methods and materials are described below. Publications, patent applications, patents, and other references mentioned herein are hereby incorporated by reference in their entirety. In case of conflict, this specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.
[0097] Other features and advantages of the present invention will be apparent from the following detailed description and from the claims. In certain embodiments, for example, the following are provided: (Item 1) An RNA guide of type V-I (CLUST.029130) or a nucleic acid encoding said type V-I RNA guide, the RNA guide comprising a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid; and A type V-I (CLUST.029130) CRISPR-Cas effector protein or a nucleic acid encoding said effector protein, the effector protein having the ability to bind to said RNA guide and having the ability to target a target nucleic acid sequence complementary to said spacer sequence comprising, wherein the target nucleic acid is DNA, an engineered non-naturally occurring clustered regularly interspaced short palindromic repeat (CRISPR)-associated (Cas) system. (Item 2) The system according to item 1, comprising two or more RNA guides. (Item 3) The system according to item 1 or 2, wherein the type V-I CRISPR-Cas effector protein is less than about 1100 amino acids in length and comprises at least one RuvC domain. (Item 4) The system according to item 1 or 2, wherein the type V-I RNA guide comprises a directory repeat sequence, the spacer sequence, and a second directory repeat, which are arranged in order within the type V-I RNA guide. (Item 5) The type V-I CRISPR-Cas effector protein is an RuvC domain comprising the amino acid sequence X1SHX4DX6X7 (SEQ ID NO: 200), wherein X1 is S or T, X4 is Q or L, X6 is P or S, and X7 is F or L; an amino acid sequence X1XDXNX6X7XXXX 11 (SEQ ID NO: 201), wherein X1 is A, G, or S, X is any amino acid, X6 is Q or I, X7 is T, S, or V, and X 10 is T or A; and an RuvC domain comprising the amino acid sequence X1X2X3E (SEQ ID NO: 210), wherein X1 is C, F, I, L, M, P, V, W, or Y, X2 is C, F, I, L, M, P, R, V, W, or Y, and X3 is C, F, G, I, L, M, P, V, W, or Y The system according to any one of items 1 to 4, comprising one or more of the above. (Item 6) The system according to any one of items 1 to 4, wherein the type V-I CRISPR-Cas effector protein comprises an amino acid sequence that is at least 80% (e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identical to the amino acid sequence in Table 4. (Item 7) The system according to item 6, wherein the type V-I CRISPR-Cas effector protein comprises an amino acid sequence having at least 80% (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identity to the amino acid sequence of SEQ ID NO: 3 or SEQ ID NO: 5. (Item 8) The V-I type RNA guide includes a directory repeat sequence containing a stem-loop structure proximal to the 3' end of the directory repeat sequence, and the stem-loop structure is a first stem nucleotide strand 5 nucleotides in length; a second stem nucleotide strand 5 nucleotides in length (the first and second stem nucleotide strands can hybridize to each other); and a loop nucleotide strand disposed between the first and second stem nucleotide strands, the loop nucleotide strand containing 6, 7, or 8 nucleotides The system according to any one of items 1 to 7, comprising. (Item 9) The directory repeat sequence is 5'-CCGUCNNNNNNUGACGG-3' (SEQ ID NO: 202) proximal to the 3' end (where N is any nucleobase); 5'-GUGCCNNNNNNUGGCAC-3' (SEQ ID NO: 203) proximal to the 3' end (where N is any nucleobase); 5'-GUGUCN 5-6 UGACAX1-3' (SEQ ID NO: 204) (where N 5-6 is any continuous sequence of 5 or 6 nucleobases, and X1 is C or T or U); 5'-UCX3UX5X6X7UUGACGG-3' (SEQ ID NO: 205) proximal to the 3' end (where X3 is C, T, or U, X5 is A, T, or U, X6 is A, C, or G, and X7 is A or G); 5'-CCX3X4X5CX7UUGGCAC-3' (SEQ ID NO: 206) proximal to the 3' end (where X3 is C, T, or U, X4 is A, T, or U, X5 is C, T, or U, and X7 is A or G) The system according to any one of items 1 to 8, comprising any one of the above. (Item 10) The system according to any one of items 1 to 9, wherein the directory repeat sequence comprises a nucleotide sequence that is at least 80% (e.g., 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identical to the nucleotide sequence provided in Table 5. (Item 11) The system according to any one of items 1 to 10, wherein the RNA guide comprises a directory repeat sequence, the spacer sequence, and a second directory repeat arranged in sequence, the RNA guide sequence comprises a nucleotide sequence that is at least 80% (e.g., 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) identical to the nucleotide sequence or a fragment thereof provided in Table 5B, the variable N region means the spacer, and may have a length as specified in Table 5B. (Item 12) The system according to any one of items 1 to 11, wherein the type V-I CRISPR-Cas effector protein has the ability to recognize a protospacer adjacent motif (PAM), and the target nucleic acid comprises a PAM containing the nucleic acid sequence 5'-TTN-3' (wherein N is any nucleotide), 5'-TTH-3', 5'-TTY-3', or 5'-TTC-3'. (Item 13) The system according to any one of items 1 to 12, wherein the target nucleic acid is DNA. (Item 14) The system according to any one of items 1 to 13, wherein targeting of the target nucleic acid by the type V-I CRISPR-Cas effector protein and the RNA guide results in modification of the target nucleic acid. (Item 15) The system according to item 14, wherein the modification of the target nucleic acid is a cleavage event. (Item 16) The system according to item 14, wherein the modification of the target nucleic acid is a nicking event. (Item 17) The system according to any one of items 13 to 16, wherein the modification causes cytotoxicity. (Item 18) The system according to any one of items 3 to 17, wherein the type V-I CRISPR-Cas effector protein contains one or more amino acid substitutions within the RuvC domain that result in a decrease in the nuclease or nickase activity of the type V-I CRISPR-Cas effector protein as compared to the nuclease or nickase activity of the type V-I CRISPR-Cas effector protein having no amino acid substitutions. (Item 19) The system according to item 18, wherein the one or more amino acid substitutions include alanine substitutions at amino acid residues corresponding to D647, E894, or D948 of SEQ ID NO: 3; or D599, E833, or D886 of SEQ ID NO: 5. (Item 20) The system according to item 18 or 19, wherein the type V-I CRISPR-Cas effector protein is fused to a base editing domain. (Item 21) The system according to item 18 or 19, wherein the type V-I CRISPR-Cas effector protein is fused to a DNA methylation domain, a histone residue modification domain, a localization factor, a transcription modification factor, a light gate control factor, a chemically inducible factor, or a chromatin visualization factor. (Item 22) The system according to any one of items 1 to 21, wherein the type V-I CRISPR-Cas effector protein contains at least one nuclear localization signal (NLS), at least one nuclear export signal (NES), or both. (Item 23) The system according to any one of items 1 to 22, comprising the nucleic acid encoding the type V-I CRISPR-Cas effector protein operably linked to a promoter. (Item 24) The system according to item 23, wherein the promoter is a constitutive promoter. (Item 25) The system according to item 23, wherein the nucleic acid encoding the type V-I CRISPR-Cas effector protein is codon-optimized for expression in cells. (Item 26) The system according to item 23, wherein the nucleic acid encoding the type V-I CRISPR-Cas effector protein, which is operably linked to a promoter, is in a vector. (Item 27) The system according to item 26, wherein the vector is selected from the group consisting of a retroviral vector, a lentiviral vector, a phage vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector. (Item 28) The system according to any one of items 1 to 26, which is present in a delivery system selected from the group consisting of nanoparticles, liposomes, exosomes, microvesicles, and gene guns. (Item 29) The system according to any one of items 1 to 28, further comprising a target DNA or a nucleic acid encoding the target DNA, wherein the target DNA comprises a sequence having the ability to hybridize to the spacer sequence of the RNA guide. (Item 30) The system according to any one of items 1 to 29, further comprising a donor template nucleic acid. (Item 31) The system according to item 30, wherein the donor template nucleic acid is DNA or RNA. (Item 32) A cell comprising the system according to any one of items 1 to 31. (Item 33) The cell according to item 32, wherein the cell is a eukaryotic cell. (Item 34) The cell according to item 32, wherein the cell is a prokaryotic cell. (Item 35) A method for targeting and editing a target nucleic acid, the method comprising contacting the target nucleic acid with the system according to any one of items 1 to 31. (Item 36) A method for non-specific degradation of single-stranded DNA according to the recognition of target DNA, the method comprising contacting the system according to any one of items 1 to 36 with the target nucleic acid. (Item 37) A method for targeting and nicking the non-spacer complementary strand of double-stranded target DNA according to the recognition of the spacer complementary strand of double-stranded target DNA, the method comprising contacting the system according to any one of items 1 to 36 with the double-stranded target DNA. (Item 38) A method for targeting and cleaving double-stranded target DNA, the method comprising contacting the system according to any one of items 1 to 31 with the double-stranded target DNA. (Item 39) The method according to item 40, wherein the non-spacer complementary strand of the double-stranded target DNA is nicked before the spacer complementary strand of the double-stranded target nucleic acid is nicked. (Item 40) A method for detecting a target nucleic acid in a sample, (a) contacting the system according to any one of items 1 to 31 and a labeled reporter nucleic acid with the sample, wherein when the crRNA hybridizes to the target nucleic acid, cleavage of the labeled reporter nucleic acid occurs; and (b) measuring a detectable signal generated by cleavage of the labeled reporter nucleic acid, thereby detecting the presence of the target nucleic acid in the sample. (Item 41) The method according to item 40, further comprising comparing the level of the detectable signal with a reference signal level and determining the amount of the target nucleic acid in the sample based on the level of the detectable signal. (Item 42) The method according to item 41, wherein the measurement is performed using gold nanoparticle detection, fluorescence polarization, colloid phase transition / dispersion, electrochemical detection, or semiconductor-based sensing. (Item 43) The method according to item 42, wherein the labeled reporter nucleic acid comprises a fluorescent dye pair, a fluorescence resonance energy transfer (FRET) pair, or a quencher / fluorophore pair, and when the labeled reporter nucleic acid is cleaved by the effector protein, an increase or decrease in the amount of signal generated by the labeled reporter nucleic acid occurs. (Item 44) A method for specifically editing a double-stranded nucleic acid, (a) A type V-I CRISPR-Cas effector and one other enzyme having sequence-specific nicking activity, and a crRNA that induces the type V-I CRISPR-Cas effector to nick the opposite strand compared to the activity of the other sequence-specific nickase; and (b) contacting the double-stranded nucleic acid for a sufficient length of time under sufficient conditions; A method in which the formation of a double-strand break occurs by the method. (Item 45) A method for editing a double-stranded nucleic acid, (a) A fusion protein comprising the type V-I CRISPR-Cas effector, a protein domain having DNA modification activity, and an RNA guide that targets the double-stranded nucleic acid; and (b) contacting the double-stranded nucleic acid for a sufficient length of time under sufficient conditions; A method in which the type V-I CRISPR-Cas effector of the fusion protein is modified to nick the non-target strand of the double-stranded nucleic acid. (Item 46) A method for inducing genotype-specific or transcription state-specific cell death or dormancy in a cell, comprising contacting the cell with the system according to any one of items 1 to 31, and when the RNA guide hybridizes to the target DNA, collateral DNase activity-mediated cell death or dormancy occurs. (Item 47) The method according to item 47, wherein the cell is a prokaryotic cell. (Item 48) The method according to item 47, wherein the cell is a eukaryotic cell. (Item 49) The method according to item 48, wherein the cell is a mammalian cell. (Item 50) The method according to item 49, wherein the cell is a cancer cell. (Item 51) The method according to item 47, wherein the cell is an infectious cell or a cell infected with an infectious pathogen. (Item 52) The method according to item 51, wherein the cell is a cell infected with a virus, a cell infected with a prion, a fungal cell, a protozoan, or a parasitic cell. (Item 53) A method for treating a condition or disease in a subject in need thereof, comprising administering to the subject the system according to any one of items 1 to 31, wherein the spacer sequence is complementary to at least 15 nucleotides of a target nucleic acid associated with the condition or disease; wherein the type V-I CRISPR-Cas effector protein associates with the RNA guide to form a complex; the complex binds to a target nucleic acid sequence complementary to the at least 15 nucleotides of the spacer sequence; and when the complex binds to the target nucleic acid sequence, the type V-I CRISPR-Cas effector protein cleaves the target nucleic acid, thereby treating the condition or disease of the subject. (Item 54) The method according to item 53, wherein the condition or disease is cancer or an infectious disease. (Item 55) The method according to item 54, wherein the condition or disease is cancer, and the cancer is selected from the group consisting of Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myelogenous leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, and bladder cancer. (Item 56) The system or cell according to any one of items 1 to 34 for use as a medicament. (Item 57) The system or cell according to any one of items 1 to 35 for use in the treatment or prevention of cancer or an infectious disease. (Item 58) The system or cell for use according to item 57, wherein the cancer is selected from the group consisting of Wilms tumor, Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid cancer, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myelogenous leukemia, Hodgkin lymphoma, non-Hodgkin lymphoma, and bladder cancer. (Item 59) In vitro or ex vivo a) A method for targeting and editing a target nucleic acid; b) A method for non-specific degradation of single-stranded DNA in response to recognition of a DNA target nucleic acid; c) A method for targeting and nicking the non-spacer complementary strand of the double-stranded target DNA in response to recognition of the spacer complementary strand of the double-stranded target DNA; d) A method for targeting and cleaving a double-stranded target DNA; e) A method for detecting a target nucleic acid in a sample; f) A method for specifically editing a double-stranded nucleic acid; g) A method for base editing of a double-stranded nucleic acid; h) Method for inducing genotype-specific or transcription state-specific cell death or dormancy in cells i) Method for creating indels in double-stranded target DNA; j) Method for inserting a sequence into double-stranded target DNA, or k) Method for deleting or inverting a sequence in double-stranded target DNA Use of the system or cell according to any one of items 1 to 35 in. (Item 60) a) Method for targeting and editing a target nucleic acid; b) Method for non-specific degradation of single-stranded DNA according to the recognition of a DNA target nucleic acid; c) Method for targeting and nicking the non-spacer complementary strand of the double-stranded target DNA according to the recognition of the spacer complementary strand of the double-stranded target DNA; d) Method for targeting and cleaving double-stranded target DNA; e) Method for detecting a target nucleic acid in a sample; f) Method for specifically editing double-stranded nucleic acids; g) Method for base editing of double-stranded nucleic acids; h) Method for inducing genotype-specific or transcription state-specific cell death or dormancy in cells; i) Method for creating indels in double-stranded target DNA; j) Method for inserting a sequence into double-stranded target DNA, or k) Method for deleting or inverting a sequence in double-stranded target DNA Use of the system or cell according to any one of items 1 to 35 in, wherein the method does not include a process of modifying the identity of human germline genes and does not include a method for treating the human or animal body. (Item 61) The method according to item 38 or 53, wherein the cleavage of the target DNA or target nucleic acid results in the formation of indels. (Item 62) The method according to item 38 or 53, wherein the cleavage of the target DNA or target nucleic acid results in the insertion of a nucleic acid sequence. (Item 63) The method according to item 38 or 53, wherein cleavage of the target DNA or target nucleic acid comprises cleavage of the target DNA or target nucleic acid at two sites, resulting in deletion or inversion of the sequence between the two sites. (Item 64) The system according to any one of items 1 to 31, wherein the system lacks tracrRNA. (Item 65) The system according to any one of items 1 to 31, wherein the type V-I CRISPR-Cas effector protein and the type V-I RNA guide form a complex that modifies the target nucleic acid by associating with the target nucleic acid. (Item 66) The system according to any one of items 1 to 31, wherein the spacer sequence is 15 to 47 nucleotides in length, for example, 20 to 40 nucleotides in length, or 24 to 38 nucleotides in length. (Item 67) A eukaryotic cell comprising a modified target locus of interest, wherein the target locus of interest has been modified by the method according to any one of items 1 to 66 or by use of a composition thereof. (Item 68) The modification of the target locus of interest is (i) the eukaryotic cell comprising a change in the expression of at least one gene product; (ii) the eukaryotic cell comprising a change in the expression of at least one gene product, wherein the expression of the at least one gene product is increased; (iii) the eukaryotic cell comprising a change in the expression of at least one gene product, wherein the expression of the at least one gene product is decreased; or (iv) the eukaryotic cell comprising an edited genome The eukaryotic cell according to item 67, which results in (Item 69) The eukaryotic cell according to item 67 or 68, wherein the eukaryotic cell comprises mammalian cells. (Item 70) The eukaryotic cell according to item 69, wherein the mammalian cell comprises human cells. (Item 71) A eukaryotic cell or a eukaryotic cell line containing the same, or a progeny thereof, according to any one of items 67 to 69. (Item 72) A multicellular organism containing one or more cells according to any one of items 67 to 69. (Item 73) A plant or animal model containing one or more cells according to any one of items 67 to 69. (Item 74) A method for producing a plant having a modified target trait encoded by a target gene, which comprises contacting a plant cell with the system according to any one of items 1 to 31, thereby performing either modification or introduction of the target gene, and regenerating a plant from the plant cell. (Item 75) A method for identifying a target trait in a plant, wherein the target trait is encoded by a target gene and comprises contacting a plant cell with the system according to any one of items 1 to 34, thereby identifying the target gene. (Item 76) The method according to item 75, further comprising introducing the identified target gene into a plant cell, a plant cell line, or a plant germplasm, and generating a plant therefrom, whereby the plant contains the target gene. (Item 77) The method according to item 76, wherein the plant exhibits the target trait. (Item 78) A method for targeting and cleaving single-stranded target DNA, which comprises contacting the target nucleic acid with the system according to any one of items 1 to 31. (Item 79) The method according to item 69, wherein the disease or disorder is infectious and the infectious pathogen is selected from the group consisting of human immunodeficiency virus (HIV), herpes simplex virus type 1 (HSV1), and herpes simplex virus type 2 (HSV2). (Item 80) The method according to item 38, wherein both strands of the target DNA are cleaved at different sites, resulting in a sticky-end type of cleavage. (Item 81) The method according to item 38, wherein both strands of the target DNA are cleaved at the same site, resulting in a blunt-end type of double-strand break (DSB). (Item 82) The method according to item 53, wherein the disease state or disorder is selected from the group consisting of cystic fibrosis, Duchenne muscular dystrophy, Becker muscular dystrophy, α1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington's disease, fragile X syndrome, Friedreich's ataxia, amyotrophic lateral sclerosis, frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, hypercholesterolemia, Leber congenital amaurosis, sickle cell disease, and β-thalassemia.
[0098] These figures include a series of schematic diagrams representing the results of locus analysis of various protein clusters, as well as nucleic acid and amino acid sequences.
Brief Description of the Drawings
[0099]
Figure 1A
Figure 1B
Figures 2A-2B
Figure 3
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 6A
Figure 6B
Figure 7A
Figure 7B
Figure 8A
Figure 8B
Figure 9A
Figure 9B
Figure 10A
Figure 10B
Figure 11A
Figure 11B
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17A
Figure 17B
Figure 18A
Figure 18B
Figure 19A
Figure 19B
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24A
Figure 24B
Figure 25A
Figure 25B
Figure 26A
Figure 26B
Figure 27A
Figure 27B
Figure 28A
Figure 28B
Figure 29
Figure 30
Figure 31A
Figure 31B
Figure 32A
Figure 32B
Figure 33A
Figure 33B
Figure 34A
Figure 34B
Figure 35A
Figure 35B
Figure 36
Figure 37
Figure 38A
Figure 38B
Figure 39A
Figure 39B
Figure 40
Figure 41A
Figure 41B
Modes for Carrying Out the Invention
[0100] The broad natural diversity of CRISPR-Cas defense systems includes various active mechanisms and functional elements that can be harnessed for programmable biotechnology. In natural systems, these mechanisms and parameters achieve efficient defense against foreign DNA and viruses while providing self-nonself discrimination to avoid self-targeting. In engineered systems, the same mechanisms and parameters also provide a diverse toolbox of molecular technologies and define the boundaries of the targeting space. For example, systems Ca9 and Ca13a have canonical DNA and RNA endonuclease activities, and their targeting spaces are defined by protospacer adjacent motifs (PAMs) on the target DNA and protospacer adjacent sites (PFSs) on the target RNA, respectively.
[0101] Using the methods described herein, additional mechanisms and parameters capable of expanding the ability of nucleic acid manipulation programmable by RNA were discovered within a single subunit class 2 effector system.
[0102] In one aspect, the present disclosure relates to the use of computational methods and algorithms for exploring and identifying novel protein families that exhibit strong co-occurrence patterns with certain other features within naturally occurring genomic sequences. In certain embodiments, these computational methods relate to identifying protein families that co-occur in close proximity to CRISPR arrays. However, the methods disclosed herein are useful in identifying proteins that naturally occur within a range in close proximity to other features, both non-coding and protein-coding (e.g., fragments of phage sequences in the non-coding regions of bacterial loci; or CRISPR Cas1 proteins). It is understood that the methods and calculations described herein may be implemented on one or more computing devices.
[0103] In some embodiments, a set of genomic sequences is obtained from a genomic or metagenomic database. The database includes short reads, or contig-level data, or assembled scaffolds, or complete genomic sequences of organisms. Similarly, the database may include genomic sequence data from prokaryotes or eukaryotes, or data from metagenomic environmental samples. Examples of database repositories include RefSeq of the National Center for Biotechnology Information (NCBI), GenBank of NCBI, Whole Genome Shotgun (WGS) of NCBI, and Integrated Microbial Genomes (IMG) of the Joint Genome Institute (JGI).
[0104] In some embodiments, a minimum size requirement is imposed for the selection of genomic sequence data of a specified minimum length. In certain exemplary embodiments, the minimum contig length may be 100 nucleotides, 500 nt, 1 kb, 1.5 kb, 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 20 kb, 40 kb, or 50 kb.
[0105] In some embodiments, known or predicted proteins are extracted from complete or a selected set of genomic sequence data. In some embodiments, known or predicted proteins are taken from extracting the coding sequence (CDS) annotations provided by a source database. In some embodiments, predicted proteins are determined by applying computational methods to identify proteins from nucleotide sequences. In some embodiments, the GeneMark suite is used to predict proteins from genomic sequences. In some embodiments, Prodigal is used to predict proteins from genomic sequences. In some embodiments, multiple protein prediction algorithms are used for the same set of sequence data, and duplicates may be removed from the resulting set of proteins.
[0106] In some embodiments, CRISPR arrays are identified from genomic sequence data. In some embodiments, PILER-CR is used to identify CRISPR arrays. In some embodiments, the CRISPR Recognition Tool (CRT) is used to identify CRISPR arrays. In some embodiments, CRISPR arrays are identified by a discovery approach that identifies nucleotide motifs that are repeated a minimum number (e.g., 2, 3, or 4 times), where the interval between consecutive occurrences of the repeated motif does not exceed a specified length (e.g., 50, 100, or 150 nucleotides). In some embodiments, multiple CRISPR array identification tools are used for the same set of sequence data, and duplicates may be removed from the resulting set of CRISPR arrays.
[0107] In some embodiments, proteins that are very close to the CRISPR array are identified. In some embodiments, proximity is defined as a nucleotide distance and may be within 20 kb, 15 kb, or 5 kb. In some embodiments, proximity is defined as the number of open reading frames (ORFs) between the protein and the CRISPR array, and specific exemplary distances can be 10, 5, 4, 3, 2, 1, or 0 ORFs. Proteins identified as being within a very close range to the CRISPR array are then grouped into homologous protein clusters. In some embodiments, blastclust is used to form protein clusters. In certain other embodiments, mmseqs2 is used to form protein clusters.
[0108] To establish a strong co-occurrence pattern with the CRISPR array among members of the protein cluster, a BLAST search of each member of the protein family may be performed against a pre-compiled complete set of known and predicted proteins. In some embodiments, UBLAST or mmseqs2 may be used to search for similar proteins. In some embodiments, the search may be performed only for a representative subset of the proteins within the family.
[0109] In some embodiments, co-occurrence is determined by ranking or filtering the clusters of proteins within a very close range to the CRISPR array by a metric. One exemplary metric is the ratio of the number of elements within the protein cluster to the number of BLAST matches up to a specific E-value threshold. In some embodiments, a fixed E-value threshold may be used. In other embodiments, the E-value threshold may be determined by the most distant member of the protein cluster. In some embodiments, a global set of proteins is clustered and the co-occurrence metric is the ratio of the number of elements in the CRISPR-related cluster to the number of elements in one or more global clusters included.
[0110] In some embodiments, by using a manual review process, the potential functionality and minimal set of components of a system engineered based on the naturally occurring locus structure of proteins in a cluster are evaluated. In some embodiments, tabular representations of protein clusters can be useful for manual review, which can include information including pairwise sequence similarity, phylogenetic trees, source organisms / environments, predicted functional domains, and graphical depictions of locus structures. In some embodiments, the graphical depiction of the locus structure can filter neighboring protein families with high representativeness. In some embodiments, representativeness may be calculated by the ratio of the number of relevant neighboring proteins to the size of one or more of the included global clusters. In certain exemplary embodiments, the tabular representation of the protein cluster can include a depiction of the CRISPR array structure of the naturally occurring locus. In some embodiments, the tabular representation of the protein cluster can include a depiction of the number of conserved direct repeats relative to the length of the putative CRISPR array, or the number of unique spacer sequences relative to the length of the putative CRISPR array. In some embodiments, the tabular representation of the protein cluster includes depictions of various co-occurrence metrics of putative effectors with the CRISPR array, can predict novel CRISPR-Cas systems, and identify their components.
[0111] Pooled screening To efficiently verify the activity of an engineered novel CRISPR-Cas system and simultaneously evaluate various activation mechanisms and functional parameters in an unbiased manner, a novel pooled screening approach is used in Escherichia coli (E. coli). First, from the computational identification of conserved proteins and non-coding elements of the novel CRISPR-Cas system, individual components are assembled into a single artificial expression vector (based on the pET-28a+ backbone in one embodiment) using DNA synthesis and molecular cloning. In a second embodiment, the effector and non-coding elements are transcribed into a single mRNA transcript, and individual effectors are translated using different ribosome binding sites.
[0112] Second, replace the native crRNA and targeting spacer with a library of unprocessed crRNAs containing non-native spacers targeting the second plasmid pACYC184. Clone this crRNA library into a vector backbone (e.g., pET-28a+) containing a protein effector and non-coding elements, and then subsequently transform this library into Escherichia coli (E. coli) together with the pACYC184 plasmid target. As a result, each resulting E. coli cell contains only one targeting spacer. In an alternative embodiment, a library of unprocessed crRNAs containing non-native spacers further targets essential E. coli genes such as those cited from sources such as those described in Baba et al. (2006) Mol. Syst. Biol. 2:2006.0008; and Gerdes et al. (2003) J. Bacteriol. 185(19):5673-84 (the entire contents of each of which are incorporated herein by reference). In this embodiment, the positive targeting activity of the novel CRISPR-Cas system that disrupts essential gene function results in cell death or growth arrest. In some embodiments, an essential gene targeting spacer can be combined with the pACYC184 target to add another dimension to the assay. In other embodiments, the non-coding sequence adjacent to the CRISPR array, putative effector or accessory open reading frame, and predicted anti-repeat that serves as an indicator of the tracrRNA element are concatenated together, cloned into pACYC184, and expressed by the lac and IPTG-inducible T7 promoter.
[0113] Third, grow Escherichia coli (E. coli) under antibiotic selection. In one embodiment, triple antibiotic selection: kanamycin to confirm the success of transformation of the pET-28a+ vector containing the engineered CRISPR-Cas effector system, and chloramphenicol and tetracycline to confirm the success of co-transformation of the pACYC184 target vector are used. pACYC184 usually confers resistance to chloramphenicol and tetracycline. Under antibiotic selection, due to the positive activity of the novel CRISPR-Cas system targeting this plasmid, cells that actively express the effector, non-coding elements, and specific active elements of the crRNA library will be eliminated. Examining the surviving cell population at a later time point compared to an earlier time point typically results in a depletion of the signal compared to the inactive crRNA. In some embodiments, dual antibiotic selection is used. For example, removing either chloramphenicol or tetracycline to remove the selection pressure can provide new information regarding the targeting substrate, sequence specificity, and efficacy. In some embodiments, only kanamycin is used to confirm the success of transformation of the pET-28a+ vector containing the engineered CRISPR-Cas effector system. This embodiment is suitable for libraries containing spacers targeting essential E. coli genes because no additional selection other than kanamycin is required to observe growth changes. In this embodiment, chloramphenicol and tetracycline dependencies are removed, and their targets (if present) in the library provide additional sources of negative or positive information regarding the targeting substrate, sequence specificity, and efficacy.
[0114] The pACYC184 plasmid contains a diverse set of features and sequences that can affect the activity of the CRISPR-Cas system. By mapping the active crRNAs from a pooled screen to pACYC184, a wide range of activity patterns that may suggest various mechanisms of action and functional parameters can be provided without being constrained by hypotheses. In this way, the features required for the reconstitution of novel CRISPR-Cas systems in heterologous prokaryotic species can be more comprehensively tested and studied.
[0115] Specific important advantages of the in vivo pooled screen described herein include the following: (1) Versatility - Plasmid design enables the expression of multiple effectors and / or non-coding elements; Library cloning strategies achieve the expression of crRNAs in both transcriptional directions predicted computationally; (2) By using comprehensive testing of mechanisms of action and functional parameters, various interference mechanisms including DNA or RNA cleavage can be evaluated; Co-occurrence of features such as transcription, plasmid DNA replication, etc.; and the flanking sequences for the crRNA library can be examined to reliably determine the PAM of 4N complexity equivalence; (3) Sensitivity - pACYC184 is a low-copy plasmid, and even with a slight interference rate, it can remove the antibiotic resistance encoded by the plasmid, thus achieving high sensitivity for CRISPR-Cas activity; and (4) Efficiency - Pooled screening includes optimized molecular biology steps that achieve higher speed and throughput for RNA sequencing, and protein expression samples can be directly collected from the viable cells in the screen.
[0116] As will be discussed in more detail in the following examples, the novel CRISPR-Cas family described herein was evaluated by using this in vivo pool-type screen to assess its operable elements, mechanisms and parameters, and its ability to be active and reprogrammed in engineered systems outside its native cellular environment.
[0117] In vitro pool-type screening In vitro pool-type screening methods can also be used, which complement the in vivo pool-type screen. In vitro pool-type screens enable rapid biochemical characterization and reduction from the CRISPR system to the essential components required for system activity. In one embodiment, by using a cell-free in vitro transcription and translation (IVTT) system, RNA and proteins are directly synthesized from the DNA encoding the non-coding and effector proteins of the CRISPR system, thus enabling a faster and higher-throughput method for evaluating a greater number of different and distinct CRISPR-Cas effector systems compared to conventional biochemical assays that rely on FPLC-purified proteins. In addition to enabling higher-throughput and more efficient biochemical reactions, in vitro screening has several advantages that complement the in vivo pool-type screening methods described above.
[0118] (1) Direct observation of both enrichment and depletion signals - In vitro pool-type screening can identify specific cleavage sites, cleavage patterns, and sequence motifs for active effector systems by directly capturing and sequencing cleavage products, i.e., cleavage enrichment, and negative signals from depletion of specific targets in the non-cleaved population, where target depletion is used as a proxy for activity, i.e., target depletion readout. Since the in vivo pool-type screen utilizes the target depletion readout, the enrichment mode provides additional insights into effector activity.
[0119] (2) Higher control of reaction components and environment - The well-defined components and activities of the proprietary IVTT enable the identification of the minimal components required for further active translation through the precise control of reaction components when compared to the complex E. coli cell environment of the in vivo screen. Additionally, non-natural modifications may be made to the reaction components so as to enhance activity or facilitate lead-out; for example, adding phosphorothioate linkages to ssDNA and dsDNA substrates reduces noise by restricting substrate degradation by exonucleases.
[0120] (3) Robustness against toxic / growth-inhibiting proteins - For proteins that can be toxic to E. coli cell growth, in vitro pool-type screens enable functional screening without exposing live cells to growth inhibition. This ultimately enables higher versatility in protein selection and screening.
[0121] The novel CRISPR-Cas family described herein was evaluated by using a combination of in vivo and in vitro pool-type screens to assess its operable elements, mechanisms, and parameters, and its ability to be active and reprogrammed in an engineered system outside its native cellular environment.
[0122] Class 2 CRISPR-Cas effector having an RuvC domain In one aspect, the disclosure provides a class 2 CRISPR-Cas system referred to herein as the CLUST.029130 (type V-I) CRISPR-Cas system. This class 2 CRISPR-Cas system includes an isolated CRISPR-associated protein having an RuvC domain and an isolated crRNA comprising a spacer sequence complementary to a target nucleic acid sequence, such as a DNA sequence, also referred to as an RNA guide, guide RNA, or gRNA.
[0123] Preferably, the CRISPR-Cas effector protein having an RuvC domain is an RuvC III motif, X1SHX4DX6X7 (SEQ ID NO: 200) (wherein X1 is S or T, X4 is Q or L, X6 is P or S, and X7 is F or L); an RuvC I motif, X1XDXNX6X7XXXX 11 (SEQ ID NO: 201) (wherein X1 is A or G or S, X is any amino acid, X6 is Q or I, X7 is T or S or V, and X 11 is T or A); and an RuvC II motif, X1X2X3E (SEQ ID NO: 210) (wherein X1 is C or F or I or L or M or P or V or W or Y, X2 is C or F or I or L or M or P or R or V or W or Y, and X3 is C or F or G or I or L or M or P or V or W or Y), and may comprise one or motifs from a set of.
[0124] Preferably, the type V-I CRISPR-Cas system comprises a CRISPR-Cas effector having an RuvC domain and a type V-I crRNA. Preferably, the Cas12i effector is about 1100 amino acids or less in length and comprises a functional PAM interaction domain that recognizes the PAM in the target DNA. The type V-I CRISPR-Cas effector protein has the ability to bind to a type V-I RNA guide to form a type V-I CRISPR-Cas system, where the type V-I RNA guide comprises a stem-loop structure of a 5-nucleotide stem and a 6-, 7-, or 8-nucleotide loop. The type V-I CRISPR-Cas system has the ability to target and bind to sequence-specific DNA even in the absence of tracrRNA.
[0125] In some embodiments, a type V-I CRISPR-Cas effector protein and a type V-I RNA guide form a binary complex that may include other components. The binary complex is activated upon binding to a nucleic acid substrate (i.e., a sequence-specific substrate or target nucleic acid) complementary to the spacer sequence in the RNA guide. In some embodiments, the sequence-specific substrate is double-stranded DNA. In some embodiments, the sequence-specific substrate is single-stranded DNA. In some embodiments, sequence specificity requires complete identity between the spacer sequence in the RNA guide (e.g., crRNA) and the target substrate. In other embodiments, sequence specificity requires partial (continuous or discontinuous) identity between the spacer sequence in the RNA guide (e.g., crRNA) and the target substrate. Sequence specificity, in certain embodiments, further requires complete identity between a protospacer adjacent motif ("PAM") sequence proximal to the spacer sequence and a canonical PAM sequence recognized by the CRISPR-associated protein. In some examples, complete PAM sequence identity is not required and partial identity is sufficient for sequence-specific association of the binary complex with the DNA substrate.
[0126] In some embodiments, the target nucleic acid substrate is double-stranded DNA (dsDNA). In some embodiments, the target nucleic acid substrate is dsDNA and includes a PAM. In some embodiments, the binary complex modifies a target sequence-specific dsDNA substrate upon binding thereto. In some embodiments, the binary complex preferentially nicks the non-target strand of the target dsDNA substrate. In some embodiments, the binary complex cleaves both strands of the target dsDNA substrate. In some embodiments, the binary complex cleaves both strands of the target dsDNA substrate with sticky ends. In some embodiments, the binary complex creates a blunt-ended double-strand break (DSB) in the target dsDNA substrate.
[0127] In some embodiments, the target nucleic acid substrate is single-stranded DNA (ssDNA). In some embodiments, the target nucleic acid substrate is ssDNA and does not contain a PAM. In some embodiments, the binary complex modifies the target sequence-specific ssDNA substrate upon binding thereto. In some embodiments, the binary complex cleaves the target ssDNA substrate.
[0128] In some embodiments, the binary complex becomes activated upon binding to the target substrate. In some embodiments, the activated complex exhibits "multiple turnover" activity, and thus upon acting on the target substrate (e.g., upon cleavage thereof), the activated complex remains in the activated state. In some embodiments, the binary complex exhibits "single turnover" activity, and thus upon acting on the target substrate, the binary complex returns to the inactive state. In some embodiments, the activated complex exhibits non-specific (i.e., "collateral") cleavage activity, and thus the activated complex cleaves nucleic acids that have no sequence similarity to the target. In some embodiments, the collateral nucleic acid substrate is ssDNA.
[0129] CRISPR Enzyme Modification Nuclease-Deficient CRISPR Enzyme When the CRISPR enzymes described herein have nuclease activity, the CRISPR enzymes can be modified to have reduced nuclease activity, e.g., at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or 100% nuclease inactivation when compared to the wild-type CRISPR enzyme. Nuclease activity can be reduced by several methods, e.g., by introducing mutations into the nuclease or PAM interaction domains of the CRISPR enzyme. In some embodiments, the catalytic residues of the nuclease activity are identified and the nuclease activity may be reduced by substituting those amino acid residues with different amino acid residues (e.g., glycine or alanine). Examples of such mutations for Cas12i1 include D647A, E894A, or D948A. Examples of such mutations for Cas12i2 include D599A, E833A, or D886A.
[0130] The inactivated CRISPR enzyme can contain one or more functional domains (e.g., via a fusion protein, a linker peptide, a Gly4Ser (GS) peptide linker, etc.) or be associated with it (e.g., through co-expression of multiple proteins). Such functional domains can have various activities, e.g., methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, DNA cleavage activity, nucleic acid binding activity, and switch activity (e.g., light-inducible). In some embodiments, the functional domains are Kruppel-associated box (KRAB), VP64, VP16, Fok1, P65, HSF1, MyoD1, and biotin-APEX.
[0131] By positioning one or more functional domains on an inactivated CRISPR enzyme, it becomes possible to achieve the correct spatial orientation for the functional effect attributed to that functional domain to target the target. For example, when the functional domain is a transcriptional activator (e.g., VP16, VP64, or p65), the transcriptional activator is placed in a spatial orientation that enables it to affect the transcription of the target. Similarly, a transcriptional repressor is positioned to affect the transcription of the target, and a nuclease (e.g., Fok1) is positioned to cleave or partially cleave the target. In some embodiments, the functional domain is located at the N-terminus of the CRISPR enzyme. In some embodiments, the functional domain is located at the C-terminus of the CRISPR enzyme. In some embodiments, the inactivated CRISPR enzyme is modified to include a first functional domain at the N-terminus and a second functional domain at the C-terminus.
[0132] Split enzyme The present disclosure also provides split versions of the CRISPR enzymes described herein. Split versions of the CRISPR enzyme may be advantageous for delivery. In some embodiments, the CRISPR enzyme is split into two parts of the enzyme, which together comprise a substantially functional CRISPR enzyme.
[0133] The splitting can be done in such a way that one or more catalytic domains are not affected. The CRISPR enzyme may function as a nuclease or may be an inactivated enzyme that is an RNA-binding protein with very little or no catalytic activity, essentially (e.g., due to one or more mutations in its catalytic domain).
[0134] In some embodiments, the nuclease lobe and the α-helix lobe are expressed as separate polypeptides. These lobes do not interact with each other per se, but an RNA guide recruits them into a complex that recapitulates the activity of the full-length CRISPR enzyme and catalyzes site-specific DNA cleavage. Using a modified RNA guide, dimerization is prevented, abolishing split enzyme activity and enabling the development of an inducible dimerization system. Split enzymes are described, for example, in Wright, Addison V., et al. “Rational design of a split-Cas9 enzyme complex,” Proc. Nat’l. Acad. Sci., 112.10 (2015):2984-2989 (incorporated herein by reference in its entirety).
[0135] In some embodiments, the split enzyme may be fused to a dimerization partner, for example, by utilizing a rapamycin-sensitive dimerization domain. This enables the creation of a chemically inducible CRISPR enzyme for temporally controlling CRISPR enzyme activity. By being split into two fragments in this way, the CRISPR enzyme can be made chemically inducible, and a rapamycin-sensitive dimerization domain can be used for the controlled reassembly of the CRISPR enzyme.
[0136] The split point is typically designed in silico and cloned into a construct. Mutations may be introduced into the split enzyme and non-functional domains may be removed during this process. In some embodiments, the two parts or fragments of the split CRISPR enzyme (i.e., the N-terminal and C-terminal fragments) can form a complete CRISPR enzyme that includes, for example, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% of the sequence of the wild-type CRISPR enzyme.
[0137] Self-activating or inactivating enzymes The CRISPR enzymes described herein may be designed to be self-activating or self-inactivating. In some embodiments, the CRISPR enzyme is self-inactivating. For example, a target sequence can be introduced into the construct encoding the CRISPR enzyme. Thus, when the CRISPR enzyme cleaves the target sequence, it can thereby self-inactivate the expression of the construct encoding the enzyme. Methods for constructing self-inactivating CRISPR systems are described, for example, in Epstein, Benjamin E., and David V. Schaffer. “Engineering a Self-Inactivating CRISPR System for AAV Vectors,” Mol. Ther., 24 (2016): S50 (incorporated herein by reference in its entirety).
[0138] In some other embodiments, additional RNA guides that are expressed under the control of a weak promoter (e.g., the 7SK promoter) can target the nucleic acid sequence encoding the CRISPR enzyme to prevent and / or block its expression (e.g., by interfering with nucleic acid transcription and / or translation). Transfecting a cell with a vector that expresses a CRISPR enzyme, an RNA guide, and an RNA guide that targets the nucleic acid encoding the CRISPR enzyme can lead to efficient destruction of the nucleic acid encoding the CRISPR enzyme, reducing the CRISPR enzyme level and thus restricting genome editing activity.
[0139] In some embodiments, the genome editing activity of the CRISPR enzyme can be regulated through endogenous RNA signatures (e.g., miRNA) in mammalian cells. By using miRNA complementary sequences in the 5'-UTR of the mRNA encoding the CRISPR enzyme, a CRISPR enzyme switch can be created. This switch selectively and efficiently responds to miRNA in the target cells. Thus, this switch can differentially control genome editing by sensing endogenous miRNA activity within a heterogeneous cell population. Therefore, this switch system can provide a framework for cell type-selective genome editing and cell engineering based on intracellular miRNA information (Hirosawa, Moe et al. “Cell-type-specific genome editing with a microRNA-responsive CRISPR-Cas9 switch,” Nucl. Acids Res., 2017 Jul 27;45(13):e118).
[0140] Inducible CRISPR enzyme The CRISPR enzyme may be inducible, for example, photoinducible or chemically inducible. This mechanism enables the activation of the functional domain in the CRISPR enzyme by a known trigger. The photoinducibility can be achieved by various methods known in the art, for example, by designing a fusion complex in which the CRY2PHR / CIBN pair is used in a split CRISPR enzyme (see, for example, Konermann et al., “Optical control of mammalian endogenous transcription and epigenetic states,” Nature, 500.7463 (2013): 472). The chemical inducibility can be achieved, for example, by designing a fusion complex in which the FKBP / FRB (FK506-binding protein / FKBP rapamycin-binding domain) pair is used in a split CRISPR enzyme. Rapamycin is required for the formation of the fusion complex and thus activates the CRISPR enzyme (see, for example, Zetsche, Volz, and Zhang, “A split-Cas9 architecture for inducible genome editing and transcription modulation,” Nature Biotech., 33.2 (2015): 139-142).
[0141] Furthermore, expression of the CRISPR enzyme can be regulated by inducible promoters such as transcriptional activation under the control of tetracycline or doxycycline (Tet-On and Tet-Off expression systems), hormone-inducible gene expression systems (e.g., ecdysone-inducible gene expression systems), and arabinose-inducible gene expression systems. When delivered as RNA, expression of the RNA-targeting effector protein may be regulated by riboswitches that can sense small molecule-like tetracyclines (see, e.g., Goldfless, Stephen J. et al., “Direct and specific chemical control of eukaryotic translation with a synthetic RNA-protein interaction,” Nucl. Acids Res., 40.9 (2012): e64-e64).
[0142] Various embodiments of inducible CRISPR enzymes and inducible CRISPR systems are described, for example, in U.S. Patent No. 8,871,445, U.S. Patent Application Publication No. 2016 / 0208243, and International Publication No. 2016 / 205764, each of which is incorporated herein by reference in its entirety.
[0143] Functional mutations Various mutations or modifications can be introduced into the CRISPR enzymes as described herein to improve specificity and / or robustness. In some embodiments, amino acid residues that recognize the protospacer adjacent motif (PAM) are identified. The CRISPR enzymes described herein may be further modified to recognize different PAMs, for example, by substituting the amino acid residues that recognize the PAM with other amino acid residues. In some embodiments, the CRISPR enzyme can recognize a different PAM, such as as described herein.
[0144] In some embodiments, the CRISPR-related protein comprises at least one (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) nuclear localization signal (NLS) attached to the N-terminus or C-terminus of the protein. Non-limiting examples of NLSs include the NLS of SV40 virus large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO: 300); the NLS from nucleoplasmin (e.g., the nucleoplasmin bipartite NLS of the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 301)); the c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 302) or RQRRNELKRSP (SEQ ID NO: 303); the hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 304); the sequence of the IBB domain from importin-α, RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 305); the sequences VSRKRPRP (SEQ ID NO: 306) and PPKKARED (SEQ ID NO: 307) of the myogenic T protein; the sequence PQPKKKPL (SEQ ID NO: 308) of human p53; the sequence SALIKKKKKMAP (SEQ ID NO: 309) of mouse c-abl IV; the sequences DRLRR (SEQ ID NO: 310) and PKQKKRK (SEQ ID NO: 311) of influenza virus NS1; the sequence RKLKKKIKKL (SEQ ID NO: 312) of hepatitis delta antigen; the sequence REKKKFLKRR (SEQ ID NO: 313) of mouse Mx1 protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 314) of human poly(ADP-ribose) polymerase; and NLS sequences derived from the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 315) of the human glucocorticoid receptor. In some embodiments, the CRISPR-related protein comprises at least one (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) nuclear export signal (NES) attached to the N-terminus or C-terminus of the protein. In preferred embodiments, the C-terminal and / or N-terminal NLS or NES is attached for optimal expression and nuclear targeting in eukaryotic cells, such as human cells.
[0145] In some embodiments, the CRISPR enzymes described herein have one or more functional activities modified by mutating one or more amino acid residues. For example, in some embodiments, the CRISPR enzyme has its helicase activity modified by mutating one or more amino acid residues. In some embodiments, the CRISPR enzyme has its nuclease activity (e.g., endonuclease activity or exonuclease activity) modified by mutating one or more amino acid residues. In some embodiments, the CRISPR enzyme has its ability to functionally associate with an RNA guide modified by mutating one or more amino acid residues. In some embodiments, the CRISPR enzyme has its ability to functionally associate with a target nucleic acid modified by mutating one or more amino acid residues.
[0146] In some embodiments, the CRISPR enzymes described herein have the ability to cleave a target nucleic acid molecule. In some embodiments, the CRISPR enzyme cleaves both strands of the target nucleic acid molecule. However, in some embodiments, the CRISPR enzyme has its cleavage activity modified by mutating one or more amino acid residues. For example, in some embodiments, the CRISPR enzyme may contain one or more mutations such that the enzyme no longer has the ability to cleave the target nucleic acid. In other embodiments, the CRISPR enzyme may contain one or more mutations such that the enzyme has the ability to cleave only one strand of the target nucleic acid (i.e., nickase activity). In some embodiments, the CRISPR enzyme has the ability to cleave the strand of the target nucleic acid that is complementary to the strand to which the RNA guide hybridizes. In some embodiments, the CRISPR enzyme has the ability to cleave the strand of the target nucleic acid to which the RNA guide hybridizes.
[0147] In some embodiments, the CRISPR enzymes described herein may be engineered to contain deletions in one or more amino acid residues in order to reduce the size of the enzyme while retaining one or more desired functional activities (e.g., nuclease activity and the ability to interact functionally with an RNA guide). This truncated CRISPR enzyme may advantageously be used in combination with delivery systems with limited payload capacity.
[0148] In one aspect, the disclosure provides nucleic acid sequences that are at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the nucleic acid sequences described herein. In another aspect, the disclosure also provides amino acid sequences that are at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequences described herein.
[0149] In some embodiments, the nucleic acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, e.g., consecutive or non - consecutive nucleotides) that is the same as the sequence described herein. In some embodiments, the nucleic acid sequence has at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, e.g., consecutive or non - consecutive nucleotides) that is different from the sequence described herein.
[0150] In some embodiments, the amino acid sequence has at least a portion that is the same as the sequences described herein (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid residues, e.g., contiguous or non-contiguous amino acid residues). In some embodiments, the amino acid sequence has at least a portion that is different from the sequences described herein (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid residues, e.g., contiguous or non-contiguous amino acid residues).
[0151] To determine the percent identity between two amino acid sequences, or two nucleic acid sequences, the sequences are aligned for optimal comparison (e.g., gaps may be introduced into one or both of the first and second amino acid or nucleic acid sequences for optimal alignment, and non-homologous sequences may be disregarded for comparison purposes). Generally, the length of the reference sequence aligned for comparison purposes should be at least 80% of the length of the reference sequence, and in some embodiments, at least 90%, 95%, or 100% of the length of the reference sequence. Next, the amino acid residues or nucleotides at the corresponding amino acid positions or nucleotide positions are compared. When a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, then the molecules are identical at that position. The percent identity between two sequences is a function of the number of identical positions shared by those sequences, taking into account the number of gaps that need to be introduced to optimally align the two sequences and the length of each gap. For the purposes of the present disclosure, sequence comparison and determination of percent identity between two sequences can be achieved using the Blosum 62 scoring matrix with a gap penalty of 12, a gap extension penalty of 4, and a frameshift gap penalty of 5.
[0152] In addition to the biochemical and diagnostic applications described herein, the programmable V-I type CRISPR-Cas systems described herein have important applications in eukaryotic cells, such as therapeutic modification of the genome. Examples of modifications include, but are not limited to, genotypic correction, gene knockout, gene sequence insertion / deletion (by homologous recombination repair or other methods), single nucleotide modification, or gene regulation. These gene modification modalities can utilize the nuclease activity of Cas12i, double nicking, or the programmable DNA binding of catalytically inactive Cas12i fused to additional effector domains.
[0153] In some embodiments, the CRISPR-related proteins and accessory proteins described herein can be fused to one or more peptide tags, including His tag, GST tag, FLAG tag, or myc tag. In some embodiments, the CRISPR-related proteins or accessory proteins described herein can be fused to a detectable moiety, such as a fluorescent protein (e.g., green fluorescent protein or yellow fluorescent protein). And in some embodiments, the CRISPR-related proteins or accessory proteins of the present disclosure are fused to a peptide or non-peptide moiety that causes the protein to enter or localize to a tissue, cell, or region of a cell. For example, the CRISPR-related proteins or accessory proteins (such as Cas12i) of the present disclosure may include a nuclear localization sequence (NLS), such as SV40 (simian virus 40) NLS, c-Myc NLS, or other suitable monopartite NLS. The NLS may be fused to the N-terminus and / or C-terminus of the CRISPR-related protein or accessory protein and may be fused alone (i.e., a single NLS) or concatenated (e.g., a chain of 2, 3, 4, etc. NLSs).
[0154] In embodiments where a tag is fused to a CRISPR-related protein, such a tag can facilitate affinity-based or charge-based purification of the CRISPR-related protein, for example, by liquid chromatography or bead separation using immobilized affinity or ion exchange reagents. As a non-limiting example, the recombinant CRISPR-related proteins (such as Cas12i) of the present disclosure include a polyhistidine (His) tag and are loaded onto a chromatography column containing immobilized metal ions for purification (e.g., Zn chelated by a chelating ligand immobilized on a resin). 2+ , Ni 2+ , Cu 2+ ions, and this resin may be a resin prepared individually, a commercially available resin, or a ready-made column such as the HisTrap FF column commercialized by GE Healthcare Life Sciences, Marlborough, Massachusetts). After the loading step, the column is optionally rinsed, for example, using one or more suitable buffer solutions, and then the protein with the His tag added is eluted using a suitable elution buffer. Alternatively or in addition, if the recombinant CRISPR-related protein of the present disclosure utilizes a FLAG tag, such a protein may be purified using immunoprecipitation methods known in the art. Other suitable purification methods for the tagged CRISPR-related proteins or accessory proteins of the present disclosure will be apparent to those skilled in the art.
[0155] The proteins described herein (e.g., CRISPR-related proteins or accessory proteins) can be delivered or used either as nucleic acid molecules or polypeptides. When using nucleic acid molecules, the nucleic acid molecules encoding CRISPR-related proteins can be codon-optimized as discussed in more detail below. The nucleic acids can be codon-optimized for use in any organism of interest, particularly human cells or bacteria. For example, the nucleic acids can be codon-optimized for any non-human eukaryote, including mice, rats, rabbits, dogs, livestock, or non-human primates. Codon usage tables are readily available in the "Codon Usage Database" available, for example, at www.kazusa.orjp / codon / , and these tables can be adapted in several ways. See Nakamura et al. Nucl. Acids Res. 28:292 (2000) (incorporated herein by reference in its entirety). Computer algorithms for codon-optimizing specific sequences for expression in specific host cells are also available, such as Gene Forge (Aptagen; Jacobus, PA).
[0156] In some examples, the nucleic acids of the present disclosure encoding CRISPR-related proteins or accessory proteins for expression in eukaryotic cells (e.g., human, or other mammalian cells) include one or more introns, i.e., one or more non-coding sequences that include a splice donor sequence at a first end (e.g., 5' end) and a splice acceptor sequence at a second end (e.g., 3' end). In various embodiments of the present disclosure, any suitable splice donor / splice acceptor can be used, including, without limitation, simian virus 40 (SV40) introns, β-globin introns, and synthetic introns. Alternatively or in addition, the nucleic acids of the present disclosure encoding CRISPR-related proteins or accessory proteins can include a transcription termination signal, such as a polyadenylation (polyA) signal, at the 3' end of the DNA coding sequence. In some examples, the polyA signal is located very close to or adjacent to an intron, such as an SV40 intron.
[0157] RNA guide In some embodiments, the CRISPR systems described herein include at least one type V-I RNA guide. Many RNA guide configurations are known in the art (see, e.g., International Publication Nos. WO 2014 / 093622 and WO 2015 / 070083, the entire contents of each of which are incorporated herein by reference). In some embodiments, the CRISPR systems described herein include multiple RNA guides (e.g., 2, 3, 4, 5, 6, 7, 8, or more RNA guides).
[0158] In some embodiments, the CRISPR systems described herein include at least one type V-I RNA guide or a nucleic acid encoding at least one type V-I RNA guide. In some embodiments, the RNA guide includes a crRNA. Generally, the crRNAs described herein include direct repeat sequences and spacer sequences. In certain embodiments, the crRNA comprises, consists essentially of, or consists of a direct repeat sequence linked to a guide sequence or spacer sequence. In some embodiments, the crRNA includes a direct repeat sequence, a spacer sequence, and a direct repeat sequence (DR-spacer-DR), which is typical of the precursor crRNA (pre-crRNA) construct in other CRISPR systems. In some embodiments, the crRNA includes a truncated direct repeat sequence and a spacer sequence, which is typical of a processed or mature crRNA. In some embodiments, the CRISPR-Cas effector protein forms a complex with the RNA guide, and the spacer sequence directs the complex to sequence-specific binding with a target nucleic acid complementary to the spacer sequence.
[0159] Preferably, the CRISPR systems described herein include at least one type V-I RNA guide or a nucleic acid encoding a type V-I RNA guide, wherein the RNA guide includes a direct repeat. Preferably, the type V-I RNA guide may form a secondary structure, such as a stem-loop structure, as described herein, for example.
[0160] A direct repeat can comprise two nucleotide stretches that may be complementary to each other, which are separated by intervening nucleotides, so that as a result of the direct repeats hybridizing to form a double-stranded RNA duplex (dsRNA duplex), two complementary nucleotide stretches form a stem, and a stem-loop structure can occur in which the intervening nucleotides form a loop or a hairpin (Figure 3). For example, the intervening nucleotides that form a "loop" have a length of about 6 nucleotides to about 8 nucleotides, or about 7 nucleotides. In different embodiments, the stem can comprise at least 2, at least 3, at least 4, or 5 base pairs.
[0161] Preferably, a direct repeat can comprise two complementary nucleotide stretches that are about 5 nucleotides in length and separated by about 7 intervening nucleotides.
[0162] Exemplary direct repeats of a type V-I system are illustrated in Figure 3, and preferably when deviating from naturally occurring type V-I direct repeats, one of ordinary skill in the art can mimic the structure of such direct repeats illustrated in Figure 3.
[0163] A direct repeat can comprise or consist of about 22 to 40 nucleotides, or about 23 to 38 nucleotides or about 23 to 36 nucleotides.
[0164] In some embodiments, the CRISPR systems described herein comprise multiple RNA guides (e.g., 2, 3, 4, 5, 10, 15, or more) or multiple nucleic acids encoding multiple RNA guides.
[0165] In some embodiments, the CRISPR systems described herein include an RNA guide or a nucleic acid encoding an RNA guide. In some embodiments, the RNA guide comprises or consists of a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid (e.g., hybridize thereto under appropriate conditions), where the direct repeat sequence comprises 5'-CCGUCNNNNNNNGACGG-3' (SEQ ID NO: 202) proximal to its 3' end and adjacent to the spacer sequence. In some embodiments, the RNA guide comprises or consists of a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid (e.g., hybridize thereto under appropriate conditions), where the direct repeat sequence comprises 5'-GUGCCNNNNNNNGGCAC-3' (SEQ ID NO: 203) proximal to its 3' end and adjacent to the spacer sequence. In some embodiments, the RNA guide comprises or consists of a direct repeat sequence and a spacer sequence having the ability to hybridize to a target nucleic acid (e.g., hybridize thereto under appropriate conditions), where the direct repeat sequence is proximal to the 3' end and adjacent to the spacer sequence 5'-GUGUCN 5-6 UGACAX1-3' (SEQ ID NO: 204) (wherein N 5-6 refers to any contiguous sequence of 5 or 6 nucleobases, and X1 refers to C or T or U) is included.
[0166] Examples of pairs of RNA guide direct repeat sequences and effector proteins are provided in Table 5A. In some embodiments, the direct repeat sequence comprises or consists of a nucleic acid sequence listed in Table 5A (e.g., SEQ ID NOs: 6-10, 19-24). In some embodiments, the direct repeat sequence comprises or consists of a nucleic acid having a nucleic acid sequence listed in Table 5A with a truncation of the first 3 5' nucleotides. In some embodiments, the direct repeat sequence comprises or consists of a nucleic acid having a nucleic acid sequence listed in Table 5A with a truncation of the first 4 5' nucleotides. In some embodiments, the direct repeat sequence comprises or consists of a nucleic acid having a nucleic acid sequence listed in Table 5A with a truncation of the first 5 5' nucleotides. In some embodiments, the direct repeat sequence comprises or consists of a nucleic acid having a nucleic acid sequence listed in Table 5A with a truncation of the first 6 5' nucleotides. In some embodiments, the direct repeat sequence comprises or consists of a nucleic acid having a nucleic acid sequence listed in Table 5A with a truncation of the first 7 5' nucleotides. In some embodiments, the direct repeat sequence comprises or consists of a nucleic acid having a nucleic acid sequence listed in Table 5A with a truncation of the first 8 5' nucleotides.
[0167] Multiplexing of RNA guides The CLUST.029130 (type V-I) CRISPR-Cas effector utilizes two or more RNA guides and, accordingly, it has been demonstrated that these effectors, as well as systems and complexes containing them, are capable of targeting multiple different nucleic acid targets. In some embodiments, the CRISPR systems described herein include multiple RNA guides (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, or more RNA guides). In some embodiments, the CRISPR systems described herein include a single RNA strand or a nucleic acid encoding a single RNA strand, where the RNA guides are arranged in tandem. The single RNA strand can include multiple copies of the same RNA guide, multiple copies of different RNA guides, or combinations thereof.
[0168] In some embodiments, the CLUST.029130 (type V-I) CRISPR-Cas effector protein is delivered complexed with multiple RNA guides directed to different target nucleic acids. In some embodiments, the CLUST.029130 (type V-I) CRISPR-Cas effector protein can be co-delivered with multiple RNA guides, each specific for a different target nucleic acid. Methods of multiplexing using CRISPR-associated proteins are described, for example, in U.S. Patent No. 9,790,490 and European Patent No. 3009511, the entire contents of each of which are hereby expressly incorporated by reference.
[0169] RNA Guide Modification Spacer Length The spacer length of the RNA guide may be in the range of about 15 to 50 nucleotides. In some embodiments, the spacer length of the RNA guide is at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, or at least 22 nucleotides. In some embodiments, the spacer length is 15 to 17 nucleotides, 15 to 23 nucleotides, 16 to 22 nucleotides, 17 to 20 nucleotides, 20 to 24 nucleotides (e.g., 20, 21, 22, 23, or 24 nucleotides), 23 to 25 nucleotides (e.g., 23, 24, or 25 nucleotides), 24 to 27 nucleotides, 27 to 30 nucleotides, 30 to 45 nucleotides (e.g., 30, 31, 32, 33, 34, 35, 40, or 45 nucleotides), 30 or 35 to 40 nucleotides, 41 to 45 nucleotides, 45 to 50 nucleotides, or more. In some embodiments, the spacer length of the RNA guide is 31 nucleotides. In some embodiments, the directory repeat length of the RNA guide is at least 21 nucleotides, or 21 to 37 nucleotides (e.g., 23, 24, 25, 30, 35, or 36 nucleotides). In some embodiments, the directory repeat length of the RNA guide is 23 nucleotides.
[0170] The RNA guide sequence may be modified in such a way that it allows for the formation of the CRISPR effector complex and successful binding to the target, but at the same time does not allow for successful nuclease activity (i.e., no nuclease activity / does not cause indels). Such modified guide sequences are referred to as "dead guides" or "dead guide sequences". Such dead guides or dead guide sequences may be catalytically inactive or conformationally inactive in terms of nuclease activity. Dead guide sequences are typically shorter than each guide sequence that results in active RNA cleavage. In some embodiments, the dead guide is 5%, 10%, 20%, 30%, 40%, or 50% shorter than each RNA guide having nuclease activity. The dead guide sequence of the RNA guide may be 13-15 nucleotides in length (e.g., 13, 14, or 15 nucleotides in length), 15-19 nucleotides in length, or 17-18 nucleotides in length (e.g., 17 nucleotides in length).
[0171] Accordingly, in one aspect, the present disclosure provides a non-naturally occurring or engineered CRISPR system comprising a functional CRISPR enzyme as described herein and an RNA guide (gRNA), wherein the gRNA comprises a dead guide sequence, and thus the gRNA has the ability to hybridize to a target sequence such that the CRISPR system is directed to a desired genomic locus in the cell without detectable cleavage activity.
[0172] A detailed description of dead guides is provided, for example, in WO 2016 / 094872 (incorporated herein by reference in its entirety).
[0173] Inducible guide The RNA guide can be made as a component of an inducible system. The inducible nature of this system allows for spatiotemporal control of gene editing or gene expression. In some embodiments, the stimuli for the inducible system include, for example, electromagnetic radiation, acoustic energy, chemical energy, and / or thermal energy.
[0174] In some embodiments, the transcription of the RNA guide can be regulated by an inducible promoter, such as transcriptional activation under the control of tetracycline or doxycycline (Tet-On and Tet-Off expression systems), a hormone-inducible gene expression system (e.g., an ecdysone-inducible gene expression system), and an arabinose-inducible gene expression system. Other examples of inducible systems include, for example, a small molecule two-hybrid transcriptional activation system (FKBP, ABA, etc.), a light-inducible system (phytochrome, LOV domain, or cryptochrome), or a light-inducible transcriptional effector (LITE). These inducible systems are described, for example, in WO 2016 / 205764 and U.S. Pat. No. 8,795,965, both of which are incorporated herein by reference in their entirety.
[0175] Chemical modification Chemical modifications can be applied to the phosphate backbone, sugar, and / or base of an RNA guide. Backbone modifications such as phosphorothioates modify the charge on the phosphate backbone and are useful for oligonucleotide delivery and nuclease resistance (see, e.g., Eckstein, “Phosphorothioates, essential components of therapeutic oligonucleotides,” Nucl. Acid Ther., 24 (2014), pp. 374-387); sugar modifications such as 2'-O-methyl (2'-OMe), 2'-F, and locked nucleic acid (LNA) enhance both base pairing and nuclease resistance (see, e.g., Allerson et al. “Fully 2‘-modified oligonucleotide duplexes with improved in vitro potency and stability compared to unmodified small interfering RNA,” J. Med. Chem., 48.4 (2005): 901-904). Chemically modified bases, particularly 2-thiouridine or N6-methyladenosine, etc., can enable either stronger or weaker base pairing (see, e.g., Bramsen et al., “Development of therapeutic-grade small interfering RNAs by chemical engineering” Front. Genet., 2012 Aug 20; 3:154). Additionally, RNA is suitable for conjugation at both the 5' and 3' ends with various functional moieties including fluorescent dyes, polyethylene glycol, or proteins.
[0176] A wide variety of modifications can be applied to chemically synthesized RNA guide molecules. For example, when oligonucleotides are modified with 2'-OMe to improve nuclease resistance, the binding energy of Watson-Crick base pairing can be changed. Furthermore, 2'-OMe modifications can affect how oligonucleotides interact with transfection reagents, proteins, or any other molecules in cells. The effects of these modifications can be determined by empirical testing.
[0177] In some embodiments, the RNA guide comprises one or more phosphorothioate modifications. In some embodiments, the RNA guide comprises one or more locked nucleic acids for the purpose of enhancing base pairing and / or increasing nuclease resistance.
[0178] For an overview of these chemical modifications, see, for example, Kelley et al., “Versatility of chemically synthesized guide RNAs for CRISPR-Cas9 genome editing,” J. Biotechnol. 2016 Sep 10;233:74-83; International Publication No. WO 2016 / 205764 pamphlet; and U.S. Patent No. 8,795,965 B2 (each of which is incorporated by reference in its entirety).
[0179] Sequence Modifications The sequences and lengths of the RNA guides and crRNAs described herein can be optimized. In some embodiments, the optimized length of the RNA guide may be determined by identifying the processed form of the crRNA or by empirical length studies for the RNA guide of the crRNA.
[0180] The RNA guide can also include one or more aptamer sequences. An aptamer is an oligonucleotide or peptide molecule that can bind to a specific target molecule. The aptamer may be specific for a gene effector, gene activator, or gene repressor. In some embodiments, the aptamer is specific for a protein, which in turn may be specific for and recruit / bind to a specific gene effector, gene activator, or gene repressor. The effector, activator, or repressor can exist in the form of a fusion protein. In some embodiments, the RNA guide has two or more aptamer sequences specific for the same adapter protein. In some embodiments, the two or more aptamer sequences are specific for different adapter proteins. Examples of adapter proteins include, for example, MS2, PP7, Qβ, F2, GA, fr, JP501, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s, and PRR1. Thus, in some embodiments, the aptamer is selected from binding proteins that specifically bind to any one of the adapter proteins as described herein. In some embodiments, the aptamer sequence is an MS2 loop. For a detailed description of aptamers, reference can be made, for example, to Nowak et al., “Guide RNA engineering for versatile Cas9 functionality,” Nucl. Acid. Res., 2016 Nov 16;44(20):9555 - 9564; and WO 2016 / 205764 (which are hereby incorporated by reference in their entirety).
[0181] Guide: Target sequence matching requirements In classical CRISPR systems, the degree of complementarity between the guide sequence and its corresponding target sequence may be about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or 100%. In some embodiments, the degree of complementarity is 100%. The RNA guide may be about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 nucleotides in length, or longer.
[0182] To reduce off-target interactions, for example, to reduce guides that interact with off-target sequences with low complementarity, mutations may be introduced into the CRISPR system so that the CRISPR system can distinguish between the target sequence and off-target sequences with a complementarity higher than 80%, 85%, 90%, or 95%. In some embodiments, the degree of complementarity is 80% - 95%, for example, about 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, or 95% (e.g., to distinguish between a target having 18 nucleotides and an off-target of 18 nucleotides having 1, 2, or 3 mismatches). Thus, in some embodiments, the degree of complementarity between the guide sequence and its corresponding target sequence is higher than 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 99.9%. In some embodiments, the degree of complementarity is 100%.
[0183] In the art, it is known that perfect complementarity is not a requirement if there is sufficient complementarity to be functional. By introducing mismatches, such as one or more mismatches, such as one or two mismatches, between the spacer array and the target array, including the position of the mismatch along the spacer / target, the cleavage efficiency can be exploited. The closer the mismatch, such as a double mismatch, is located towards the center (i.e., not at the 3' or 5' end), the greater the effect on the cleavage efficiency. Thus, the cleavage efficiency can be adjusted by selecting the position of the mismatch along the spacer array. For example, if less than 100% cleavage of the target is desired (e.g., in a cell population), one or two mismatches between the spacer and the target sequence may be introduced into the spacer sequence.
[0184] Optimization of the CRISPR system for use in selected organisms Codon optimization The present invention contemplates all possible variants of nucleic acids such as cDNA that can be created by selecting combinations based on possible codon selections. These combinations are created according to the standard triplet genetic code as applied to polynucleotides encoding naturally occurring variants, and all such variants should be considered specifically disclosed. Nucleotide sequences encoding V-I type CRISPR-Cas related effector protein variants codon-optimized for expression in bacteria (e.g., Escherichia coli (E. coli)) and human cells are disclosed herein. For example, a sequence codon-optimized for human cells can be created by using codons that occur frequently in human cells instead of codons in nucleotide sequences that occur infrequently in human cells. The codon frequencies can be computationally determined by methods known in the art. Computational examples of such codon frequencies for various host cells (e.g., Escherichia coli (E. coli), yeast, insects, Caenorhabditis elegans (C. elegans), Drosophila melanogaster (D. melanogaster), human, mouse, rat, pig, Pichia pastoris (P. pastoris), Arabidopsis thaliana (A. thaliana), maize, and tobacco) are publicly available or can be accessed through resources such as the GenScript® Codon Usage Frequency Table tool (Exemplary codon usage tables for Escherichia coli (E. coli) and human are included below.
[0185]
Table 1
[0186]
Table 2
[0187] Methods of using the CRISPR system The CRISPR systems described herein have a wide variety of utilities, including modification (e.g., deletion, insertion, translocation, inactivation, or activation) of target polynucleotides in a very large number of cell types. This CRISPR system has a wide range of applications, for example, in DNA / RNA detection (e.g., specific high sensitivity enzymatic reporter unlocking: SHERLOCK), nucleic acid tracking and labeling, enrichment assays (extraction of a desired sequence from background), detection of circulating tumor DNA, preparation of next-generation libraries, drug screening, disease diagnosis and prognosis determination, and treatment of various genetic disorders. Without wishing to be bound by any particular theory, CRISPR systems containing Cas12i protein may exhibit increased activity or may be preferentially active upon targeting in certain environments, such as DNA plasmids, supercoiled DNA, or transcriptionally active genomic loci.
[0188] Overview of Genome Editing Systems The term "genome editing system" refers to the engineered CRISPR systems of the present disclosure that have RNA-guided DNA editing activity. The genome editing systems of the present disclosure include at least two components of the CRISPR systems described above: an RNA guide and a cognate CRISPR effector protein. In certain embodiments of the present disclosure, the effector is a Cas12i protein and the RNA guide is a cognate type V-I RNA guide. As described above, these two components form a complex that has the ability to associate with a specific nucleic acid sequence and edit the DNA by creating one or more of, for example, single-strand breaks (SSB or nicks), double-strand breaks (DSB), nucleobase modifications, DNA methylation or demethylation, chromatin modifications, etc., within or around the nucleic acid sequence.
[0189] In certain embodiments, the genome editing system is transiently active (e.g., incorporating an inducible CRISPR effector as discussed above), while in other embodiments, the system is constitutive (e.g., encoded by a nucleic acid in which the expression of the CRISPR system components is controlled by one or more strong promoters).
[0190] The genome editing systems of the present disclosure, upon introduction into a cell, can modify (a) endogenous genomic DNA (gDNA), including without limitation DNA encoding, for example, a gene target of interest, an exon sequence of a gene, an intron sequence of a gene, a regulatory element of a gene or gene group, etc.; (b) extra-genomic endogenous DNA such as mitochondrial DNA (mtDNA); and / or (c) exogenous DNA such as an unintegrated viral genome, plasmid, artificial chromosome, etc. Throughout the present disclosure, these DNA substrates are referred to as "target DNA".
[0191] In examples where genome editing works by the generation of an SSB or DSB, the modifications caused by the system can take the form of short DNA insertions or deletions, which are collectively referred to as "indels". These indels may typically be formed within or proximal to the predicted cleavage site that is proximal to the PAM sequence and / or within the complementary region of the spacer sequence, however in some cases an indel may occur outside of such predicted cleavage sites. Without wishing to be bound by any theory, indels are often thought to be the result of the repair of SSBs or DSBs by an "error-prone" DNA damage repair pathway such as non-homologous end joining (NHEJ).
[0192] In some cases, genome editing is used to generate two DSBs within 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1250, 1500, 1750, or 2000 base pairs of each other, which results in one or more outcomes including the formation of indels at one or both cleavage sites, and the deletion or inversion of the DNA sequence placed between those DSBs.
[0193] Alternatively, the genome editing system of the present disclosure can modify target DNA through the integration of a novel sequence. These novel sequences may be different from the existing sequences of the target DNA (for non-limiting examples, integrated by NHEJ by blunt-end ligation), or may correspond to a DNA template having one or more regions homologous to the region of the DNA to be targeted. The integration of the homologous sequence serving as the template is also referred to as "homologous recombination repair" or "HDR". The template DNA for HDR may be endogenous to the cell, including, without limitation, in the form of a homologous sequence located on the other copy of the same chromosome as the target DNA, a homologous sequence from the same gene cluster as the target DNA, etc. Alternatively, or in addition, the template DNA may be provided exogenously, including, without limitation, as free linear or circular DNA, as DNA (covalently or non-covalently) bound to one or more genome editing system components, or as part of a vector genome.
[0194] In some examples, the editing includes temporary or permanent silencing of genes by CRISPR-mediated interference, as described by Matthew H. Larson et al., “CRISPR interference (CRISPRi) for sequence-specific control of gene expression,” Nature Protocols 8, 2180 - 2196 (2013) (which is incorporated by reference in its entirety and for all purposes).
[0195] The genome editing system can include other components, including, without limitation, one or more heterologous functional domains that mediate site-specific nucleic acid base modification, DNA methylation or demethylation, or chromatin modification. In some cases, the heterologous functional domain covalently binds to a CRISPR-related protein such as Cas12i, for example, by a direct peptide bond or an intervening peptide linker. This type of fusion is described in more detail below. In some embodiments, the heterologous functional domain covalently binds to the crRNA, for example, by chemical crosslinking. And in some embodiments, one or more functional groups may non-covalently associate with the CRISPR-related protein and / or the crRNA. This can be done in various ways, such as by an aptamer added to the crRNA and / or the heterologous functional group, a binding domain configured to bind to a peptide motif fused to the CRISPR-related protein and such a motif fused to the heterologous functional domain, or vice versa.
[0196] The genome editing system design and the results of genome editing are described in more detail in other parts of this specification.
[0197] DNA / RNA Detection In one aspect, the CRISPR-Cas systems described herein can be used in DNA / RNA detection by DNA sensing. By reprogramming a single effector RNA-guided DNase with an RNA guide, a platform for specific single-stranded DNA (ssDNA) sensing can be provided. Upon recognition of its DNA target, the activated CRISPR V-I type effector protein is involved in the "collateral" cleavage of neighboring ssDNA that has no sequence similarity to the target sequence. This collateral cleavage activity programmed by the RNA enables the CRISPR system to detect the presence of specific DNA by non-specific degradation of the labeled ssDNA.
[0198] In the context of DNA detection applications, collateral ssDNase activity can be combined with a reporter, for example, in methods such as the DNA Endonuclease-Targeted CRISPR trans reporter:DETECTR method, which, when combined with amplification, achieves attomolar DNA detection sensitivity (see, e.g., Chen et al., Science, 360(6387):436-439, 2018, which is hereby incorporated by reference in its entirety). One application of the enzymes described herein is the degradation of non-target ssDNA in an in vitro environment. "Reporter" ssDNA molecules linked to fluorophores and quenchers can also be added to this in vitro system along with an unknown DNA sample (either single-stranded or double-stranded). When the target sequence is recognized in the unknown DNA fragment, this surveillance complex containing a V-I type effector cleaves the reporter ssDNA, resulting in a fluorescent readout.
[0199] In other embodiments, the SHERLOCK method (Specific High Sensitivity Enzymatic Reporter UnLOCKing) also provides an in vitro nucleic acid detection platform with attomolar (or single molecule) sensitivity based on nucleic acid amplification and collateral cleavage of reporter ssDNA, enabling real-time detection of targets. The use of CRISPR in SHERLOCK is described in detail, for example, in Gootenberg, et al. “Nucleic acid detection with CRISPR-Cas13a / C2c2,” Science, 356(6336):438-442 (2017), which is hereby incorporated by reference in its entirety.
[0200] In some embodiments, the CRISPR systems described herein can be used in multiplexed error-robust fluorescence in situ hybridization (MERFISH). For such methods, see, for example, Chen et al., “Spatially resolved, highly multiplexed RNA profiling in single cells,” Science, 2015 Apr 24;348(6233):aaa6090 (which is hereby incorporated by reference in its entirety).
[0201] In some embodiments, the CRISPR systems described herein can be used to detect target DNA in a sample (e.g., a clinical sample, a cell, or a cell lysate). The collateral DNase activity of the CLUST.029130 (type V-I) CRISPR-Cas effector protein described herein is activated when the effector protein binds to the target nucleic acid. When bound to the target DNA of interest, a signal is generated or changed (e.g., an increase or decrease in signal) because the effector protein cleaves the labeled detector ssDNA, thereby enabling qualitative and quantitative detection of the target DNA in the sample. Specific detection and quantification of DNA in a sample enables numerous applications, including diagnostic methods.
[0202] In some embodiments, the method comprises: a) contacting the sample with (i) an RNA guide (e.g., crRNA) and / or a nucleic acid encoding the RNA guide, the RNA guide comprising a direct repeat sequence and a spacer sequence capable of hybridizing to a target RNA; (ii) a CLUST.029130 (type V-I) CRISPR-Cas effector protein and / or a nucleic acid encoding the effector protein; and (iii) a labeled detector ssDNA, wherein the effector protein associates with the RNA guide to form a surveillance complex, the surveillance complex hybridizes to the target DNA, and when the surveillance complex binds to the target DNA, the effector protein exhibits collateral DNase activity and cleaves the labeled detector ssDNA; and b) measuring a detectable signal generated by the cleavage of the labeled detector ssDNA, wherein said measuring comprises providing detection of the target DNA in the sample.
[0203] In some embodiments, the method further comprises comparing the detectable signal to a reference signal and determining the amount of target DNA in the sample. In some embodiments, the measuring is performed using gold nanoparticle detection, fluorescence polarization, colloid phase transition / dispersion, electrochemical detection, and semiconductor-based sensing. In some embodiments, the labeled detector ssDNA comprises a pair of fluorescent emission dyes, a pair of fluorescence resonance energy transfer (FRET), or a quencher / fluorophore pair. In some embodiments, when the labeled detector ssDNA is cleaved by the effector protein, the amount of the detectable signal generated by the labeled detector ssDNA decreases or increases. In some embodiments, the labeled detector ssDNA generates a first detectable signal before cleavage by the effector protein and a second detectable signal after cleavage by the effector protein.
[0204] In some embodiments, the detectable signal is generated when the labeled detector ssDNA is cleaved by an effector protein. In some embodiments, the labeled detector ssDNA includes modified nucleobases, modified sugar moieties, modified nucleic acid linkages, or combinations thereof.
[0205] In some embodiments, the method includes multi-channel detection of multiple independent target DNAs (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, or more target RNAs) in a sample, which is by using multiple CLUST.029130 (type V-I) CRISPR-Cas systems, each including a different orthologous effector protein and a corresponding RNA guide, thereby enabling discrimination of multiple target DNAs in the sample. In some embodiments, the method includes multi-channel detection of multiple independent target DNAs in a sample, which is by using multiple examples of CLUST.029130 (type V-I) CRISPR-Cas systems, each including an orthologous effector protein having a distinguishable collateral ssDNase substrate. Methods for detecting DNA in a sample using CRISPR-related proteins are described, for example, in U.S. Patent Application Publication No. 2017 / 0362644, the entire contents of which are incorporated herein by reference.
[0206] Tracking and Labeling of Nucleic Acids Cellular processes rely on a network of molecular interactions among proteins, RNAs, and DNA. Accurate detection of protein-DNA and protein-RNA interactions is key to understanding such processes. In vitro proximity labeling techniques use an affinity tag combined with a reporter group, such as a photoactivatable group, to label polypeptides and DNA near a protein or DNA of interest in vitro. After ultraviolet irradiation, the photoactivatable group reacts with proteins and other molecules in close proximity to the tagged molecule, thereby labeling them. The labeled interacting molecules can then be recovered and identified. This DNA targeting effector protein can be used, for example, to target a probe to a selected DNA sequence. Such applications can also be applied to in vivo imaging of diseases or cell types that are difficult to culture in animal models. Methods for tracking and labeling nucleic acids are described, for example, in U.S. Patent No. 8,795,965; International Publication No. WO 2016 / 205764; and International Publication No. WO 2017 / 070605 (each of which is incorporated herein by reference in its entirety).
[0207] Genome editing using paired CRISPR nickases The CRISPR systems described herein are used in tandem such that two Cas12i nickase enzymes, or one Cas12i enzyme and one other CRISPR Cas enzyme having nickase activity, targeted to opposite strands of a target locus by a pair of RNA guides, can generate double-strand breaks with overhangs. According to this method, double-strand breaks are expected to occur only at the locus where both enzymes generate nicks, thus reducing the likelihood of off-target modifications and therefore increasing the specificity of genome editing. This method is referred to as the “double nicking” or “pair nickase” strategy and is described, for example, in Ran et al., “Double nicking by RNA-guided CRISPR Cas9 for enhanced genome editing specificity,” Cell, 2013 Sep 12;154(6):1380-1389, and Mali et al., “CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering,” Nature Biotechnology, 2013 Aug 01;31:833-838 (both of which are incorporated herein by reference in their entirety).
[0208] The first application of pair nickases demonstrated the utility of this strategy in mammalian cell lines. Applications of pair nickases have been made in the model plant Arabidopsis thaliana (e.g., Fauser et al., “Both CRISPR / Cas-based nucleases and nickases can be used efficiently for genome engineering in Arabidopsis thaliana,” The Plant Journal 79(2):348-59(2014), and Shiml et al., ““The CRISPR / Cas system can be used as nuclease for in planta gene targeting and as paired nickases for directed mutagenesis in Arabidopsis resulting in heritable progeny,”The Plant Journal 80(6):1139-50(2014); crops such as rice (e.g., Mikami et al., “Precision Targeted Mutagenesis via Cas9 Paired Nickases in Rice,”Plant and Cell Physiology 57(5):1058-68(2016) and wheat (e.g., Czermak et al., “A Multipurpose Toolkit to Enable Advanced Genome Engineering in Plants,”Plant Cell 29:1196-1217(2017); bacteria (e.g., Standage-Beier et al., “Targeted Large-Scale Deletion of Bacterial "Genomes Using CRISPR-Nickases," ACS Synthetic Biology 4(11):1217-25(2015); and for therapeutic purposes in primary human cells (e.g., Dabrowska et al., "Precise Excision of the CAG Tract from the Huntingtin Gene by Cas9 Nickases," Frontiers in Neuroscience 12:75(2018), and Kocher et al., "Cut and Paste: Efficient Homology-Directed Repair of a Dominant Negative KRT14 Mutation via CRISPR / Cas9 Nickases," Molecular Therapy 25(11):2585-2598(2017)) (all of which are hereby incorporated by reference in their entirety).
[0209] The CRISPR systems described herein can also be used as pair nickases to detect splice junctions, as described, for example, in Santo & Paik, "A splice junction-targeted CRISPR approach (spJCRISPR) reveals human FOXO3B to be a protein-coding gene," Gene 673:95-101(2018).
[0210] The CRISPR systems described herein can also be used, for example, in Wang et al, "Therapeutic Genome Editing for Myotonic As described in “Dystrophy Type 1 Using CRISPR / Cas9,” Molecular Therapy 26(11):2617-2630(2018), it can also be used as a pair nicking enzyme to insert a DNA molecule into a target locus. The CRISPR systems described herein can also be used, for example, as a single nickase to insert a gene as described in Gao et al, “Single Cas9 nickase induced generation of NRAMP1 knockin cattle with reduced off-target effects,” Genome Biology 18(1):13(2017).
[0211] Enhancement of Base Editing Using CRISPR Nickase The efficiency of CRISPR base editing can be enhanced using the CRISPR systems described herein. In base editing, a protein domain having DNA nucleotide modification activity (e.g., cytidine deamination) is fused to a programmable CRISPR Cas enzyme that has been inactivated by mutation so as to no longer have double-stranded DNA cleavage activity. In some embodiments, using a nickase as the programmable Cas protein, for example, as described in Komor et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage,” Nature 533:420-424(2016), and Nishida et al., “Targeted nucleotide editing using hybrid As described in “Prokaryotic and vertebrate adaptive immune systems,” Science 353(6305):aaf8729(2016)(both of which are hereby incorporated by reference in their entirety), improved base editing efficiency has been shown. Nickases that nick the non-edited strand of the target locus stimulate endogenous DNA repair pathways - such as mismatch repair or long patch base excision repair - that preferentially resolve mismatches generated by base editing on the desired allele, or it is hypothesized that the catalytic editing domain is made more accessible to the target DNA.
[0212] Targeted mutagenesis and DNA labeling by nickases and DNA polymerases The CRISPR systems described herein can be used in conjunction with proteins that act on nicked DNA. One such class of proteins are nick translation DNA polymerases such as E. coli DNA polymerase I or Taq DNA polymerase.
[0213] In some embodiments, a CRISPR system (e.g., a CRISPR nickase) can be fused to error-prone DNA polymerase I. This fusion protein can be targeted by an RNA guide and generate a nick at the target DNA site. When DNA polymerase then initiates DNA synthesis at the nick, downstream nucleotides are substituted, and mutagenesis of the target locus occurs because an error-prone polymerase is used. Polymerase mutants with various processing capabilities, fidelities, and incorporation error biases can be used to affect the properties of the generated mutants. For this method, called EvolvR, see, for example, Halperin et al., “CRISPR-guided DNA polymerases enable diversification of all nucleotides in a tunable window,”Nature 560,248-252(2018)(which is hereby incorporated by reference in its entirety).
[0214] In some embodiments, CRISPR nickases can be used in nick translation DNA labeling protocols. Nick translation, first described by Rigby et al. in 1977, involves incubating DNA with a DNA nicking enzyme, such as DNase I, that creates one or more nicks in the DNA molecule. Next, a nick translation DNA polymerase, such as DNA polymerase I, is used to incorporate nucleic acid residues labeled at the nicking site. A method for covalently tagging telomeric repeat sequences with a fluorescent dye using a CRISPR nickase, taking advantage of its programmability, using a variant of the classical nick translation labeling protocol, is described in detail, for example, in McCaffery et al., “High-throughput single-molecule telomere characterization,” Genome Research 27:1904-1915(2017)(which is hereby incorporated by reference in its entirety). This method enables haplotype-resolved telomere length analysis at the single-molecule level.
[0215] Tracking and Labeling of Nucleic Acids Cell processes rely on a network of molecular interactions among proteins, RNAs, and DNAs. Accurate detection of protein-DNA and protein-RNA interactions is key to understanding such processes. In vitro proximity labeling techniques use an affinity tag combined with a reporter group, such as a photoactivatable group, to label polypeptides and RNAs that are near a protein or RNA of interest in vitro. After ultraviolet irradiation, the photoactivatable group reacts with proteins and other molecules in close proximity to the tagged molecule, thereby labeling them. The labeled interacting molecules can then be recovered and identified. This RNA targeting effector protein can be used, for example, to target a probe to a selected RNA sequence. Such applications can also be applied to in vivo imaging of diseases or cell types that are difficult to culture in animal models. Methods for tracking and labeling nucleic acids are described, for example, in U.S. Patent No. 8,795,965; International Publication No. WO 2016 / 205764; and International Publication No. WO 2017 / 070605 (each of which is incorporated herein by reference in its entirety).
[0216] High-throughput screening The CRISPR systems described herein can be used in the preparation of next-generation sequencing (NGS) libraries. For example, to create a cost-effective NGS library, the CRISPR system can be used to disrupt the coding sequence of a target gene, and at the same time, clones transfected with the CRISPR enzyme can be screened by next-generation sequencing (e.g., on an Ion Torrent PGM system). For a detailed description of methods for preparing NGS libraries, see, for example, Bell et al., “A high-throughput screening "Strategy for detecting CRISPR-Cas9 induced mutations using next-generation sequencing," BMC Genomics, 15.1 (2014): 1002 (which is hereby incorporated by reference in its entirety) can be referred to.
[0217] Engineered microorganisms Microorganisms (e.g., Escherichia coli (E. coli), yeast, and microalgae) are widely used in synthetic biology. The development of synthetic biology has a wide range of utilities, including various clinical applications. For example, using the programmable CRISPR system described herein, a protein of a toxic domain can be split for targeted cell death using, for example, cancer-related RNA as a target transcript. Further, a fusion complex with an appropriate effector such as a kinase or an enzyme can affect a pathway involving protein-protein interaction in a synthetic biological system.
[0218] In some embodiments, an RNA guide sequence targeting a phage sequence can be introduced into a microorganism. Accordingly, the present disclosure also provides a method of inoculating a microorganism (e.g., a production strain) with a vaccine against phage infection.
[0219] In some embodiments, engineering a microorganism using the CRISPR systems provided herein can, for example, improve yield or fermentation efficiency. For example, engineering a microorganism such as yeast using the CRISPR systems described herein can produce biofuels or biopolymers from fermentable sugars or degrade plant-derived lignocellulose derived from agricultural waste as a fermentable sugar source. More specifically, the methods described herein can be used to modify the expression of endogenous genes required for biofuel production and / or modify endogenous genes that may interfere with biofuel synthesis. These methods of microorganism engineering are described, for example, in Verwaal et al., “CRISPR / Cpf1 enables fast and simple genome editing of Saccharomyces cerevisiae,” Yeast, 2017 Sep 8. doi:10.1002 / yea.3278; and Hlavova et al., “Improving microalgae for biotechnology - from genetics to synthetic biology,” Biotechnol. Adv., 2015 Nov 1;33:1194 - 203 (both of which are incorporated herein by reference in their entireties).
[0220] In some embodiments, the CRISPR systems described herein can be used to engineer microorganisms, such as the mesophilic cellulolytic bacterium Clostridium cellulolyticum, a model organism for bioenergy research, that are deficient in repair pathways. In some embodiments, CRISPR nickase can be used to introduce a single nick at a target locus, which may result in insertion by homologous recombination of an exogenously supplied DNA template. For details on how to use CRISPR nickase to edit repair-deficient microorganisms, see, for example, Xu et al., “Efficient Genome Editing in Clostridium cellulolyticum via CRISPR-Cas9 Nickase,” Appl Environ Microbiol 81:4423-4431(2015)(which is incorporated herein by reference in its entirety).
[0221] In some embodiments, the CRISPR systems provided herein can be used to induce cell death or dormancy in cells (e.g., microorganisms such as engineered microorganisms). Using these methods, but not limited to, mammalian cells (e.g., cancer cells, or tissue culture cells), protozoa, fungal cells, virus-infected cells, intracellular bacteria-infected cells, intracellular protozoa-infected cells, prion-infected cells, bacteria (e.g., pathogenic and non-pathogenic bacteria), protozoa, and single-celled and multicellular parasites, the dormancy or death of many cell types, including prokaryotic and eukaryotic cells, can be induced. For example, in the field of synthetic biology, it is highly desirable to have a mechanism to control engineered microorganisms (e.g., bacteria) to prevent their spread or seeding. The systems described herein can be used as a “kill switch” to regulate and / or prevent the spread or seeding of engineered microorganisms. Furthermore, in the art, alternatives to current antibiotic therapies are needed.
[0222] The systems described herein can also be used in applications where it is desirable to kill or control a particular microbial population (e.g., a bacterial population). For example, the systems described herein may include an RNA guide (e.g., a crRNA) that targets genus-specific, species-specific, or strain-specific nucleic acids (e.g., DNA) and can be delivered to cells. When complexed and bound to the target nucleic acid, the nuclease activity of the CLUST.029130 (type V-I) CRISPR-Cas effector protein disrupts essential functions within the microbe, ultimately resulting in dormancy or death. In some embodiments, the method comprises contacting a cell with a system described herein comprising a CLUST.029130 (type V-I) CRISPR-Cas effector protein or a nucleic acid encoding the effector protein and an RNA guide (e.g., a crRNA) or a nucleic acid encoding the RNA guide, wherein the spacer sequence is complementary to at least 15 nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 nucleotides or more) of the target nucleic acid.
[0223] While not wishing to be bound by any particular theory, the nuclease activity of the CLUST.029130 (type V-I) CRISPR-Cas effector protein can induce programmed cell death, cytotoxicity, apoptosis, necrosis, necroptosis, cell death, cell cycle arrest, cell anergy, reduced cell growth, or reduced cell proliferation. For example, in bacteria, DNA cleavage by the CLUST.029130 (type V-I) CRISPR-Cas effector protein can be bacteriostatic or bactericidal.
[0224] Applications in plants The CRISPR systems described herein are useful for a wide variety of applications in plants. In some embodiments, the CRISPR system can be used to engineer the genome of a plant (e.g., to improve production, to produce a product with a desired post-translational modification, or to introduce a gene for industrial production). In some embodiments, the CRISPR system can be used to introduce a desired trait into a plant (e.g., with or without heritable modification to the genome), or to regulate the expression of an endogenous gene in a plant cell or whole plant. Plants that can be edited using the CRISPR systems of the present disclosure (e.g., the Cas12i system) can be monocots or dicots, and include, without limitation, sunflower, corn, hemp, rice, sugarcane, canola, sorghum, tobacco, rye grass, barley, wheat, oats, triticale, peanut, potato, switchgrass, turf grass, soybean, alfalfa, sunflower, cotton, and Arabidopsis. The present disclosure also encompasses plants having traits created by the methods of the present disclosure and / or utilizing the CRISPR systems of the present disclosure.
[0225] In some embodiments, the CRISPR system can be used to identify, edit, and / or silence genes encoding specific proteins, such as allergen proteins (e.g., allergen proteins in peanut, soybean, lentil, pea, green bean, and lupin). Details regarding methods for identifying, editing, and / or silencing genes encoding proteins are described, for example, in Nicolaou et al., “Molecular diagnosis of peanut and legume allergy,” Curr. Opin. Allergy Clin. Immunol., 11(3):222-8 (2011), and International Publication No. WO 2016 / 205764 A1, both of which are incorporated herein by reference in their entirety.
[0226] Gene drive A gene drive is a phenomenon in which there is a favorable bias in the genetic traits of a specific gene or group of genes. The CRISPR system described herein can be used to construct a gene drive. For example, by targeting and disrupting a specific allele of a gene, the CRISPR system can be designed to cause a cell to copy a second allele and fix the sequence. Due to this copying, the first allele will be converted to the second allele, increasing the likelihood that the second allele will be inherited by offspring. For detailed methods regarding how to construct a gene drive using the CRISPR system described herein, see, for example, Hammond et al., “A CRISPR-Cas9 gene drive system targeting female reproduction in the malaria mosquito vector Anopheles gambiae,” Nat. Biotechnol., 2016 Jan;34(1):78-83 (which is hereby incorporated by reference in its entirety).
[0227] Pooled screening As described herein, pooled CRISPR screening is a powerful tool for identifying genes involved in biological mechanisms such as cell proliferation, drug resistance, and viral infection. Cells are transduced in bulk with a library of RNA-guided (gRNA) encoding vectors as described herein, and the gRNA distribution is measured before and after application of a selective challenge. Pooled CRISPR screens function well against mechanisms that affect cell survival and proliferation and can be extended to measure the activity of individual genes (e.g., by using engineered reporter cell lines). Arrayed CRISPR screens, where only one gene is targeted at a time, enable the use of RNA-seq as a readout. In some embodiments, the CRISPR systems as described herein can be used for single-cell CRISPR screening. For a detailed description of pooled CRISPR screening, see, for example, Datlinger et al., “Pooled CRISPR screening with single-cell transcriptome read-out,” Nat. Methods., 2017 Mar;14(3):297-301, which is incorporated herein by reference in its entirety.
[0228] Saturation mutagenesis (“bashing”) The CRISPR systems described herein can be used for in situ saturating mutagenesis. In some embodiments, a pooled RNA guide library can be used to perform in situ saturating mutagenesis on a particular gene or regulatory element. In such methods, the definitive minimal features and the individual vulnerabilities of those genes or regulatory elements (e.g., enhancers) can be revealed. These methods are described, for example, in Canver et al., “BCL11A enhancer dissection by Cas9-mediated in situ saturating mutagenesis,” Nature, 2015 Nov 12;527(7577):192-7, which is incorporated herein by reference in its entirety.
[0229] Therapeutic applications The CRISPR systems described herein that are active in a mammalian cell context (e.g., Cas12i2) can have a wide range of therapeutic applications. Furthermore, since each nuclease ortholog can have unique properties (e.g., size, PAM, etc.) that favor it for a particular targeting, therapeutic, or delivery modality, ortholog selection is important in assigning the nuclease that confers the greatest therapeutic benefit.
[0230] There are numerous factors that affect the suitability of gene editing as a therapy for a particular disease. In nuclease-based gene therapy, the major therapeutic editing techniques are considered to be gene disruption and gene correction. In the former, gene disruption generally occurs with events that activate the endogenous non-homologous end joining DNA repair mechanism of target cells (such as nuclease-induced targeted double-strand breaks), and the resulting indels often lead to loss-of-function mutations, which are intended to be beneficial to the patient. In the latter, gene correction utilizes nuclease activity to induce alternative DNA repair pathways (such as homologous recombination repair, or HDR) with the help of a template DNA (regardless of whether it is endogenous or exogenous, single-stranded or double-stranded). The DNA to be templated is either the endogenous correction of the mutation causing the disease or, alternatively, the insertion of a therapeutic transgene into a selected locus (generally a safe harbor locus such as AAVS1). Methods for designing exogenous donor template nucleic acids are described, for example, in International Publication No. WO 2016 / 094874 A1 (the entire content of which is hereby expressly incorporated by reference). What is required for a therapy using any of such editing modalities is an understanding of the gene regulators of a particular disease; the disease does not necessarily have to be a single-gene disease, but insights into how mutations can lead to the progression or outcome of the disease are important in providing guidance regarding the potential effectiveness of gene therapy.
[0231] Although not desired to be limited, the CRISPR systems described herein can be used to treat the following diseases, where in addition to the relevant references that help to adapt the type V-I CRISPR system to specific disease areas, specific gene targets are identified; cystic fibrosis by targeting CFTR (International Publication No. WO2015157070A2 pamphlet), Duchenne muscular dystrophy and Becker muscular dystrophy by targeting dystrophin (DMD) (International Publication No. WO2016161380A1 pamphlet), alpha-1-antitrypsin deficiency by targeting alpha-1-antitrypsin (A1AT) (International Publication No. WO2017165862A1 pamphlet), lysosomal storage disorders such as Pompe disease, also known as glycogenosis type II, by targeting acid alpha-glucosidase (GAA), myotonic dystrophy by targeting DMPK, Huntington's disease by targeting HTT, fragile X by targeting FMR1, Friedreich's ataxia by targeting frataxin, amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (FTD) by targeting C9orf72, hereditary chronic kidney disease by targeting ApoL1, cardiovascular diseases and hyperlipidemia by targeting PCSK9, APOC3, ANGPTL3, LPA (Nature 555, S23-S25 (2018)), and congenital blindness such as Leber congenital amaurosis 10 (LCA10) by targeting CEP290 (Maeder et al., Nat Med. 2019 Feb; 25(2): 229-233). The majority of the aforementioned diseases are best treated with in vivo gene editing techniques, where the cell types and tissues involved in the disease need to be edited in situ at a dosage and efficiency sufficient to produce a therapeutic benefit. Some of the challenges of in vivo delivery are described in the following section "Delivery of CRISPR Systems", but generally, due to the small gene size of type V-I CRISPR effectors, more versatile packaging into viral vectors with payload limitations, such as adeno-associated viruses, is achieved.
[0232] Ex vivo editing, where cells are removed from a patient's body, then edited, and then transplanted back into the patient, presents valuable therapeutic opportunities for gene editing technologies. The ability to manipulate cells outside the body begins with the use of techniques that are not suitable for the in vivo context, such as electroporation and nucleofection, to deliver proteins, DNA, and RNA to cells with high efficiency, and extends to the ability to determine toxicity (from off-target effects, etc.), and then further, to successfully select and expand the edited cells to produce a population that provides a therapeutic benefit. These advantages are offset by the relatively small number of cell types and populations that can be successfully recovered, processed while maintaining function, and then returned to the body. Although not desired to be limited, there are serious diseases that are suitable for ex vivo genome editing using the systems described herein. For example, sickle cell disease (SCD) as described in WO 2015 / 148863A2 pamphlet, and β-thalassemia as described in WO 2015 / 148860A1 pamphlet are both examples of diseases in which several different editing modalities in hematopoietic stem cells for disease treatment are enabled by an understanding of the pathophysiology. Both β-thalassemia and SCD can be treated by increasing the level of fetal hemoglobin by disrupting the BCL11A erythroid enhancer (as exemplified by using zinc finger nucleases according to Psatha et al. Mol Ther Methods Clin Dev. 2018 Sep 21). In addition, genetic modification methods can be used to revert the deleterious mutations in SCD and β-thalassemia. In another example, adding β-globin expressed from a safe harbor locus provides another alternative therapeutic strategy for ex vivo gene editing.
[0233] As a natural consequence of ex vivo editing of hematopoietic stem cells, immune cells can also be edited. In cancer immunotherapy, one treatment modality is to modify immune cells such as T cells to recognize and fight cancer, as described in WO 2015 / 161276 A2. To reduce costs while increasing effectiveness and ease of use, the creation of off-the-shelf allogeneic T cell therapies is attractive, and gene editing has the potential to modify surface antigens to minimize any immunological side effects (Jung et al., Mol Cell. 2018 Aug 31).
[0234] In another embodiment, the invention is used to target a virus or other pathogen at the double-stranded DNA intermediate stage of its life cycle. Specifically, targeting viruses that remain as latent infections where the initial infection persists permanently can have significant therapeutic value. In the following examples, for HSV-1 and HSV-2, WO 2015 / 153789 A1, WO 2015 / 153791 A1, and WO 2017 / 075475 A1, and for HIV, WO 2015 / 148670 A1 and WO 2016 / 183236 A1, the type V-I CRISPR system can be used for direct targeting of viral genomes (such as those associated with HSV-1, HSV-2, or HIV), or for editing host cells to reduce or eliminate receptors that enable infection and render the cells resistant to the virus (HIV).
[0235] In another aspect, the CRISPR systems described herein can be engineered to utilize enzymatically inactive Cas12i as a basis and attach protein domains thereto to achieve additional functions that can confer activities such as transcriptional activation, repression, base editing, and methylation / demethylation.
[0236] Accordingly, the present disclosure provides a CRISPR-Cas system and cells for use in the treatment or prevention of any of the diseases disclosed herein.
[0237] Delivery of the CRISPR System The CRISPR systems described herein, or components thereof, nucleic acid molecules thereof, or nucleic acid molecules encoding or providing components thereof, can be delivered by various delivery systems such as vectors, e.g., plasmids, viral delivery vectors, e.g., adeno-associated virus (AAV), lentivirus, adenovirus, and other viral vectors, or by methods such as nucleofection or electroporation of ribonucleoprotein complexes consisting of a type V-I effector and one or more cognate RNA guides. The protein and one or more RNA guides can be packaged into one or more vectors, e.g., plasmids or viral vectors. For bacterial applications, the nucleic acid encoding any of the components of the CRISPR systems described herein can be delivered to bacteria using phages. Exemplary phages include, but are not limited to, T4 phage, Mu, λ phage, T5 phage, T7 phage, T3 phage, φ29, M13, MS2, Qβ, and φX174.
[0238] In some embodiments, the vector, e.g., plasmid or viral vector, is delivered to the target tissue by, e.g., intramuscular injection, intravenous administration, transdermal administration, intranasal administration, oral administration, or mucosal administration. Such delivery may be by either a single dose or multiple doses. One of ordinary skill in the art will understand that the actual dosage delivered herein can vary widely depending on various factors such as the choice of vector, target cells, organism, tissue, general condition of the subject being treated, degree of transformation / modification required, route of administration, mode of administration, type of transformation / modification required, and the like.
[0239] In certain embodiments, the delivery is by adeno-associated virus (AAV), e.g., AAV2, AAV8, or AAV9, which is at least 1×10 5It can be administered in a single dose containing an adenovirus or adeno-associated virus of particles (also referred to as particle units, pu). In some embodiments, the dose is at least about 1×10 6 particles, at least about 1×10 7 particles, at least about 1×10 8 particles, or at least about 1×10 9 particles of adeno-associated virus. For delivery methods and doses, see, for example, WO 2016 / 205764 and US Pat. No. 8,454,972, both of which are incorporated herein by reference in their entirety. Since the genomic payload of recombinant AAV is limited, the small size of the V-I type CRISP-Cas effector proteins described herein allows for higher versatility when packaging the effector and RNA guide with appropriate control sequences (e.g., promoters) necessary for efficient and cell-type specific expression.
[0240] In some embodiments, delivery is by a recombinant adeno-associated virus (rAAV) vector. For example, in some embodiments, a modified AAV vector may be used for delivery. The modified AAV vector may be based on one or more of several capsid types including AAV1, AAV2, AAV5, AAV6, AAV8, AAV8.2, AAV9, AAV rh10, modified AAV vectors (e.g., modified AAV2, modified AAV3, modified AAV6), and pseudotype AAVs (e.g., AAV2 / 8, AAV2 / 5, and AAV2 / 6). Exemplary AAV vectors and techniques that can be used to generate rAAV particles are known in the art (e.g., Aponte-Ubillus et al. (2018) Appl. Microbiol. Biotechnol. 102(3):1045-54; Zhong et al. (2012) J. Genet. Syndr. Gene Ther. S1:008; West et al. (1987) Virology 160:38-47 (1987); Tratschin et al. (1985) Mol. Cell. Biol. 5:3251-60); U.S. Pat. Nos. 4,797,368 and 5,173,414; and International Publication Nos. WO 2015 / 054653 and WO 93 / 24641 (each of which is incorporated by reference)).
[0241] In some embodiments, delivery is by plasmid. The dosage may be an amount of plasmid sufficient to elicit a response. In some cases, a suitable amount of plasmid DNA in the plasmid composition may be from about 0.1 to about 2 mg. The plasmid generally will comprise: (i) a promoter; (ii) a sequence encoding a nucleic acid targeting CRISPR enzyme operably linked to the promoter; (iii) a selectable marker; (iv) an origin of replication; and (v) a transcription terminator downstream of and operably linked to (ii). The plasmid can also encode the RNA components of the CRISPR-Cas system, although alternatively one or more of these may be encoded by a different vector. The frequency of administration is within the purview of a medical or veterinary practitioner (e.g., a physician, veterinarian), or one of ordinary skill in the art.
[0242] In another embodiment, delivery is by liposomes or lipofectin formulations, etc., and can be prepared by methods known to those of ordinary skill in the art. Such methods are described, for example, in WO 2016205764 pamphlet and U.S. Patent Nos. 5,593,972; 5,589,466; and 5,580,859 (each of which is incorporated herein by reference in its entirety).
[0243]
[0244] A further means of introducing one or more components of this novel CRISPR system into cells is by use of a cell-penetrating peptide (CPP). In some embodiments, the cell-penetrating peptide is linked to the CRISPR enzyme. In some embodiments, the CRISPR enzyme and / or RNA guide is coupled with one or more CPPs and effectively transported into the cell interior (e.g., plant protoplasts). In some embodiments, the CRISPR enzyme and / or one or more RNA guides are encoded by one or more circular or linear DNA molecules that are coupled to one or more CPPs for cell delivery.
[0245] A CPP is a short-chain peptide of less than 35 amino acids derived from either a protein or chimeric sequence having the ability to transport biomolecules across the cell membrane in a receptor-independent manner. The CPP may be a cationic peptide, a peptide having a hydrophobic sequence, an amphiphilic peptide, a peptide having a proline-rich antimicrobial sequence, and a chimeric or bipartite peptide. Examples of CPPs include, for example, Tat (which is a transcriptional activator protein required for viral replication by human immunodeficiency virus type 1), penetratin, Kaposi fibroblast growth factor (FGF) signal peptide sequence, integrin β3 signal peptide sequence, polyarginine peptide Arg sequence, guanine-rich molecular transporter, and sweet arrow peptide. For CPPs and methods of their use, see, for example, Haellbrink et al., “Prediction of cell-penetrating peptides,” Methods Mol. Biol., 2015; 1324: 39-58; Ramakrishna et al., “Gene disruption by cell-penetrating peptide-mediated delivery of Cas9 protein and guide RNA,” Genome Res., 2014 Jun; 24(6): 1020-7; and International Publication No. WO 2016 / 205764 A1, each of which is incorporated herein by reference in its entirety.
[0246] Delivery of the type V-I CRISPR system as a ribonucleoprotein complex by electroporation or nucleofection involves pre-incubating the purified Cas12i protein with the RNA guide and electroporating (or nucleofecting) the target cells, and is another method for efficiently introducing the CRISPR system into cells for gene editing. This is particularly useful for ex vivo genome editing and the development of cell therapies, and for such methods, see Roth et al. “Reprogramming human T cell function and specificity with non-viral genome targeting,” Nature, 2018 Jul;559(7714):405-409.
[0247] Various delivery methods for the CRISPR systems described herein are also described, for example, in U.S. Patent No. 8,795,965, European Patent No. 3,009,511, International Publication No. 2016 / 205764, and International Publication No. 2017 / 070605, each of which is incorporated herein by reference in its entirety.
[0248] Kit The disclosure also encompasses kits for performing the various methods of the disclosure using the CRISPR systems described herein. One exemplary kit of the disclosure includes (a) one or more nucleic acids encoding a CRISPR-associated protein and cognate crRNA, and / or (b) a ribonucleoprotein complex of a CRISPR-associated protein and cognate crRNA. In some embodiments, the kit includes a Cas12i protein and a Cas12i guide RNA. As described above, the complex of the protein and the guide RNA has editing activities such as SSB formation, DSB formation, CRISPR interference, nucleobase modification, DNA methylation or demethylation, chromatin modification, etc. In certain embodiments, the CRISPR-associated protein is a mutant, such as a mutant with reduced endonuclease activity.
[0249] The kits of the present disclosure also optionally include one or more additional reagents, such as a reaction buffer, a wash buffer, one or more control materials (e.g., a nucleic acid encoding a substrate or a CRISPR system component). The kits of the present disclosure also optionally include instructions for performing the methods of the present disclosure using the materials provided in the kit. These instructions are provided in a physical form, such as a printed document physically packaged together with the other items of the kit, and / or in a digital form, such as a digitally issued document downloadable from a website or provided on a computer-readable medium.
Examples
[0250] The present invention is further described in the following examples, which do not limit the scope of the present invention described in the claims.
[0251] Example 1: Identification of the minimal components of the CLUST.029130 (V-I type) CRISPR-Cas system (Figures 1-3) This protein family represents large single effectors associated with CRISPR systems found in uncultured metagenomic sequences retrieved from freshwater environments (Table 3). The CLUST.029130 (V-I type) effector, designated Cas12i, includes exemplary proteins detailed in Tables 3 and 4. Exemplary direct repeat sequences of these systems are shown in Table 5.
[0252] NCBI (Benson et al. (2013) GenBank. Nucleic Acids Res. 41, D36-42; Pruitt et al. (2012) NCBI Reference Sequences (RefSeq): current status, new features and genome annotation policy. Nucleic Acids Res. 40, D130-135), NCBI whole-genome sequencing (WGS), and DOE JGI Integrated Microbial Genomes (Markowitz et al. (2012) IMG: the Integrated Microbial Genomes database and comparative analysis system. Nucleic Acids Res. 40, D115-122) by downloading and aggregating genome and metagenome sequences, a database of 293,985 putative CRISPR-Cas systems was constructed, within which the inventors identified novel nuclease systems. This pipeline engineering approach expands the search space and reduces bias for novel CRISPR effector discovery by performing minimal filtering at intermediate stages.
[0253] By comparing sequence profiles extracted from multiple alignments of readily alignable Cas12 protein groups, the classification trees shown in FIGS. 1A - 1B were created. Profile - to - profile comparisons were performed using HHsearch (Soeding et al. (2005) Protein homology detection by HMM - HMM comparison. Bioinforma. Oxf. Engl. 21, 951 - 960); the score between two profiles was normalized by the minimum of the self - scores and converted into a distance matrix on a natural logarithm scale. The UPGMA dendrogram was re - created from this distance matrix. Trees at a depth of two distance units (corresponding to pairwise HHsearch scores with e -2D = 0.02 compared to the self - score) typically reliably recover profile similarity and can serve as a guide for subtype classification (Shmakov et al., 2017).
[0254] The domain composition of Cas12i shown in FIGS. 2A and 2B indicates that the effector contains the catalytic residues of the RuvC nuclease domain. Additionally, the predicted secondary structure of the most frequent direct repeats for the V-I type locus shown in FIG. 3 indicates a stem-loop structure conserved in the crRNAs of many exemplary V-I type CRISPR-Cas systems.
[0255] [Table 3]
[0256] [Table 4-1]
[0257] [Table 4-2]
[0258] [Table 4-3]
[0259] [Table 4-4]
[0260] [Table 4-5]
[0261] [Table 4-6]
[0262] [Table 4-7]
[0263]
Table 4-8
[0264]
Table 4-9
[0265]
Table 4-10
[0266]
Table 5A
[0267]
Table 5B-1
[0268]
Table 5B-2
[0269] Example 2: In Vivo Bacterial Validation of Engineered CLUST.029130 (Type V-I) CRISPR-Cas System (Figures 4A-10B) By identifying the minimal components of the Type V-I CRISPR-Cas system, the inventors selected two systems for functional verification, one containing an effector called Cas12i1 (SEQ ID NO: 3), and the other containing an effector called Cas12i2 (SEQ ID NO: 5).
[0270] Method Gene Synthesis and Oligo Library Cloning The E. coli codon-optimized protein sequences of the CRISPR effector and accessory proteins were cloned into pET-28a(+) (EMD-Millipore) to generate effector plasmids. The Cas gene (including the terminal CDS coding sequence of 150 nt) or the non-coding sequence adjacent to the CRISPR array was synthesized into pACYC184 (New England Biolabs) (Genscript) to generate non-coding plasmids (Figure 4A). Effector mutant (e.g., D513A or A513D) plasmids were cloned by site-directed mutagenesis using the primers indicated in the Sequence Listing: the sequence changes were first introduced into PCR fragments and then reassembled into plasmids using NEBuilder HiFi DNA Assembly Master Mix or NEB Gibson Assembly Master Mix (New England Biolabs) according to the manufacturer's instructions.
[0271] For the pooled spacer library, the inventors first computationally designed an oligonucleotide library synthesis (OLS) pool (Agilent) to express the minimal CRISPR array of the "repeat-spacer-repeat" sequence. The "repeat" element was derived from the consensus direct repeat sequence found in the CRISPR array associated with the effector, and the "spacer" corresponded to approximately 8,900 sequences targeting the pACYC184 plasmid and essential genes of E. coli, or a negative control that did not target sequences. The spacer length was determined by the most frequent value of the spacer lengths found in the endogenous CRISPR array. Adjacent to the minimal CRISPR array, there were unique PCR priming sites that enabled the amplification of specific libraries from larger-scale oligo synthesis pools.
[0272] The inventors next cloned the minimal CRISPR array library into effector plasmids to create an effector plasmid library. The inventors added adjacent restriction sites, which are unique molecular identifiers, and the J23119 promoter for array expression on the oligo library using PCR (NEBNext High-Fidelity 2x PCR Master Mix), and then assembled a complete plasmid library of effectors with its targeting array using NEB Golden Gate Assembly Master Mix (New England Biolabs). This corresponded to the "input library" of the screen.
[0273] In vivo E. coli screen The inventors performed an in vivo screen using electrocompetent E. cloni EXPRESS BL21(DE3) E. coli cells (Lucigen) unless otherwise instructed. Competent cells were co-transformed with the effector plasmid and / or non-coding (Figure 4B). Cells were electroporated with the "input library" using Gene Pulser Xcell® (Bio-rad) in a 1.0 mm cuvette according to the manufacturer's protocol. Cells were plated on bioassay plates containing both chloramphenicol (Fisher) and kanamycin (Alfa Aesar), grown for 11 hours, and then the inventors confirmed sufficient library representation by estimating the approximate colony number and harvested the cells.
[0274] An "output library" was created by extracting the plasmid DNA fraction from the harvested cells using the QIAprep® Spin Miniprep Kit (Qiagen), while the harvested cells were lysed in Direct-zol® (Zymo Research) and then total RNA = 17 nt was recovered by extraction using the Direct-zol RNA Miniprep Kit (Zymo Research).
[0275] For both the input and output libraries, PCR was performed using custom primers containing barcodes and handles adjacent to the CRISPR array cassette of the effector plasmid library and compatible with Illumina sequencing chemistry to prepare a next-generation sequencing library for DNA depletion signals. This library was then normalized, pooled, and loaded onto a Nextseq 550 (Illumina) to evaluate the activity of the effector.
[0276] Bacterial screen sequencing analysis Next-generation sequencing data for the screen input and output libraries were demultiplexed using Illumina bcl2fastq. Reads in the fastq files obtained for each sample contained CRISPR array elements for the screening plasmid library. The orientation of the array was determined using the direct repeat sequences of the CRISPR array, and the corresponding targets were determined by mapping the spacer sequences to the source (pACYC184 or E. coli essential genes) or negative control sequences (GFP). For each sample, the total number of reads (r a ) for each unique array element in a given plasmid library was counted and normalized as follows: (r a + 1) / total number of reads for all library array elements. The depletion score was calculated by dividing the normalized output read count for a given array element by the normalized input read count.
[0277] To identify specific parameters that result in enzyme activity and bacterial cell death, the inventors quantified and compared the representation of individual CRISPR arrays (i.e., repeat-spacer-repeat) in the PCR products of input and output plasmid libraries using next-generation sequencing (NGS). The inventors defined the depletion fold for each CRISPR array as the normalized input read count divided by the normalized output read count (plus 1 to avoid division by zero). Arrays were considered "strongly depleted" if the depletion fold was greater than 3. When calculating the array depletion fold across biological replicates, the inventors took the maximum depletion fold value for a given CRISPR array across all experiments (i.e., strongly depleted arrays must be strongly depleted in all biological replicates). The inventors created a matrix for each spacer target that included the array depletion fold and the following features: target strand, transcript targeting, ORI targeting, target sequence motif, flanking sequence motif, and target secondary structure. The inventors investigated the extent to which different features in this matrix explained target depletion for type V-I systems, thereby obtaining a broad survey of functional parameters within a single screen.
[0278] Results Figures 5A - 5D show the positions of strongly depleted targets for Cas12i1 and Cas12i2 targeting pACYC184 and essential genes of Escherichia coli (E. coli) E.cloni®. Notably, the positions of strongly depleted targets appear to be dispersed across the entire potential target space.
[0279] The inventors have found that the dsDNA interference activities of V-I type effectors, Cas12i1 (1094aa) and Cas12i2 (1054aa), are abolished by mutations of conserved aspartic acids in the RuvC I motif (Figure 6A and Figure 6B). The RuvC-dependent dsDNA interference activity of Cas12i indicates that there are no requirements for non-coding sequences adjacent to the CRISPR array or cas genes (Figure 7A and Figure 7B), and it is shown that the minimal V-I interference module contains only the effector and crRNA (Figure 8A and Figure 8B).
[0280] Analysis of target flanking sequences corresponding to arrays strongly depleted from the in vivo screen indicates that dsDNA interference by Cas12i is PAM-dependent. Specifically, the inventors have found that both Cas12i1 and Cas12i2 show a 5’ TTN PAM preference (Figure 9A - Figure 9B and Figure 10A - Figure 10B). These results suggest that the compact Cas12i effector has autonomous PAM-dependent dsDNA interference ability.
[0281] Example 3: Characterization of the biochemical mechanism of the engineered CLUST.029130 (V-I type) CRISPR-Cas system (Figure 11A - Figure 13, Figure 15 - Figure 17B) Cas12i processes pre-crRNA in vivo To examine crRNA biogenesis for the V-I type CRISPR-Cas system, the inventors purified and sequenced small RNAs from Escherichia coli (E. coli) expressing Cas12i and a minimal CRISPR array library from a bacterial screen. Figure 11A and Figure 11B show the cumulative RNA-sequencing reads, indicating the strong consensus forms of Cas12i1 and Cas12i2 mature crRNAs, respectively, and the spacer length distribution. The most common spacer length observed was 21, with length variations from 16 nt to 22 nt.
[0282] For the type V-I CRISPR-Cas system containing Cas12i1, the mature crRNA can take the form of 5'-AUUUUUGUGCCCAUCGUUGGCAC[spacer]-3' (SEQ ID NO: 100).
[0283] For the type V-I CRISPR-Cas system containing Cas12i2, the mature crRNA can take the form of 5'-AGAAAUCCGUCUUUCAUUGACGG[spacer]-3' (SEQ ID NO: 101).
[0284] Small RNA sequencing from in vivo bacterial screens was performed by extracting total RNA from the recovered bacteria using Direct-zol RNA MiniPrep Plus with TRI Reagent (Zymo Research). After removing ribosomal RNA using the Ribo-Zero rRNA Removal Kit for Bacteria, it was subsequently cleaned up using the RNA Clean and Concentrator-5 kit. The resulting ribosomal RNA-depleted total RNA was treated with ATP-free T4 PNK for 3 hours to enrich the 3'-P ends, and then ATP was added and the reaction was incubated for an additional 1 hour to enrich the 5'-OH ends. The sample was then column-purified, incubated with RNA 5' polyphosphatase (Lucigen), column-purified again, and then prepared for next-generation sequencing using the NEBNext Multiplex Small RNA Library Prep Set for Illumina (New England Biolabs). This library was subjected to paired-end sequencing on a Nextseq 550 (Illumina), and the resulting paired-end alignments were analyzed using Geneious 11.0.2 (Biomatters).
[0285] Cas12i effector purification The effector vector was transformed into Escherichia coli (E. coli) NiCo21(DE3) (New England BioLabs) and expressed under the T7 promoter. The transformed cells were first grown overnight in 3 mL of Luria broth (Sigma) + 50 μg / mL kanamycin, and then 1 mL of the overnight culture was inoculated into 1 L of terrific broth medium (Sigma) + 50 μg / mL kanamycin. The cells were grown at 37 °C until an OD600 of 1 - 1.5 was reached, and then protein expression was induced with 0.2 mM IPTG. The culture was then grown at 20 °C for an additional 14 - 18 hours. The culture was harvested, pelleted by centrifugation, and then resuspended in 80 mL of lysis buffer (50 mM HEPES pH 7.6, 0.5 M NaCl, 10 mM imidazole, 14 mM 2-mercaptoethanol, and 5% glycerol) + protease inhibitor (Sigma). The cells were lysed with a cell disruptor (Constant System Limited), and the lysate was clarified by centrifugation at 28,000×g for 20 minutes twice at 4 °C. This lysate was loaded onto a 5 mL HisTrap FF column (GE Life Sciences) and purified by an imidazole gradient from 10 mM to 250 mM using FPLC (AKTA Pure, GE Life Sciences). Cas12i1 was purified in a low salt concentration buffer (50 mM HEPES-KOH pH 7.8, 500 mM KCl, 10 mM MgCl2, 14 mM 2-mercaptoethanol, and 5% glycerol). After purification, the fractions were run on an SDS-PAGE gel, and the fractions containing the protein of the appropriate size were pooled and concentrated using a 10 kD Amicon Ultra-15 centrifugal unit. The protein concentration was determined by the Qubit protein assay (Thermo Fisher).
[0286] Cas12i processes pre-crRNA in vitro To determine whether Cas12i1 has the ability for autonomous crRNA biogenesis, the inventors incubated the effector protein purified from Escherichia coli (E. coli) with pre-crRNA expressed from a minimal CRISPR array (repeat-spacer-repeat-spacer-repeat). The inventors observed that the purified Cas12i1 processes the pre-crRNA into fragments that match the mature crRNAs identified from in vivo small RNAseq, suggesting that Cas12i1 has the ability for autonomous pre-crRNA processing (Figure 12).
[0287] The pre-crRNA processing assay of Cas12i1 was performed at 37 °C for 30 min at a final pre-crRNA concentration of 100 nM in cleavage buffer. This reaction was carried out in a cleavage buffer optimized for Cas12i (50 mM Tris-HCl pH 8.0, 50 mM NaCl, 1 mM DTT, 10 mM MgCl2, 50 μg / ml BSA). The reaction was quenched by adding 1 μg / μL Proteinase K (Ambion) and incubated at 37 °C for 15 min. After adding 50 mM EDTA to the reaction, it was mixed with an equal volume of 2× TBE-urea sample buffer (Invitrogen) and denatured at 65 °C for 3 min. Samples were analyzed on a 15% TBE-urea gel (Invitrogen). The gel was stained with SYBR Gold nucleic acid stain (Invitrogen) for 5 min and imaged with a Gel Doc EZ (Biorad). The gel containing the labeled pre-crRNA was first imaged with an Odyssey CLx scanner (LI-COR Biosciences) and then SYBR stained.
[0288] Cas12i1 DNA manipulation using a strongly depleted array To explore the interference mechanism of Cas12i1, the inventors selected a CRISPR array sequence that was strongly depleted from an in vivo negative selection screen and generated a pre-crRNA with a DR-spacer-DR-spacer-DR sequence. The pre-crRNA was designed to target Cas12i1 to a 128-nt ssDNA and dsDNA substrate containing a target sequence complementary to the second spacer of the pre-crRNA. The inventors observed that the Cas12i1 binary complex consisting of the effector protein and the pre-crRNA cleaved 100 nM of target ssDNA up to saturation at a complex concentration of 62.5 nM (Figure 13). Further degradation from the cleaved ssDNA to short fragments or single nucleotides was observed at high complex concentrations, suggesting collateral ssDNA cleavage activated by the binding of the binary complex to the ssDNA target (Figure 13).
[0289] To explore the dsDNA interference activity of Cas12i, the inventors targeted the Cas12i1 binary complex to a target dsDNA substrate containing a 5'-end label on the non-spacer complementary strand. To comprehensively evaluate both dsDNA cleavage and nicking activities, the resulting dsDNA cleavage reaction was divided into three fractions for different analyses. The first two fractions were quenched and analyzed by denaturing or non-denaturing gel electrophoresis conditions, respectively. The third fraction was treated with 0.1 U of S1 nuclease to convert any dsDNA nicks to double-strand breaks, quenched, and analyzed by non-denaturing gel electrophoresis.
[0290] The inventors observed dose-dependent cleavage under denaturing conditions, which suggests either target nicking or dsDNA cleavage (Figure 15). Under non-denaturing conditions without S1 nuclease treatment, the inventors observed a dose-dependent increase in primary products that migrated with slightly lower electrophoretic mobility than the input dsDNA, which suggests nicked dsDNA products (Figure 16). Incubation of these products with S1 nuclease converted the shifted-up bands into smaller dsDNA products, indicating S1-mediated conversion from nicked dsDNA to double-strand breaks (Figure 16). The inventors also observed that dsDNA cleavage products were less at high concentrations and long incubation times, indicating that Cas12i1 is a dsDNA nuclease that cleaves the spacer complementary ("SC") strand and the non-spacer complementary ("NSC") strand of the target dsDNA with substantially different efficiencies (Figure 17A).
[0291] Observation of nicking activity with 5'-labeling of the spacer complementary strand of the dsDNA substrate suggested that Cas12i1 preferentially nicks the DNA strand opposite to the crRNA-target DNA hybrid. To verify this bias in DNA strand cleavage by Cas12i1, we generated dsDNA substrates labeled with the IR800 dye at either the 5'-end of the spacer complementary strand or the 5'-end of the non-spacer complementary strand. At low effector complex concentrations, we observed cleavage of only the NSC strand of the DNA duplex, while at high effector complex concentrations, cleavage of both the NSC and SC strands was observed (Figures 17A-17B). Comparing SYBR staining that labels all nucleic acid products with strand-specific labeling using the IR800 dye reveals differences in the strand product formation ratios compared to the overall cleavage product accumulation. These results suggest that an ordered series of events leads to dsDNA interference, and thus the Cas12i1 binary complex first nicks the NSC strand and then cleaves the SC strand with lower efficiency to result in dsDNA cleavage. Collectively, these findings indicate that Cas12i is an effector with the ability for autonomous pre-crRNA processing, ssDNA targeting and collateral cleavage, and dsDNA cleavage. This series of catalytic activities is quite similar to Cas12a and Cas12b, except that preferential dsDNA nicking occurs due to a significant bias towards non-spacer complementary strand cleavage.
[0292] crRNA and Substrate RNA Preparation Single-stranded DNA oligo templates for crRNA and substrate RNA were ordered from IDT. Double-stranded in vitro transcription (IVT) template DNA was generated by PCR amplification of substrate RNA and pre-crRNA templates using NEBNEXT Hifi 2× Master Mix (New England Biolabs). After annealing the T7 primer to the template, followed by extension using DNA polymerase I, large (Klenow) fragment (New England Biolabs), a double-stranded DNA template for mature cr-RNA was generated. Annealing was performed by incubating at 95 °C for 5 min, followed by incubation to 4 °C with a decrease of -5 °C / min. In vitro transcription was performed by incubating the dsDNA template with T7 RNA polymerase at 37 °C for 3 h using the HiScribe T7 Quick High Yield RNA Kit (New England Biolabs). After incubation, the IVT samples were treated with Turbo DNase (registered trademark) (Thermo Scientific) and then purified using the RNA Clean & Concentrator Kit (Zymo Research). Mature cr-RNA generated from IVT was treated with calf intestinal alkaline phosphatase (Thermo Fisher) or RNA 5'-polyphosphatase (Lucigen) at 37 °C for 2 h to generate 5'-hydroxyl or 5'-monophosphate, respectively, followed by cleanup with the RNA Clean & Concentrator Kit (Zymo Research). Concentrations were measured by Nanodrop 2000 (Thermo Fisher).
[0293] The pre-crRNA sequences used for biochemical characterization of Cas12i are included in Table 6. The oligonucleotide templates and primers for crRNA preparation are included in Table 9.
[0294] Preparation of IR-800-labeled substrate RNA and DNA The RNA substrate from IVT was treated with calf intestinal alkaline phosphatase (Thermo Fisher) at 37 °C for 30 minutes to convert 5'-triphosphate to a 5'-terminal hydroxyl group and purified using an RNA Clean & Concentrator kit (Zymo Research). Thiol-terminal groups were added to the 5'-terminal hydroxyl groups of the DNA and RNA substrates using a 5’ EndTag Labeling Kit (Vector Labs), and the substrates were then labeled with IRDye 800CW maleimide (LI-COR Biosciences). The substrates were purified using a DNA Clean & Concentrator kit or an RNA Clean & Concentrator kit (Zymo Research). Unlabeled (non-spacer-complementary) ssDNA strands were labeled, annealed to primers, and then extended at 25 °C for 15 minutes by DNA polymerase I, large (Klenow) fragment (New England Biolabs) to generate labeled dsDNA substrates. These substrates were purified using a DNA Clean & Concentrator kit (Zymo Research). Concentrations were measured using a Nanodrop 2000 (Thermo Fisher).
[0295] The RNA and DNA substrate sequences used for the biochemical characterization of Cas12i are included in Tables 7 and 8.
[0296] Target cleavage assay with Cas12i ssDNA: Optimized cleavage buffer (50 mM Tris-HCl pH 8.0, 50 mM The Cas12i target cleavage assay with ssDNA was performed in (NaCl, 1 mM DTT, 10 mM MgCl2, 50 μg / ml BSA). A binary complex was formed by incubating Cas12i:pre-crRNA at a 1:2 molar concentration ratio at 37 °C for 10 minutes and then transferring it to ice. All further complex dilutions were performed on ice while maintaining a constant protein:RNA ratio. This complex was added to a 100 nM IR800-labeled substrate and incubated at 37 °C for 30 minutes. The reaction was treated with an RNase cocktail and Proteinase K and analyzed as described above.
[0297] dsDNA: The dsDNA target cleavage assay was set up at 37 °C for 1 hour in the optimized cleavage buffer. A binary complex was formed as described above and added to a 100 nM dsDNA substrate. First, the reaction was treated with an RNase cocktail by incubation at 37 °C for 15 minutes. Next, this was treated with Proteinase K by incubation at 37 °C for 15 minutes. To detect dsDNA cleavage products, this reaction was analyzed on a 15% TBE-urea gel as described above. To detect the nicking activity of Cas12i, the reaction was SPRI purified after Proteinase K treatment and divided into three fractions. One fraction was analyzed on a 15% TBE-urea gel as described above. Another fraction was mixed with 5× hi-density TBE sample buffer and analyzed on a non-denaturing 4–20% TBE gel to detect nicked dsDNA products. The last fraction was incubated with 0.01 U / μL of S1 nuclease (Thermo Scientific) at 50 °C for 1 hour to convert the nick to a double-strand break, followed by mixing with 5× hi-density TBE sample buffer and analyzing on a non-denaturing 4–20% TBE gel. All gels were imaged with an Odyssey CLx scanner, then stained with SYBR for 5 minutes and imaged with a Gel Doc imager.
[0298] To identify the nicked strand, dsDNA was prepared by labeling either the target strand (complementary to crRNA) or the non-target strand (non-spacer complementary, same sequence as crRNA). The cleavage reaction was performed as described. Next, the labeled strand was annealed with the corresponding primer and extended at 25 °C for 15 minutes by DNA polymerase I, large (Klenow) fragment (New England Biolabs). Then, the dsDNA substrate was purified using SPRI purification.
[0299]
Table 6
[0300]
Table 7
[0301]
Table 8
[0302]
Table 9
[0303] Example 4: In Vitro Pooled Screening for the Rapid Evaluation of CRISPR-Cas Systems (Figures 20 - 25) As described herein, in vitro pool-type screening serves as a high-throughput method efficient for performing biochemical evaluations. Briefly, the inventors begin with an in vitro reconstitution of the CRISPR-Cas system (Figure 20). In one embodiment, a dsDNA template containing a T7-RNA polymerase promoter that drives the expression of one or more effector proteins is used, and effector proteins are produced using in vitro transcription and translation reagents that create the proteins for the reaction. In another embodiment, the minimal CRISPR array and tracrRNA include a T7 promoter sequence added in either the top-strand or bottom-strand transcription direction using PCR to examine all possible RNA orientations. As shown in Figure 20, the apo-type contains only the effector, the binary-type contains the effector protein and the T7 transcript minimal CRISPR array, and the binary + tracrRNA-type adds any T7-transcribed tracrRNA element to the complex for incubation.
[0304] In one embodiment, the nucleotide strand cleavage activity of the CRISPR-Cas system is the major biochemical activity being assayed. Figure 21 shows one form of ssDNA and dsDNA substrates where the target sequence is flanked by 6-degenerate bases on both sides, creating a pool of possible PAM sequences that can gate ssDNA and dsDNA cleavage activity. In addition to the PAM sequence, the substrates include 5' and 3' reference marks designed to facilitate a downstream next-generation sequencing library preparation protocol that selectively enriches the substrate ssDNA or dsDNA and to facilitate mapping of the cleavage products. In one embodiment, the dsDNA substrate is generated by second-strand synthesis in the 5' to 3' direction using short DNA primers and DNA polymerase I. Similar reactions can be carried out using a pool of different targets in the minimal CRISPR array, as well as libraries of different ssDNA and dsDNA sequences.
[0305] The CRISPR-Cas cleavage reaction is carried out by mixing a pre-formed Apo / dual / dual-tracrRNA complex with either a targeting substrate or a non-targeting substrate and incubating. While other methods such as gel electrophoresis are possible, an embodiment useful for cleavage capture with the highest sensitivity and base pair resolution is next-generation sequencing of ssDNA or dsDNA substrates after incubation with the effector complex. Figure 22 is a schematic diagram illustrating the preparation of a library for enrichment of ssDNA substrates. By annealing a primer to a well-defined sequence within a reference mark, second-strand synthesis and end repair occur, generating dsDNA fragments corresponding to both the cut ssDNA and the uncut ssDNA. Subsequently, this newly formed dsDNA molecule serves as a substrate for adapter ligation, and then selective PCR is performed using one primer (I5 / P5) complementary to the ligation adapter and another (I7 / P7) complementary to the 3' reference of the original ssDNA substrate. This ultimately produces a sequencing library containing both full-length and cleaved and fragmented ssDNA products as shown in Figure 24A. The dsDNA readout NGS library preparation starts without the need for primer annealing and second-strand synthesis, so end repair and subsequent adapter ligation can be carried out directly. Figure 23 illustrates a general overview of library preparation that labels both cleaved / fragmented and uncleaved fragments, similar to ssDNA preparation. Notably, either end of the dsDNA cleavage fragment can be enriched based on the selection of PCR primers. In one embodiment shown in Figure 24A, the dsDNA-operated next-generation sequencing library for readout can be prepared with a first primer complementary to the handle ligated to the 5' end of the full-length or cleaved substrate (and containing the I5 / P5 sequence) and a second primer complementary to the 3' reference sequence of the substrate (and containing the I7 / P7 sequence).In one embodiment shown in FIG. 24B, a DNA manipulation next-generation sequencing library for readout can be prepared with a first primer complementary to the 5' reference sequence of the substrate (and containing the I5 / P5 sequence) and a second primer complementary to the handle ligated to the 3' end of the full-length substrate or the cleaved substrate (and containing the I7 / P7 sequence). As shown in FIGS. 25A-25B respectively, the target length and the substrate length can be extracted from the NGS reads obtained from the RNA / ssDNA / dsDNA manipulation experiment. Using the extracted target length and substrate length, the presence of RNA / ssDNA / dsDNA nicking or cleavage can be examined.
[0306] Example 5: Characterization of the dsDNA cleavage activity of the V-I1 type CRISPR-Cas system (FIGS. 26-32) By computationally identifying the minimal components of the V-I type CRISPR-Cas system, the inventors examined the double-stranded DNA (dsDNA) cleavage activity from the V-I1 type system containing the effector Cas12i1.
[0307] As shown in FIGS. 26A-26B, Cas12i1 by in vitro transcription translation (IVTT) expression in a complex with crRNA by top strand expression targeting dsDNA resulted in a truncated target length population absent in the apo (effector only) control. A library prepared using a 5' ligation adapter and selected with respect to the 3' reference (as shown in FIG. 24A) showed cleavage products absent in the Apo control at position +24 within the target sequence. This result indicates nicking of either the non-target dsDNA strand or both strands of the dsDNA between nucleotides +24 and +25 relative to the PAM. Since the target length analysis shows a peak at +24, truncation of the target between nucleotides +24 and +25 is indicated (FIG. 27A). Since this truncated target sequence population coincides with the substrate length, cleavage of the non-target dsDNA strand between nucleotides +24 and +25 of the target sequence is indicated (FIG. 28A).
[0308] The library prepared using the 3’ ligation adapter and selected with respect to the 5’ reference (as shown in Fig. 24B) showed cleavage products that were not present in the control at the -9 position ( +19 considering the 28nt target) within the target sequence. This result indicates nicking of either the target dsDNA strand or both strands of the dsDNA between the +19 and +20 nucleotides relative to the PAM. Since the target length analysis shows a peak from the PAM to the -9 nucleotide (28nt full-length target), truncation of the target between nucleotides +19 and +20 is indicated (Fig. 27B). Since this truncated target sequence population matches the substrate length, cleavage of the target dsDNA strand between nucleotides +19 and +20 of the target sequence is indicated (Fig. 28B).
[0309] Sequence motif analysis of substrates showing non-target strand cleavage between the +24 / +25 nucleotides relative to the PAM revealed the 5’ TTN PAM motif to the left of the target sequence for Cas12i1 (Fig. 29). No PAM sequence requirements were observed for the right side of the Cas12i1 target. In summary, in vitro screening of Cas12i1 showed predominant nicking between the +24 / +25 nucleotides of the non-target strand relative to the TTN PAM, and a significant proportion of these products were converted to double-strand cleavage with a 5nt 3’ overhang by cleavage of the target strand between the +19 / +20 nucleotides relative to the PAM (Fig. 30).
[0310] Targeting of Cas12i1 in complexes with non-target crRNAs by top strand expression showed that Cas12i1 cleavage specificity is conferred by the crRNA spacer, as no relative dsDNA manipulation occurred (Figs. 31A - 31B). Cas12i1 has been shown to have no cleavage activity in the presence of crRNA by bottom strand expression targeting the dsDNA substrate, indicating that a crRNA in the top strand orientation is required for the formation of an active Cas12i1 complex (Figs. 32A - 32B).
[0311] Example 6: Characterization of dsDNA cleavage activity of the V-I2 type CRISPR-Cas system (Figs. 33-39) By computationally identifying the minimal components of the V-I type CRISPR-Cas system, the inventors examined the double-stranded DNA (dsDNA) cleavage activity from the V-I2 type system containing the effector Cas12i2.
[0312] As shown in FIGS. 33A-33B, Cas12i2 by in vitro transcription / translation (IVTT) expression in a complex with crRNA having a top strand expression targeting dsDNA resulted in a population of truncated target lengths not present in the apo (effector only) control. A library prepared using a 5' ligation adapter and selected with respect to the 3' reference (as illustrated in Fig. 24A) showed cleavage products not present in the Apo control at position +24 within the target sequence. This result indicates nicking of either the non-target dsDNA strand or both strands of the dsDNA between nucleotides +24 and +25 relative to the PAM. Since the target length analysis shows a peak at +24, truncation of the target between nucleotides +24 and +25 is indicated (Fig. 34A). Since this population of truncated target sequences matches the substrate length, cleavage of the non-target dsDNA strand between nucleotides +24 and +25 of the target sequence is indicated (Fig. 35A).
[0313] A library prepared using a 3' ligation adapter and selected with respect to the 5' reference (as illustrated in Fig. 33B) showed cleavage products not present in the Apo control at position -7 (i.e., +24 considering a 31 nt target) within the target sequence. This result indicates nicking of either the target dsDNA strand or both strands of the dsDNA between nucleotides +24 and +25 relative to the PAM. Since the target length analysis shows a peak at -7 nucleotides from the PAM (28 nt full-length target), truncation of the target between nucleotides +24 and +25 is indicated (Fig. 34B). Since this population of truncated target sequences matches the substrate length, cleavage of the target dsDNA strand between nucleotides +24 and +25 of the target sequence is indicated (Fig. 35B).
[0314] From the sequence motif analysis of a substrate showing non-target strand cleavage between +24 / +25 nucleotides relative to the PAM for Cas12i2, the 5'TTN PAM motif to the left of the target sequence was revealed (Figure 36). No PAM sequence requirement was recognized for the right side of the Cas12i2 target. In summary, in vitro screening of Cas12i2 showed dominant nicking between +24 / +25 nucleotides of the non-target strand relative to the TTN PAM, and a significant proportion of these products were converted to blunt-ended double-strand breaks by cleavage of the target strand between +24 / +25 nucleotides relative to the PAM (Figure 37).
[0315] Since no relative dsDNA manipulation occurred upon targeting Cas12i2 in a complex with non-target crRNA by top strand expression, it is shown that Cas12i2 cleavage specificity is conferred by the crRNA spacer (Figures 38A - 38B). Cas12i2 has been shown to have no cleavage activity in the presence of crRNA by bottom strand expression targeting the dsDNA substrate, indicating that a crRNA in the top strand orientation is required for the formation of an active Cas12i2 complex (Figures 39A - 39B).
[0316] Example 7: The CLUST.029130 (V-I type) CRISPR Cas system can be used for gene silencing in vitro To rapidly verify the activity of the novel CRISPR-Cas system, an in vitro gene silencing assay (Figures 18A and 18B) that mimics in vivo gene silencing activity was developed. This assay can simultaneously evaluate various activation mechanisms and functional parameters in a non-biased manner outside the natural cellular environment.
[0317] First, the reconstituted IVTT (in vitro transcription and translation) system was supplemented with E. coli RNA polymerase core enzyme, enabling gene expression (protein synthesis) to occur not only from the T7 promoter but also from any E. coli promoter as long as the corresponding E. coli sigma factor is present.
[0318] Second, to facilitate rapid high-throughput experimental methods, linear DNA templates generated from PCR reactions were used directly. These linear DNA templates included those encoding type V-I effectors, RNA guides, and E. coli sigma factor 28. Incubating these DNA templates with the reconstituted IVTT reagents results in co-expression of the type V-I effector and the RNA guide and formation of RNP (ribonucleoprotein complex). E. coli sigma factor 28 was also expressed for subsequent GFP and RFP expression as described below.
[0319] Third, as a target substrate, linear or plasmid DNA encoding GFP expressed from the sigma factor 28 promoter was included in the above incubation reaction to allow newly synthesized RNP to reach the target substrate immediately. As an internal control, non-target linear DNA encoding RFP expressed from the sigma factor 28 promoter was also included. RNA polymerase core enzyme alone does not recognize the sigma factor 28 promoter until sufficient sigma factor 28 protein is synthesized. This delay in GFP and RFP expression allows newly synthesized RNP to interfere with the GFP target substrate, resulting in a potential decrease in GFP expression and depletion of GFP fluorescence. On the other hand, since RFP expression was not affected negatively, this serves as an internal control for protein synthesis and fluorescence measurement.
[0320] Specific important advantages of the in vitro gene silencing assay described herein include the following: (1) modularity - The reconstituted IVTT is a synthetic system composed of individually purified components, enabling the custom design of assays for various controls and activities. Since each component of the CRISPR-Cas system is encoded on a separate linear DNA template, rapid assays of different effector, effector mutant, and RNA guide combinations are possible; (2) complexity - Since this assay includes all the components essential for RNA transcription and protein synthesis, it is possible to test various interference mechanisms, such as DNA and RNA cleavage and transcription-dependent interference, in a single one-pot reaction. The kinetic fluorescence readout of this assay provides significantly more data points compared to endpoint activity assays; (3) sensitivity - Since this assay combines effector and RNA guide synthesis with substrate interference, newly synthesized RNPs (ribonucleoprotein complexes of effector proteins and RNA guides) can immediately interact with substrates within the same reaction. Without separate purification steps, it is potentially possible to generate signals with a small amount of RNP. Furthermore, due to the combined GFP transcription and translation that can generate over 100 GFP proteins per DNA template, the interference of GFP expression is amplified. (4) efficiency - This assay is designed to be highly compatible with high-throughput platforms. Due to its modularity, all components of the assay can be added in 96, 384, and 1536 well formats using commonly available liquid handling equipment, and fluorescence can be measured using commonly available plate fluorometers. (5) relevance - This assay tests the ability of CRISPR-Cas effector proteins to interfere with gene expression during transcription and translation in a system where they are engineered in vitro outside their native cellular environment. Highly active CRISPR-Cas effectors measured by this gene silencing assay may also be highly efficient for gene editing in mammalian cells.
[0321] This assay is used to measure the gene silencing effect exerted by the Cas12i effector complex as shown herein when targeting GFP encoded by plasmid DNA. Multiple V-I type RNA guides are designed, one containing a spacer sequence complementary to the template strand of the GFP sequence and the other containing a spacer sequence complementary to the coding strand of the GFP sequence. Next, the degree of gene silencing by the Cas12i1 effector protein was compared to the mutants Cas12i1 D647A, Cas12i1 E894A, and Cas12i1 D948A.
[0322] Figure 19A shows the depletion fold of each of the four tested Cas12i effectors when complexed with an RNA guide complementary to the template strand. In this case, the non-target strand that is preferentially nicked is the coding strand. Cas12i1 shows GFP expression with a depletion fold of approximately 2 after 400 minutes, while each of the three mutants shows a lower degree of depletion.
[0323] Figure 19B shows the depletion fold of each of the four tested Cas12i1 effectors when complexed with an RNA guide complementary to the coding strand. In this case, the non-target strand that is preferentially nicked is the template strand. In this configuration, the ability of RNA polymerase to produce a functional RNA transcript appears to be significantly impaired by Cas12i1, with depletion exceeding 4-fold in the case of Cas12i. The gene silencing ability of the three mutants appears to be significantly reduced.
[0324] In summary, the data shown in FIGS. 19A and 19B indicate that this assay is effective for detecting the gene silencing activity of Cas12i1 when using an RNA guide that targets both the coding strand and the template strand. Since there is significantly higher depletion when targeting the coding strand rather than the template strand, it is suggested that Cas12i1 interferes with GFP expression by preferentially nicking the non-target strand. All three Cas12i1 mutants substitute the putative catalytic residues (aspartic acid (D) and glutamic acid (E)) with alanine (A). The decrease in the silencing activity of these Cas12i1 mutants further supports that DNA strand cleavage, rather than just binding, underlies the gene silencing mechanism by Cas12i1.
[0325] Example 8: The CLUST.029130 (V-I type) CRISPR-Cas system can be used for the specific detection of nucleic acid species together with a fluorescent reporter Due to the nuclease activity of the Cas12i protein (i.e., the non-specific collateral DNase activity activated by a target ssDNA substrate complementary to the crRNA spacer), such effectors are promising candidates for use in the detection of nucleic acid species. Some of these methods have already been described (see, for example, East-Seletsky et al. “Two distinct RNase activities of CRISPR-C2c2 enable guide-RNA processing and RNA detection,” Nature. 2016 Oct 13;538(7624):270-273), Gootenberg et al. (2017), Chen et al. 2018, and Gootenberg et al. (2018) “Multiplexed and portable nucleic acid detection platform with Cas13, Cas12a, and Csm6” Science 15 Feb 2018: eaaq0179) described the general RNA detection principle using Cas13a (East-Seletsky et al. (2016)), which was supplemented by amplification and further optimization of the Cas13a enzyme to increase detection sensitivity (Gootenberg et al. (2017)), and more recently, by including additional RNA targets, orthologous and paralogous enzymes, and the Csm6 activator to achieve multiplexed nucleic acid detection with increased detection sensitivity (Gootenberg et al. (2018)). Adding Cas12i to this toolkit provides an additional channel of orthogonal activity for nucleic acid detection.
[0326] Considering that the dye-labeled collateral DNA was efficiently cleaved at low target ssDNA concentrations and that the background nuclease activity was limited to non-targeting substrates, the in vitro biochemical activity of Cas12i1 suggests that it may be promising for application to sensitive nucleic acid detection (Figure 14). To adapt Cas12i1 for sensitive nucleic acid detection applications, several steps are required, including, but not limited to, optimizing the substrate for sensitive readout of collateral activity and determining the per-base mismatch tolerance between the spacer and the target substrate.
[0327] Identification of substrates optimal for nucleic acid detection may be informed by performing next-generation sequencing (NGS) on cleavage products of Cas12i collateral activity against both DNA substrates. It may be necessary to titrate the enzyme concentration or adjust the incubation time to generate cleavage fragments that are still large enough to be prepared into next-generation sequencing libraries. NGS data reveals the preference of enzyme cleavage sites and adjacent bases. It has been demonstrated that individual effectors within the Cas13a and b families have different dinucleotide base preferences for RNA cleavage, resulting in significantly different cleavage sizes and signal-to-noise ratios (Gootenberg et al. (2018)). Thus, collateral NGS data enables better insight into the preferences for Cas12i. A separate experimental approach for identifying the dinucleotide preferences of Cas12i collateral cleavage is to generate collateral DNA substrates with degenerate Ns in consecutive positions such that they have a broader sequence space compared to defined sequences. If library preparation and analysis of NGS data proceed similarly, the base preferences for cleavage can be identified. To confirm the preferences, synthetic short DNA can be introduced into the cleavage reaction with collateral substrates containing 5'- and 3'-terminal fluorophore / quencher pairs to evaluate the signal-to-noise ratio. Further optimization can be performed with respect to the length of the collateral DNA substrate to determine whether Cas12i1 has a length preference.
[0328] Although a preferred substrate has been identified, another important parameter to determine is the mismatch tolerance of the Cas12i system, as this has implications for guide design that affects the enzyme's ability to distinguish single base pair mismatches. Mismatch tolerance may be determined by designing a panel of targets carrying different positions and types of mismatches (e.g., insertions / deletions, single base pair mismatches, adjacent double mismatches, distant double mismatches, triple mismatches, and more). Mismatch tolerance may be measured by assessing the amount of collateral DNA cleavage for targets containing varying amounts of mismatches. As an example, the collateral DNA substrate may be a short ssDNA probe containing a fluorophore and a quencher on opposite sides. For reactions containing a Cas12i effector, an RNA guide, and a target substrate containing different numbers of mismatches, insertions, and deletions in the target sequence, when activation of the Cas12i system by targeting the modified target DNA sequence is successful, collateral cleavage of the fluorescent probe will occur. Consequently, for the resulting fluorescence measurements representing the cleaved collateral substrate, the background can be subtracted using a negative control sample and normalized to the signal from a perfectly matched target to estimate the effect of target modifications on the collateral cleavage efficiency by Cas12i. Using the obtained map of mismatch, insertion, and deletion tolerances by the Cas12i enzyme relative to the target length from the PAM, optimal RNA guides can be designed to distinguish between different DNA sequences or genotypes for specific detection or discrimination between different nucleic acid species. With fluorescence quantitative cleavage readouts and preferred collateral substrates, the position and type of mismatches at which the enzyme exhibits the highest sensitivity can be determined by comparing the fluorescence activity to a perfectly matched sequence.
[0329] This optimization process can be further applied to other Cas12i orthologs to generate other systems that may have different properties. For example, the dinucleotide preference for orthogonality of collateral cleavage may help generate separate detection channels.
[0330] Example 9. Highly specific dsDNA manipulation can be achieved using the CLUST.029130 (type V-I) CRISPR Cas system in a paired nicking manner. The CLUST.029130 effector Cas12i has the ability to manipulate dsDNA through nicking of the non-target strand (Figures 15, 16, 17A - 17B). A catalytically inactivated Cas12i may also be fused to the FokI nuclease domain to create a fusion protein with the ability to bind and nick dsDNA. Some of these methods have been previously described. Ran et al. (2013) “Double Nicking by RNA-Guided CRISPR Cas9 for Enhanced Genome Editing Specificity” Science 29 Aug 2013 describes the general principles and optimization of double nicking using Cas9; Guilinger et al. (2014) “Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification” Science 25 Apr 2014 describes the principle of double nicking using the FokI-dCas9 fusion.
[0331] Using a pair of Cas12i nickases, highly specific dsDNA manipulation as follows can be achieved. A first Cas12i complex having a crRNA targeting one strand of the dsDNA target region and a second Cas12i-crRNA complex targeting the reverse strand of the dsDNA are introduced together to effect a dsDNA cleavage reaction. By targeting the Cas12i complexes to different dsDNA strands, the first and second Cas12i complexes cleave the opposing dsDNA strands, resulting in a double-strand break.
[0332] To optimize the formation efficiency of dsDNA double-strand breaks by double nicking, select crRNA spacer sequence pairs whose predicted nuclease cleavage sites are separated by different lengths. Cleavage of the top and bottom strands of dsDNA by different Cas12i nickases with different target substitutions results in sequence overhangs of different lengths, and thus the formation efficiency of double-strand breaks will be different. Nickase targets can be selected in a specific orientation to generate double-strand breaks with either 3' or 5' overhangs, or blunt ends (overhang length is 0).
[0333] For nicking applications where Cas12i1 and Cas12i2-WT enzymes contain a 5’TTN PAM, the orientation of the nickase target with PAM “out” (PAM is outside the paired target) results in a 5’ overhang, while the pair of nickase targets with PAM inside the target pair results in a 3’ overhang. In some examples, the 3’ and 5’ overhangs range from 1 to 200 nt. In some examples, the 3’ and 5’ overhangs are 20 to 100 nt.
[0334] Autonomous pre-crRNA processing facilitates the delivery of Cas12i for double nicking applications, as two separate genomic loci can be targeted from a single crRNA transcript (Figure 12). In that regard, a CRISPR array containing two spacer sequences that target Cas12i to nick both opposing strands of dsDNA can be expressed from a single viral vector or plasmid. Cas12i and the CRISPR array can also be delivered with separate plasmids or viral vectors. Next, the Cas12i protein processes the CRISPR array into two cognate crRNAs, which leads to the formation of a nicking complex. Viral vectors can include phage or adeno-associated viruses for delivery to bacterial or mammalian cells, respectively.
[0335] In addition to viral or plasmid delivery methods, the paired nicking complex can be directly delivered using nanoparticles or other direct protein delivery methods such that complexes containing both paired crRNA elements are co-delivered. Further, the protein can be delivered by a viral vector or directly to the cell, followed by direct delivery of a CRISPR array containing two paired spacers for double nicking. In some examples, for direct delivery of RNA, the RNA may be conjugated to at least one sugar moiety such as N-acetylgalactosamine (GalNAc), particularly a triantennary type GalNAc.
[0336] Example 10: Adaptation of the CLUST.029130 (V-I type) CRISPR Cas system effector to eukaryotic and mammalian activity To develop a CLUST.029130 (V-I type) CRISPR Cas system for eukaryotic applications, a construct encoding the protein effector was first codon-optimized for expression in mammalian cells and optionally a specific localization tag was added to either or both the N-terminus or C-terminus of the effector protein. Such localization tags can include sequences such as a nuclear localization signal (NLS) sequence that localizes the effector to the nucleus for genome DNA modification. These sequences are described in the "Functional Mutations" section above. Some examples of engineered non-naturally occurring nucleotide sequences encoding a mammalian codon-optimized Cas12i effector containing a localization tag are provided in Table 10. Other accessory proteins such as fluorescent proteins may be further added. Addition of a robust "superfolding" protein, such as a superfolder green fluorescent protein (GFP), has been demonstrated to increase the activity of the CRISPR enzyme in mammalian cells when added to the effector (Abudayyeh et al. (2017) Nature 550(7675):280-4, and Cox et al. (2017) Science 358(6366):1019-27).
[0337] Next, the codon-optimized sequences encoding Cas12i, the added accessory proteins, and the localization signals were cloned into a eukaryotic expression vector together with an appropriate 5' Kozak eukaryotic translation initiation sequence, a eukaryotic promoter, and a polyadenylation signal. In mammalian expression vectors, these promoters can include, for example, common promoters such as CMV, EF1a, EFS, CAG, SV40, etc., as well as cell-type specific RNA polymerase II promoters such as Syn and CamKIIa for neuronal expression, and thyroxine-binding globulin (TBG) for hepatocyte expression. Similarly, useful polyadenylation signals include, but are not limited to, SV40, hGH, and BGH. To increase the expression of such constructs, additional transcript stabilization or transcript nuclear export elements such as WPRE can be used. For the expression of pre-crRNA or mature crRNA, an RNA polymerase III promoter such as H1 or U6 can be used.
[0338] Depending on the application and packaging method, the eukaryotic expression vector may be a lentiviral plasmid backbone, an adeno-associated virus (AAV) plasmid backbone, or a similar plasmid backbone that can be used in the production of recombinant viral vectors. Notably, since the CLUST.029130 (V-I type) CRISPR Cas effector protein, such as the Cas12i protein, is small in size, it is ideally suitable for packaging into a single adeno-associated virus particle together with its crRNA and appropriate control sequences; due to the 4.7 kb packaging size limit of AAV, the use of large effectors may not be possible, especially when a large cell-type specific promoter is used for expression control.
[0339] After adapting the arrays, delivery vectors, and methods for use in eukaryotes and mammals, different Cas12i constructs as described herein were characterized with respect to performance. Initial characterization was performed by lipofection of DNA constructs expressing the minimal components of the Cas12i system adapted for use in eukaryotes as described above. In one embodiment, the Cas12i effector is codon-optimized for mammals and a nucleoplasmin nuclear localization sequence (npNLS) is added to the C-terminus of the protein. Expression of the effector is driven by the elongation factor 1α short (EFS) promoter and terminated using the bGH poly(A) signal (Table 10). For the expression of the cognate RNA guide for the Cas12i system, a double-stranded linear PCR product containing the U6 promoter was used, referring to (Ran et al. “Genome engineering using the CRISPR-Cas9 system,” Nat Protoc. 2013 Nov;8(11):2281-2308). This approach is well-suited for testing a large number of sgRNAs by subjecting them to plasmid cloning and sequence verification. (Figure 40) The effector plasmid and the U6-guide PCR fragment were co-transfected into 293T cells in a 24-well plate format as 400 ng of the effector plasmid and 30 ng of the U6-guide PCR product at a plasmid-to-PCR product molar concentration ratio of approximately 1:2. The resulting gene editing events were evaluated using next-generation sequencing of the targeted PCR amplicons around the target site (Hsu et al., “DNA targeting specificity of RNA-guided Cas9 nucleases,” Nat Biotechnol. 2013 Sep;31(9):827-32).
[0340] From the initial evaluation of Cas12i2, 13% indel activity was obtained at the VEGFA locus at the target site containing the TTC PAM. The inventors tested different RNA guide designs as described in FIG. 41, where the most potent indel efficiency was achieved using pre-crRNA, and the indel rate decreased as the spacer length decreased. Examination of the indels created by Cas12i2 revealed that the dominant position of the indels centered around +20 relative to the PAM sequence.
[0341] Multiplexing of type V-I effectors is achieved using pre-crRNA processing with effector capabilities, where multiple targets with different sequences can be programmed onto a single RNA guide. Thus, multiple genes or DNA targets can be manipulated simultaneously for therapeutic applications. One embodiment of the RNA guide design is a pre-crRNA expressed from a CRISPR array consisting of target sequences interspersed with unprocessed DR sequences repeated to enable simultaneous targeting of one, two, or more loci by endogenous pre-crRNA processing of the effector.
[0342] In addition to testing various construct configurations and accessory arrays against individual targets, the use of a pooled library-based approach determines 1) the targeting dependence of any given Cas12i protein in mammalian cells, and 2) the effects of mismatch positions and combinations along the length of the targeting crRNA. Briefly, a pooled library contains plasmids that express target DNA containing different flanking sequences and mismatches with one or more guides used in the screening experiment, and thus depletion of sequences from the library occurs upon successful target recognition and cleavage. Furthermore, the specificity of the CLUST.029130 (V-I type) CRISPR-Cas system can be evaluated using targeted indel sequencing or unbiased genome-wide cleavage assays (Hsu et al. (2013), Tsai et al. “GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases.” Nat Biotechnol. 2015 Feb;33(2):187-197, Kim et al. “Digenome-seq: genome-wide profiling of CRISPR-Cas9 off-target effects in human cells,” Nat Methods. 2015 Mar;12(3):237-43, Tsai et al., “CIRCLE-seq: a highly sensitive in vitro screen for genome-wide CRISPR-Cas9 nuclease off-targets,” Nat Methods. 2017 Jun;14(6):607-614).
[0343] In addition, mutations are created to expand the functional scope of the Cas12i protein. In some embodiments, conserved residues of the RuvC domain are mutated to alanine (such as the D647A mutation for Cas12i1 and the D599A mutation for Cas12i2), enabling the production of catalytically inactive Cas12i proteins. Catalytically inactive Cas12i versions (referred to as dCas12i) retain their programmable DNA-binding activity but can no longer cleave target or collateral ssDNA or dsDNA. Direct uses of dCas12i include immunoprecipitation and transcriptional repression. By adding other domains to the dCas12i protein, additional functionality is provided.
[0344] Activities of these domains include, but are not limited to, DNA base modification (e.g., ecTAD and its evolved forms, APOBEC), DNA methylation (m 6 A methyltransferase and demethylase), localization factors (KDEL retention sequence, mitochondrial targeting signal), transcriptional modification factors (e.g., KRAB, VP64). In addition, domains such as light-gated control (cryptochrome) and chemically inducible components (FKBP-FRB chemically inducible dimerization) can be added to provide further control.
[0345] Optimization of the activity of such fusion proteins requires a systematic approach to compare linkers that connect dCas12i to the added domain. These linkers include, but are not limited to, flexible glycine-serine (GS) linkers of various combinations and lengths, non-flexible linkers such as the α-helix-forming EAAAK sequence, XTEN linkers (Schellenberger V, et al. Nat. Biotechnol. 2009;27:1186-1190), and various combinations thereof (see Table 11). Next, various designs are assayed in parallel with the same crRNA target complex and functional readout to determine which produces the desired properties.
[0346] To adapt Cas12i for use in targeted DNA base modification (see, e.g., Gaudelli et al. (2017) “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Science 25 Oct 2017), the inventors begin with a combination of the Cas12i ortholog and NLS that yielded the highest endogenous mammalian DNA cleavage activity and mutate conserved residues of the RuvC domain to create a catalytically inactive enzyme (dCas12i). Next, a linker is used between the dCas12i-NLS and the base editing domain to create a fusion protein. Initially, this domain will consist of the ecTadA(wt) / ecTadA*(7.10) heterodimer (hereinafter referred to as the dCas12i-TadA heterodimer) that was previously engineered for hyperactivity and modification of dsDNA A·T dinucleotides to G·C (Table 11). Considering the possible structural differences between the small Cas12i and the Cas9 effectors characterized to date, different linker designs and lengths may result in an optimal design of the base editing fusion protein.
[0347] To assess the activity of base editors derived from dCas12i, transiently transfect HEK 293T cells with a dCas12i-TadA heterodimer construct, a plasmid expressing crRNA, and optionally a reporter plasmid if targeting a reporter rather than an endogenous locus. Harvest the cells 48 hours after transient transfection, extract the DNA, and prepare it for next-generation sequencing. Analysis of the base composition of the locus of samples containing targeting crRNA compared to a negative control non-targeting crRNA provides information regarding editing efficiency, and analysis of broader changes in the transcriptome provides information regarding off-target activity.
[0348] One particular advantage of developing a DNA base editing system using Cas12i is that its small size, being smaller than existing Cas9 and Cas12a effectors, enables easier packaging into AAV together with its crRNA and regulatory elements of the dCas12i-TadA heterodimer without the need to truncate the protein. This all-in-one AAV vector enables higher in vivo base editing efficiency in tissues, which is particularly relevant as a path towards therapeutic applications of Cas12i.
[0349] In addition to editing using Cas12i and RNA guides, additional template DNA sequences can be co-delivered by a vector such as an AAV viral vector or as linear single-stranded or double-stranded DNA fragments. For insertion of template DNA by homologous recombination repair (HDR), a template sequence is designed that contains the payload sequence to be inserted at the target locus and flanking sequences homologous to the endogenous sequences adjacent to the desired insertion site. In some examples, for insertion of short DNA payloads (e.g., less than 1 kb in length), the flanking homologous sequences may be short (e.g., in the range of 15 - 200 nt in length). In other cases, for insertion of long DNA payloads (e.g., 1 kb or more in length), longer homologous flanking sequences are required to promote efficient HDR (e.g., longer than 200 nt in length). Cleavage of the target genomic locus for HDR between sequences homologous in the template DNA flanking regions can significantly increase HDR frequency. Cas12i cleavage events that promote HDR include, but are not limited to, dsDNA cleavage, double nicking, and single-strand nicking activities.
[0350] The dsDNA fragment can contain overhang sequences complementary to the overhangs generated by double nicking using Cas12i. Pairing of the insert with the double nicking overhangs and subsequent ligation by the endogenous DNA repair machinery results in seamless insertion of the template DNA at the double nicking site.
[0351]
Table 10-1
[0352]
Table 10-2
[0353]
Table 11
[0354] These results suggest that members of the compact V-I type CRISPR family can be engineered for activity in eukaryotic cells, specifically for genome editing in mammalian cells. Mammalian functional V-I type effectors enable the development of additional technologies based on further engineering on the basis of DNA binding.
[0355] Example 11. Using a V-I type CRISPR-Cas system, genotype gating control of genome replication, viral propagation, plasmid propagation, cell death, or cell dormancy can be provided When the V-I type CRISPR-Cas effector protein and crRNA hybridize to a specific ssDNA or dsDNA target, nicking or cleavage of the substrate occurs. The fact that such activity depends on the presence of a specific DNA target in the cell is valuable for realizing targeting of specific genomic materials or cell populations based on specific genotypes. There are numerous applications for such control of genome replication, cell death, or cell dormancy in both eukaryotic, prokaryotic, and viral / plasmid settings.
[0356] For prokaryote, virus, and plasmid applications, a type V-I CRISPR-Cas system (e.g., including a type V-I effector and an RNA guide) can be delivered (e.g., in vitro or in vivo) to genotype-specifically arrest the genomic replication of a particular prokaryotic population (e.g., a bacterial population) and / or induce cell death or dormancy. For example, a type V-I CRISPR-Cas system can include one or more RNA guides that specifically target a particular virus, plasmid, or prokaryotic genus, species, or strain. As shown in FIGS. 5A-5D, cleavage, nicking, or interference with the E. coli genome or plasmid DNA conferring antibiotic resistance in E. coli by the type V-I system results in specific depletion of E. coli containing those sequences. Specific targeting of a virus, plasmid, or prokaryote has many therapeutic benefits as it can be used to induce the death or dormancy of unwanted bacteria (e.g., pathogenic bacteria such as Clostridium difficile). In addition, the type V-I systems provided herein may be used to target prokaryotic cells having a specific genotype. Among the microbial diversity colonizing humans, only a very small number of bacterial strains can induce pathogenicity. Furthermore, even among pathogenic strains such as Clostridium difficile, not all members of the bacterial population are always present in a disease-causing state that is active. Thus, targeting of the type V-I system based on the genotype of a virus, plasmid, or prokaryotic cell allows for specific control over which genome or cell population is targeted without disrupting the entire microbiome.
[0357] In addition, the engineered bacterial strain can be readily engineered to generate a genetic kill switch that restricts the growth, colony formation, and / or excretion of the engineered bacterial strain using genetic circuits or environmentally controlled expression elements. For example, the expression of type V-I effectors and specific crRNAs can be controlled using promoters derived from the regulatory regions of genes encoding proteins that are expressed in response to external stimuli, such as cold shock protein (PcspA), heat shock protein (Hsp), and chemically inducible systems (Tet, Lac, AraC). Controlled expression of one or more elements of the type V-I system allows for the expression of a fully functional system only upon exposure to environmental stimuli, which results in the genotype-specific DNA interference activity of the system and, consequently, induction of cell death or dormancy. A kill switch comprising a Cas12i effector, such as those described herein, may be advantageous because it does not depend on the relative protein expression rate that can be affected by leaky expression from a promoter (e.g., an environmentally stimulated-dependent promoter) compared to conventional kill switch designs, such as toxin / antitoxin systems (e.g., the CcdB / CcdA type II toxin / antitoxin system), and thus allows for more precise control of the kill switch.
[0358] Other embodiments Although the invention has been described with a detailed description, it should be understood that the foregoing description is illustrative only and is not intended to limit the scope of the invention as defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
Claims
**Claim 1**: An engineered, non-naturally occurring clustered regularly interspaced short palindromic repeat (CRISPR)-associated (Cas) system, comprising: (a) an RNA guide or a nucleic acid encoding said RNA guide, said RNA guide comprising a direct repeat sequence and a spacer sequence, said RNA guide or the nucleic acid encoding said RNA guide; and (b) a CRISPR-Cas effector protein or a nucleic acid encoding said CRISPR-Cas effector protein, said CRISPR-Cas effector protein comprising an amino acid sequence having at least 95% identity with SEQ ID NO: 5 and a nuclear localization sequence (NLS) having an amino acid sequence at least 90% identical to KRPAATKKAGQAKKKKK (SEQ ID NO: 301), said CRISPR-Cas effector protein or the nucleic acid encoding said CRISPR-Cas effector protein wherein said CRISPR-Cas effector protein binds to said RNA guide and said spacer sequence binds to a target nucleic acid. A system. **Claim 2**: The system according to claim 1, wherein said system does not contain tracrRNA. **Claim 3**: The CRISPR-Cas effector protein is (a) an amino acid sequence X1SHX4DX6X7 (SEQ ID NO: 200) (wherein X1 is S or T, X4 is Q or L, X6 is P or S, and X7 is F or L) containing a RuvC domain, (b) an amino acid sequence X1XDXN X6X7XXXX11 (SEQ ID NO: 201) (wherein X1 is A, G, or S, X is any amino acid, X6 is Q or I, X7 is T, S, or V, and X11 is T or A) containing a RuvC domain; and (c) an amino acid sequence X1X2X3E (SEQ ID NO: 210) (wherein X1 is C, F, I, L, M, P, V, W, or Y, X2 is C, F, I, L, M, P, R, V, W, or Y, and X3 is C, F, G, I, L, M, P, V, W, or Y) containing a RuvC domain The system according to claim 1, comprising one or more of the above. **Claim 4**: The direct repeat sequence is (a) 5'-CCGUCNNNNNNNUGACGG-3' (SEQ ID NO: 202) adjacent to the spacer array (where N is any nucleobase); '(SEQ ID NO: 202) (wherein N is any nucleobase); (b) 5'-GUGCCNNNNNNUGGCAC-3' (SEQ ID NO: 203) adjacent to the spacer array (where N is any nucleobase); (c) 5'-GUGUC N5-6UGACA X1-3' (SEQ ID NO: 204) adjacent to the spacer array (where N5-6 is any continuous sequence of 5 or 6 nucleobases, and X1 is C or T or U); (d) 5'-UCX3UX5X6X7UUGACGG-3' (SEQ ID NO: 205) adjacent to the spacer array (where X3 is C, T, or U, X5 is A, T, or U, X6 is A, C, or G, and X7 is A or G); or (e) 5'-CCX3X4X5CX7UUGGCAC-3' (SEQ ID NO: 206) adjacent to the spacer array (where X3 is C, T, or U, X4 is A, T, or U, X5 is C, T, or U, and X7 is A or G) The system according to claim 1, comprising any one of the above.
5. The system according to claim 1 or claim 4, wherein the direct repeat sequence comprises an RNA transcript of a nucleotide sequence having at least 95% sequence identity with SEQ ID NO: 9 or SEQ ID NO:
10.
6. The system according to claim 1 or claim 4, wherein the direct repeat sequence comprises an RNA transcript of the nucleotide sequence set forth in SEQ ID NO: 9 or SEQ ID NO:
10.
7. The system according to claim 1, wherein the CRISPR-Cas effector protein comprises the amino acid sequence set forth in SEQ ID NO:
5.
8. The direct repeat sequence comprises a stem-loop structure proximal to the 3' end of the direct repeat sequence, and the stem-loop structure (a) a first stem nucleotide strand 5 nucleotides in length; (b) a second stem nucleotide strand 5 nucleotides in length (the first and second stem nucleotide strands are bonded to each other); and (c) a loop nucleotide strand disposed between the first and second stem nucleotide strands and comprising 6, 7, or 8 nucleotides The system according to claim 1, comprising the above.
9. The system according to claim 1, wherein the spacer array comprises nucleotides of length 15 to 47.
10. The system according to claim 9, wherein the spacer array comprises nucleotides of length 24 to 38.
11. The system according to claim 1, wherein the target nucleic acid comprises a sequence complementary to the nucleotide sequence in the spacer array.
12. The system according to claim 1, wherein the CRISPR-Cas effector protein recognizes a protospacer adjacent motif (PAM) sequence, and the PAM sequence comprises a nucleotide sequence described as 5'-TTN-3' (wherein N is any nucleotide).
13. The system according to claim 12, wherein the PAM sequence comprises a nucleotide sequence described as 5'-TTY-3' (wherein Y is C or T) or 5'-TTH-3' (wherein H is A or C or T).
14. The system according to claim 1, wherein the CRISPR-Cas effector protein cleaves the target nucleic acid.
15. The system according to claim 1, wherein the CRISPR-Cas effector protein further comprises a second NLS or at least one nuclear export signal (NES).
16. The system according to claim 1, wherein the CRISPR-Cas effector protein further comprises a peptide tag, a fluorescent protein, a base editing domain, a DNA methylation domain, a histone residue modification domain, a localization factor, a transcriptional modification factor, a light gate control factor, a chemical inducibility factor, or a chromatin visualization factor.
17. The system according to claim 1, wherein the nucleic acid encoding the CRISPR-Cas effector protein is codon-optimized for expression in cells.
18. The system according to claim 1, wherein the nucleic acid encoding the CRISPR-Cas effector protein is operably linked to a promoter.
19. The system according to claim 1, wherein the nucleic acid encoding the CRISPR-Cas effector protein is in a vector.
20. The system according to claim 19, wherein the vector comprises a retroviral vector, a lentiviral vector, a phage vector, an adenoviral vector, an adeno-associated vector, or a herpes simplex vector.
21. The system according to claim 1, which is present in a delivery system comprising nanoparticles, liposomes, exosomes, microvesicles, or a gene gun.
22. An ex vivo or in vitro cell comprising the system according to claim 1.
23. The ex vivo or in vitro cell according to claim 22, wherein the cell is a prokaryotic cell or a eukaryotic cell.
24. The ex vivo or in vitro cell according to claim 23, wherein the cell is a mammalian cell or a plant cell.
25. The ex vivo or in vitro cell according to claim 24, wherein the cell is a human cell.
26. The system according to claim 1 for use in a method of binding the system to a target nucleic acid in a cell, the method comprising: (a) providing the system; and (b) delivering the system to the cell wherein the cell comprises the target nucleic acid, the CRISPR-Cas effector protein binds to the RNA guide, and the spacer sequence binds to the target nucleic acid.
27. The system according to claim 26, wherein the target nucleic acid is single-stranded DNA or double-stranded DNA.
28. The system according to claim 26, wherein binding the system to the target nucleic acid results in modification of the target nucleic acid.
29. The system according to claim 26, wherein binding the system to the target nucleic acid results in cleavage of the target nucleic acid.
30. The system according to claim 29, wherein cleavage of the target nucleic acid results in the formation of an insertion or deletion in the target nucleic acid.
Citation Information
Patent Citations
JPP7457653B
Crystal structure of crispr CPF1
WO2017127807A1
Methods for identifying class 2 crispr-CAS systems
WO2018035250A1
Novel crispr enzymes and systems
WO2018035388A1