Type vi crispr ortholog and system
Patent Information
- Application Number
- JP2025087622
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-04-12
- Filing Date
- 2025-05-27
- Publication Date
- 2025-11-04
AI Technical Summary
There is a need for novel genome and transcriptome engineering technologies that are inexpensive, easy to set up, scalable, and suitable for targeting multiple locations within eukaryotic genomes and transcriptomes, as existing methods like designer zinc fingers and TALEs are limited in versatility and efficiency.
Development of CRISPR-Cas13a systems, particularly using LwaCas13a, for precise RNA targeting and knockdown in mammalian cells, with catalytically inactive versions (dCas13a) for imaging and binding, and active versions for transcript destruction, leveraging the RNA-guided RNase activity for targeted RNA manipulation.
The CRISPR-Cas13a systems enable robust and specific RNA targeting and manipulation in eukaryotic systems, facilitating applications such as transcript knockdown, imaging, and targeted destruction, with potential for therapeutic applications like inducing programmed cell death in undesirable cells.
Abstract
Description
[Technical Field]
[0001] Related Applications and Incorporation by Reference This application claims priority to U.S. Provisional Patent Application Nos. 62 / 351,662 and 62 / 351,803, filed June 17, 2016, U.S. Provisional Patent Application No. 62 / 376,377, filed August 17, 2016, U.S. Provisional Patent Application No. 62 / 410,366, filed October 19, 2016, U.S. Provisional Patent Application No. 62 / 432,240, filed December 9, 2016, U.S. Provisional Patent Application No. 62 / 471,792, filed March 15, 2017, and U.S. Provisional Patent Application No. 62 / 484,786, filed April 12, 2017.
[0002] Reference is made to U.S. Provisional Patent Application No. 62 / 471,710, filed March 15, 2017, entitled "Novel Cas13B Orthologues CRISPR Enzymes and Systems," Attorney Docket No. BI-10157 VP 47627.04.2149. Further reference is made to U.S. Provisional Patent Application No. 62 / 432,553, filed December 9, 2016; U.S. Provisional Patent Application No. 62 / 456,645, filed February 8, 2017; and U.S. Provisional Patent Application No. 62 / 471,930, filed March 15, 2017 (entitled "CRISPR Effector System Based Diagnostics"), Attorney Docket No. BI-10121 BROD 0842P), and a prospectively assigned U.S. Provisional Patent Application No. BI-10121 BROD 0843P, filed April 12, 2017 (entitled "CRISPR Effector System Based Diagnostics"), Attorney Docket No. BI-10121 BROD 0843P.
[0003] All documents cited or referenced in the documents cited herein, together with any manufacturer's instructions, manuals, product specifications, and product sheets for any products mentioned herein or in any document incorporated by reference herein, are hereby incorporated by reference and may be used in the practice of the present invention. More specifically, all documents referenced are incorporated by reference to the same extent as if each individual document were specifically and individually indicated to be incorporated by reference.
[0004] Federally Sponsored Research Statement This invention was made with federal support under Grant Nos. MH100706 and MH110049 awarded by the National Institutes of Health. The federal government has certain rights in this invention.
[0005] The present invention relates generally to systems, methods, and compositions used to control gene expression, including gene transcript perturbation or nucleic acid editing, including sequence targeting, which may use vector systems related to clustered regularly interspaced short palindromic repeats (CRISPR) and its components. [Background technology]
[0006] Recent advances in genome sequencing technologies and analytical methods have rapidly improved our ability to catalog and map genetic factors associated with various biological functions and diseases. Precise genome targeting technologies are needed to enable systematic reverse engineering of causal genetic variations by selectively perturbing individual genetic elements and advance synthetic biology, biotechnological, and medical applications. While genome editing techniques such as designer zinc fingers, transcription activator-like effectors (TALEs), or homing meganucleases are available to generate targeted genome perturbations, there remains a need for novel genome and transcriptome engineering technologies that utilize novel strategies and molecular mechanisms and are inexpensive, easy to set up, scalable, and suitable for targeting multiple locations within eukaryotic genomes and transcriptomes. This will serve as a major resource for novel applications in genome engineering and biotechnology.
[0007] CRISPR-Cas systems in bacterial and archaeal adaptive immunity exhibit a high degree of diversity in terms of protein composition and genomic locus organization. CRISPR-Cas loci contain over 50 gene families, with no strictly universal genes, suggesting rapid evolution and a high degree of locus organization diversity. To date, a multidisciplinary approach has comprehensively identified approximately 395 cas gene profiles for 93 Cas proteins. The classification includes signature gene profiles and locus organization signatures. A novel classification of CRISPR-Cas systems has been proposed, broadly dividing these systems into two classes: Class 1, which has multisubunit effector complexes, and Class 2, which has single-subunit effector modules, as exemplified by the Cas9 protein. Novel effector proteins associated with Class 2 CRISPR-Cas systems can be developed as powerful genome engineering tools, and the prediction, engineering, and optimization of putative novel effector proteins are crucial.
[0008] The CRISPR-Cas adaptive immune system defends microorganisms against foreign genetic elements through DNA or RNA-DNA interference. Recently, the class 2 type VI single-component CRISPR-Cas effector C2c2 (Shmakov et al. (2015) “Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems”; Molecular Cell 60:1–13; doi:http: / / dx.doi.org / 10.1016 / j.molcel.2015.10.008) has been characterized as an RNA-guided RNase (Abudayyeh et al. (2016) Science, [Epub ahead of print], June 2; “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector”; doi:10.1126 / science.aaf5573). C2c2 (e.g., from Leptotrichia shahii) has been demonstrated to confer robust interference against RNA phage infection. Through in vitro biochemical analysis and in vivo assays, it was shown that C2c2 can be programmed to cleave ssRNA targets with protospacers flanked by 3'H (non-G) PAMs. Cleavage is mediated by catalytic residues in the two conserved HEPN domains of C2c2, and mutations in these domains result in catalytically inactive RNA-binding proteins. Guided by a single guide, C2c2 can be reprogrammed to deplete specific mRNAs in vivo. LshC2c2 was shown to be capable of targeting specific sites of interest and, once primed with the cognate target RNA, to perform nonspecific RNase activity.These results expand our understanding of the CRISPR-Cas system and suggest that C2c2 may be harnessed to develop a broad array of RNA-targeting tools.
[0009] C2c2 is now known as Cas13a. It will be understood that the term "C2c2" is used synonymously with "Cas13a" herein.
[0010] Citation or identification of any document in this application is not an admission that such document is available as prior art to the present invention. Summary of the Invention [Means for solving the problem]
[0011] There is an urgent need for alternative and robust systems and techniques for targeting nucleic acids or polynucleotides (e.g., DNA or RNA, or any hybrid or derivative thereof) with a wide variety of applications, particularly in eukaryotic systems, and more particularly in mammalian systems. The present invention addresses this need and provides related advantages. The addition of the present novel RNA targeting system to the repertoire of genome, transcriptome, and epigenome targeting technologies can transform the study and perturbation or editing of specific target sites through direct detection, analysis, and manipulation, particularly in eukaryotic systems, more particularly in mammalian systems (including cells, organs, tissues, or organisms), and plant systems. To effectively utilize the present RNA targeting system for RNA targeting without adverse effects, it is critical to understand the engineering and optimization aspects of these RNA targeting tools.
[0012] The CRISPR-Cas13 family was discovered by computational mining of bacterial genomes for CRISPR system signatures (Shmakov, S. et al., "Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems." Mol Cell 60, 385-397, doi:10.1016 / j.molcel.2015.10.008(2015)), and is comprised of the single-effector RNA-guided RNase Cas13a / C2c2 (Abudayyeh, O.O. et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector." Science 353, aaf5573, doi:10.1126 / science.aaf5573(2016)) and later the single-effector RNA-guided RNase Cas13b (Shmakov, S. et al. "Diversity and evolution of class 2 CRISPR-Cas systems". Nat Rev Microbiol 15, 169-182, doi:10.1038 / nrmicro.2016.184(2017); Smargon, A. A. et al. "Cas13b Is a Type VI-B CRISPR-Associated RNA-Guided RNase Differentially Regulated by Accessory Proteins Csx27 and Csx28" Csx28). Mol Cell 65, 618-630 e617, doi:10.1016 / j.molcel.2016.12.023(2017)).The class 2 type VI effector protein C2c2, also known as Cas13a, is an RNA-guided RNase that can be efficiently programmed to degrade ssRNA. In contrast to the catalytic mechanisms of other known RNases found in CRISPR-Cas systems, C2c2 (Cas13a) achieves RNA cleavage through conserved basic residues within its two HEPN domains. Mutation of the HEPN domain, such as alanine substitution, at any of the four predicted HEPN domain catalytic residues converts C2c2 into an inactive programmable RNA-binding protein (dC2c2, similar to dCas9).
[0013] The RNA-guided RNase Cas13, due to its programmability and specificity, may be an ideal platform for transcriptome engineering. We develop Cas13a for use as a mammalian transcript knockdown and binding tool. Leptotrichia shahii Cas13a (LshCas13a) was capable of robust RNA cleavage and binding in a catalytically inactive version using programmable crRNA, with cleavage dependent on a motif immediately 3' adjacent to it, known as the protospacer adjacent site (PFS) with identity H (non-guanine) (Abudayyeh, OO et al. "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector." Science 353, aaf5573, doi:10.1126 / science.aaf5573 (2016)). Once the RNA is cleaved, activated LshCas13a engages in "collateral activity," where non-target RNAs are cleaved by its constitutive RNase activity (Abudayyeh, OO et al. "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector." Science 353, aaf5573, doi:10.1126 / science.aaf5573 (2016)).This crRNA-programmed collateral activity prevents the spread of infection by enabling bacterial programmed cell death in vivo (Abudayyeh, OO et al. "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector." Science 353, aaf5573, doi:10.1126 / science.aaf5573(2016)) and has been applied to the specific detection of nucleic acids in vitro (Abudayyeh, OO et al. "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector." Science 353, aaf5573, doi:10.1126 / science.aaf5573(2016); East-Seletsky, A. et al. "Two distinct RNase activities of CRISPR-C2c2 enable guide-RNA processing and RNA detection". Nature 538, 270-273, doi:10.1038 / nature19802(2016)). Collateral activity has recently been utilized in a highly sensitive and specific nucleic acid detection platform called SHERLOCK, which is useful for many clinical diagnostic applications (Gootenberg, J.S. et al. "Nucleic acid detection with CRISPR-Cas13a / C2c2". Science 356, 438-442 (2017)).
[0014] After screening Cas13a orthologs bacteriologically and subsequent biochemical characterization, we selected the optimal ortholog for RNA endonuclease activity, Cas13a (LwaCas13a) from Leptotrichia wadeii. LwaCas13a can be stably expressed in mammalian cells and retargeted to effectively knockdown both reporter and endogenous transcripts in cells, achieving a high level of targeting specificity compared to RNAi without observable collateral activity. Furthermore, we demonstrate that catalytically inactive LwaCas13a (dCas13a) can programmably bind to RNA transcripts in vivo and be used to image transcripts in cells. By engineering a dCas13a-based negative feedback imaging system, we can track the formation of stress granules in live cells.
[0015] The ability of dC2c2 (dCas13a) to bind to designated sequences can be used in several embodiments of the present invention to (i) deliver effector modules to specific transcripts to modulate their function or translation (potentially useful for large-scale screening, construction of synthetic regulatory circuits, and other purposes); (ii) fluorescently tag specific RNAs to visualize their transport and / or localization; (iii) alter RNA localization with domains that have affinity for specific subcellular compartments; and (iv) capture specific transcripts (either by directly pulling down dC2c2 or by using dC2c2 to localize biotin ligase activity to specific transcripts) to enrich for nearby molecular partners, including RNAs and proteins.
[0016] Active C2c2 should also have many applications. Some embodiments of the present invention involve targeting specific transcripts for destruction, similar to RFP herein. In addition, once primed by its cognate target, C2c2 can cleave other (non-complementary) RNA molecules in vitro and inhibit cell growth in vivo. Biologically, this promiscuous RNase activity may reflect a defense mechanism based on programmed cell death / dormancy (PCD / D) of the type VI CRISPR-Cas system. Thus, in some embodiments of the present invention, C2c2 could potentially be used to induce PCD or dormancy in specific cells—e.g., cancer cells expressing particular transcripts, certain classes of neurons, cells infected with specific pathogens, or other abnormal cells or cells whose presence is otherwise undesirable.
[0017] The present invention provides a method for modifying a nucleic acid sequence associated with or at a target locus of interest, particularly a eukaryotic cell, tissue, organ, or organism, more particularly a mammalian cell, tissue, organ, or organism, comprising delivering to the locus a non-naturally occurring or engineered composition comprising a Type VI CRISPR-Cas locus effector protein and one or more nucleic acid components, wherein the effector protein forms a complex with the one or more nucleic acid components, and upon binding of the complex to the locus of interest, the effector protein induces modification of the sequence associated with or at the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break. In a preferred embodiment, the sequence associated with or at the target locus of interest comprises RNA, and the effector protein is encoded by a Type VI CRISPR-Cas locus. The complex may be formed in vitro or ex vivo and introduced into a cell or contacted with RNA; or may be formed in vivo.
[0018] The terms Cas enzyme, CRISPR enzyme, CRISPR protein, Cas protein and CRISPR Cas are generally used interchangeably and will be understood to refer by analogy to the novel CRISPR effector proteins described further herein unless otherwise clear, such as by specific reference to Cas9. The CRISPR effector proteins described herein are preferably C2c2 effector proteins.
[0019] The present invention provides a method for targeting (e.g., modifying) a sequence associated with or at a target locus of interest, comprising delivering a non-naturally occurring or engineered composition comprising a C2c2 locus effector protein (which may be catalytically active or catalytically inactive) and one or more nucleic acid components to the sequence associated with or at the locus, wherein the C2c2 effector protein forms a complex with the one or more nucleic acid components, and upon binding of the complex to the locus of interest, the effector protein induces modification of the sequence associated with or at the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break. In a preferred embodiment, the C2c2 effector protein forms a complex with one nucleic acid component, advantageously an engineered or non-naturally occurring nucleic acid component. The complex may be formed in vitro or ex vivo and introduced into a cell or contacted with RNA; or may be formed in vivo. Induction of modification of a sequence associated with or at a target locus of interest can be under C2c2 effector protein-nucleic acid guidance. In a preferred embodiment, one nucleic acid component is a CRISPR RNA (crRNA). In a preferred embodiment, one nucleic acid component is a mature crRNA or guide RNA, wherein the mature crRNA or guide RNA comprises a spacer sequence (or guide sequence) and a direct repeat sequence or a derivative thereof. In a preferred embodiment, the spacer sequence or a derivative thereof comprises a seed sequence, wherein the seed sequence is critical for recognition and / or hybridization with a sequence at a target locus.
[0020] Aspects of the present invention relate to C2c2 effector protein complexes having one or more non-naturally occurring, engineered, modified, or optimized nucleic acid components. In a preferred embodiment, the nucleic acid component of the complex can include a guide sequence linked to a direct repeat sequence, wherein the direct repeat sequence comprises one or more stem-loops or optimized secondary structures. In a specific embodiment, the direct repeat has a minimum length of 16 nt, e.g., at least 28 nt, and a single stem-loop. In a further embodiment, the direct repeat is longer than 16 nt, preferably longer than 17 nt, e.g., at least 28 nt, and has two or more stem-loops or optimized secondary structures. In a specific embodiment, the direct repeat is 25 nt or longer, e.g., 26 nt, 27 nt, 28 nt, or longer, and has one or more stem-loop structures. In a preferred embodiment, the direct repeat can be modified to include one or more protein-binding RNA aptamers. In a preferred embodiment, the direct repeat can be modified to include one or more protein-binding RNA aptamers. In a preferred embodiment, one or more aptamers may be included as part of an optimized secondary structure. Such aptamers may be capable of binding to a bacteriophage coat protein. The bacteriophage coat protein may be selected from the group including Qβ, F2, GA, fr, JP501, MS2, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s, and PRR1. In a preferred embodiment, the bacteriophage coat protein is MS2. The present invention also provides nucleic acid components of the complex that are 30 or more, 40 or more, or 50 or more nucleotides in length.
[0021] The present invention provides a cell comprising a Type VI effector protein and / or a complex with a guide and / or its target nucleic acid, in certain embodiments, the cell is a eukaryotic cell, including but not limited to a yeast cell, a plant cell, a mammalian cell, an animal cell, or a human cell.
[0022] The present invention also provides a method for modifying a target locus of interest, particularly a eukaryotic cell, tissue, organ, or organism, more particularly a mammalian cell, tissue, organ, or organism, comprising delivering to the locus a non-naturally occurring or engineered composition comprising a C2c2 locus effector protein and one or more nucleic acid components, wherein the C2c2 effector protein forms a complex with the one or more nucleic acid components, and upon binding of the complex to the locus of interest, the effector protein induces modification of the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break. The complex may be formed in vitro or ex vivo and introduced into a cell or contacted with RNA; or may be formed in vivo.
[0023] In such methods, the target locus of interest may be comprised within an RNA molecule. Alternatively, the target locus of interest may be comprised within a DNA molecule, and in certain embodiments, within a transcribed DNA molecule. In such methods, the target locus of interest may be comprised within an in vitro nucleic acid molecule.
[0024] In such methods, the target locus of interest may be contained in a nucleic acid molecule in a cell, particularly a eukaryotic cell, such as a mammalian cell or a plant cell. The mammalian cell may be a non-human primate, bovine, porcine, rodent, or murine cell. The cell may be a non-mammalian eukaryotic cell, such as a poultry, fish, or shrimp cell. The plant cell may be a crop plant, such as cassava, corn, sorghum, wheat, or rice. The plant cell may also be an algae, tree, or vegetable. The modification introduced into the cell by the present invention may be such that the cell and its progeny are altered for improved production of a biological product, such as an antibody, starch, alcohol, or other desired cellular product. The modification introduced into the cell by the present invention may be such that the cell and its progeny contain a change that alters the biological product produced.
[0025] The mammalian cell may be a non-human mammalian cell, such as a primate, bovine, ovine, porcine, canine, rodent, or Leporidae cell, such as a monkey, cow, sheep, pig, dog, rabbit, rat, or mouse cell. The cell may also be a non-mammalian eukaryotic cell, such as a poultry avian (e.g., chicken), vertebrate fish (e.g., salmon), or crustacean (e.g., oyster, clam, lobster, shrimp) cell. The cell may also be a plant cell. The plant cell may be from a monocotyledonous or dicotyledonous plant or from a crop or cereal plant, such as cassava, corn, sorghum, soybean, wheat, oat, or rice. The plant cell may also be that of an algae, a tree or productive plant, a fruit or vegetable (e.g., a citrus tree, such as an orange, grapefruit or lemon tree; a peach or nectarine tree; an apple or pear tree; a nut tree, such as an almond, walnut or pistachio tree; a Solanaceae plant; a Brassica plant; a Lactuca plant; a Spinacia plant; a Capsicum plant; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.).
[0026] The present invention provides a method of modifying a target locus of interest, the method comprising delivering to the locus a non-naturally occurring or engineered composition comprising a Type VI CRISPR-Cas locus effector protein and one or more nucleic acid components, wherein the effector protein forms a complex with the one or more nucleic acid components, and upon binding of the complex to the locus of interest, the effector protein induces modification of the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break.
[0027] The present invention also provides a method of modifying a target locus of interest, the method comprising delivering to the locus a non-naturally occurring or engineered composition comprising a C2c2 locus effector protein and one or more nucleic acid components, wherein the C2c2 effector protein forms a complex with the one or more nucleic acid components, and upon binding of the complex to the locus of interest, the effector protein induces modification of the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break.
[0028] In such methods, the target locus of interest may be comprised in a nucleic acid molecule in vitro. In such methods, the target locus of interest may be comprised in a nucleic acid molecule in a cell. Preferably, in such methods, the target locus of interest may be comprised in an RNA molecule in vitro. Also preferably, in such methods, the target locus of interest may be comprised in an RNA molecule in a cell. The cell may be a prokaryotic or eukaryotic cell. The cell may be a mammalian cell. The cell may be a rodent cell. The cell may be a mouse cell.
[0029] In any of the described methods, the target locus of interest may be a genomic or epigenomic locus of interest. In any of the described methods, the complex may be delivered with multiple guides for multiplexed use. In any of the described methods, more than one protein may be used.
[0030] In a further embodiment of the present invention, the nucleic acid component may comprise a CRISPR RNA (crRNA) sequence. Without limitation, Applicants hypothesize that in such instances, the pre-crRNA may comprise a secondary structure sufficient for processing to yield the mature crRNA and for loading of the crRNA onto an effector protein. By way of example and not limitation, such secondary structure may comprise, consist essentially of, or consist of a stem-loop within the pre-crRNA, more particularly within a direct repeat.
[0031] In any of the described methods, the effector protein and nucleic acid components may be provided by one or more polynucleotide molecules encoding the protein and / or one or more nucleic acid components, and wherein the one or more polynucleotide molecules are operably configured to express the protein and / or one or more nucleic acid components. The one or more polynucleotide molecules may comprise one or more regulatory elements operably configured to express the protein and / or one or more nucleic acid components. The one or more polynucleotide molecules may be contained within one or more vectors. In any of the described methods, the target locus of interest may be a genomic or epigenomic locus of interest. In any of the described methods, the complex may be delivered with multiple guides for multiplexed applications. In any of the described methods, more than one protein may be used.
[0032] The regulatory element may comprise an inducible promoter. The polynucleotide and / or vector system may comprise an inducible system.
[0033] In any of the methods described, one or more polynucleotide molecules may be included in the delivery system, or one or more vectors may be included in the delivery system.
[0034] In any of the methods described, the non-naturally occurring or engineered composition may be delivered by liposomes, particles including nanoparticles, exosomes, microvesicles, gene guns, or one or more viral vectors.
[0035] The present invention also provides non-naturally occurring or engineered compositions that have the characteristics as discussed herein or are compositions defined in any of the methods described herein.
[0036] In certain embodiments, the invention therefore provides non-naturally occurring or engineered compositions, such as compositions specifically capable of or configured to modify a target locus of interest, said compositions comprising a Type VI CRISPR-Cas locus effector protein and one or more nucleic acid components, wherein the effector protein forms a complex with the one or more nucleic acid components, and wherein the effector protein induces modification of the target locus of interest upon binding of the complex to the locus of interest. In certain embodiments, the effector protein may be a C2c2 locus effector protein.
[0037] In a further aspect, the present invention also provides non-naturally occurring or engineered compositions, such as compositions capable of or configured to modify a target locus of interest, specifically comprising: (a) a guide RNA molecule (or a combination of guide RNA molecules, e.g., a first guide RNA molecule and a second guide RNA molecule, e.g., for multiplexing) or a nucleic acid encoding a guide RNA molecule (or one or more nucleic acids encoding a combination of guide RNA molecules); (b) a Type VI CRISPR-Cas locus effector protein or a nucleic acid encoding a Type VI CRISPR-Cas locus effector protein. In a specific embodiment, the effector protein may be a C2c2 locus effector protein.
[0038] In a further aspect, the present invention also provides a non-naturally occurring or engineered composition comprising: (a) a guide RNA molecule (or a combination of guide RNA molecules, e.g., a first guide RNA molecule and a second guide RNA molecule) or a nucleic acid encoding a guide RNA molecule (or one or more nucleic acids encoding a combination of guide RNA molecules); and (b) a C2c2 locus effector protein.
[0039] The present invention also provides vector systems comprising one or more vectors, wherein the one or more vectors comprise one or more polynucleotide molecules encoding components of a non-naturally occurring or engineered composition, the composition having the characteristics as defined in the methods described herein.
[0040] The present invention also provides delivery systems comprising one or more vectors or one or more polynucleotide molecules, wherein the one or more vectors or polynucleotide molecules comprise one or more polynucleotide molecules encoding components of a non-naturally occurring or engineered composition having the characteristics as discussed herein or being a composition defined by any of the methods described herein.
[0041] The present invention also provides non-naturally occurring or engineered compositions, or one or more polynucleotides encoding components of said compositions, or vectors or delivery systems comprising one or more polynucleotides encoding components of said compositions, for use in therapeutic treatment methods, which may include gene or transcriptome editing, or gene therapy.
[0042] The present invention also provides methods and compositions in which one or more amino acid residues in an effector protein, e.g., an engineered or non-naturally occurring effector protein or C2c2, may be modified. In some embodiments, the modification may include a mutation of one or more amino acid residues in the effector protein. The one or more mutations may be in one or more catalytic domains of the effector protein. The effector protein may have reduced or eliminated nuclease activity compared to an effector protein lacking the one or more mutations. The effector protein may not induce RNA strand cleavage at a target locus of interest. In a preferred embodiment, the one or more mutations may include two mutations. In a preferred embodiment, one or more amino acid residues are modified in a C2c2 effector protein, e.g., an engineered or non-naturally occurring effector protein or C2c2. In particular embodiments, the one or more modified or mutated amino acid residues are one or more of the amino acid residues in C2c2 corresponding to R597, H602, R1278 and H1283 (referenced to the Lsh C2c2 amino acids), such as mutations R597A, H602A, R1278A and H1283A, or the corresponding amino acid residues in an Lsh C2c2 orthologue.
[0043] In particular embodiments, the one or more modified or mutated amino acid residues are K2, K39, V40, E479, L514, V518, N524, G534, K535, E580, L597, V602, D630, F676, L709, I713, R717(HEPN), N718, H722(HEPN), E773, P823, V828, I879, Y880, F884, Y997, L1001, F100, based on the C2c2 consensus numbering. In certain embodiments, the one or more modified or mutated amino acid residues are one or more of the amino acid residues in C2c2 corresponding to R717 and R1509. In certain embodiments, the one or more modified or mutated amino acid residues are one or more of the amino acid residues of C2c2 corresponding to K2, K39, K535, K1261, R1362, R1372, K1546, and K1548. In certain embodiments, the mutations result in a protein with altered or modified activity. In certain embodiments, the mutations result in a protein with increased activity, such as increased specificity. In certain embodiments, the mutations result in a protein with decreased activity, such as decreased specificity. In certain embodiments, the mutations result in a protein without catalytic activity (i.e., a "dead" C2c2). In some embodiments, the amino acid residues correspond to Lsh C2c2 amino acid residues or the corresponding amino acid residues in a C2c2 protein of another species.
[0044] In certain embodiments, the one or more modified or mutated amino acid residues are as shown in Figure 3, i.e., Leptotrichia wadei F0279 ("Lew2" or "Lw2") and Listeria newyorkensis FSL M6-0635 (Listeriaceae bacterium FSL M6-0635 (also known as "Lib" or "LbFSL")) as a standard, the following amino acids are listed: M35, K36, T38, K39, I57, E65, G66, L68, N84, T86, E88, I103, N105, E123, R128, R129, K139, L152, L194, N196, K198, N201, Y222, D253, I266, F267, S280, I303, N306, R331, Y338, K389, Y390, K391, I434, K435, L458, D459, E462, L463, I478, E479, K494, R495, N498, S501, E519, N524, Y529, V530, G534, K5 35, Y539, T549, D551, R577, E580, A581, F582, I587, A593, L597, I601, L602, E611, E613, D630, I631, G633, K641, N646, V669, F676, S678, N695, E703, A707, I709, I713, I716, R717, H722, F740, F742, K768, I774, K778, I783, L787, S789, V792, Y796, D7 99, F812, N818, P820, F821, V822, P823, S824, F825, Y829, K831, D837, L852, F858, E867, A871, L875, K877, Y880, Y881, F884 , F888, F896, N901, V903, N915, K916, R918, Q920, E951, P956, Y959, Q964, I969, N994, F1000, I10001, Q1003, F10005, K1007 , G1008, F1009, N1019, L1020, K1021, I1023, N1028, E1070, I1075, K1076, F1092, K1097, L1099, L1104, L1107, K1113, Y1114,E1149, E1151, I1153, L1155, L1158, D1166, L1203, D1222, G1224, I1228, R1236, K1243, Y 1244, G1245, D1255, K1261, S1263, L1267, E1269, K1274, I1277, E1278, L1289, H1290, A12 94, N1320, K1325, E1327, Y1328, I1334, Y1337, K1341, N1342, K1343, N1350, L1352, L1355 , L1356, I1359, L1360, R1362, V1363, G1364, Y1365, I1369, R1371, D1372, F1385, E1391, D 1459, K1463, K1466, R1509, N1510, I1512, A1513, H1514, N1516, Y1517, L1529, L1530, E15 34, L1536, R1537, Y1543, D1544, R1545, K1546, L1547, K1548, N1549, A1550, K1553, S1554 , D1557, I1558, L1559, G1563, F1568, I1612, L1651, E1652, K1655, H1658, L1659, K1663, T1673, S1677, E1678, E1679, C1681, V1684, K1685, E1689. As noted above, in certain embodiments, the above list of amino acid residues excludes residues corresponding to R597, H602, R1278, and H1283 (based on the Lsh C2c2 amino acids).
[0045] In certain embodiments, the one or more modified or mutated amino acid residues are one or more conserved charged amino acid residues, which may be mutated to alanine.
[0046] In certain embodiments, the one or more modified or mutated amino acid residues are one or more of the amino acid residues of C2c2 corresponding to K28, K31, R44, E162, E184, K262, E288, K357, E360, K338, R441(HEPN), H446(HEPN), E471, K482, K525, K558, D707, R790, K811, R833, E839, R885, E894, R895, D896, K942, R960(HEPN), H965(HEPN), D990, K992, K994, based on the consensus sequence as shown in Figure 2, i.e., based on the alignment of C2c2 orthologs as shown in Figure 1. As noted above, in certain embodiments, the above list of amino acid residues excludes residues corresponding to R597, H602, R1278 and H1283 (based on the Lsh C2c2 amino acid).
[0047] The present invention also provides one or more mutations or two or more mutations in the catalytically active domain of an effector protein. In certain embodiments, the one or more mutations or two or more mutations may be in a catalytically active domain of an effector protein comprising a HEPN domain or a catalytically active domain homologous to a HEPN domain. The effector protein may include one or more heterologous functional domains. The one or more heterologous functional domains may include one or more nuclear localization signal (NLS) domains. The one or more heterologous functional domains may include at least two or more NLS domains. The one or more NLS domains may be located at, near, or adjacent to the terminus of the effector protein (e.g., C2c2), and in the case of two or more NLSs, each of the two NLSs may be located at, near, or adjacent to the terminus of the effector protein (e.g., C2c2). The one or more heterologous functional domains may include one or more translation activation domains. In other embodiments, the functional domain may include a transcription activation domain, e.g., VP64. The one or more heterologous functional domains may include one or more transcription repression domains. In certain embodiments, the transcriptional repression domain may comprise a KRAB domain or a SID domain (e.g., SID4X). The one or more heterologous functional domains may comprise one or more nuclease domains. In a preferred embodiment, the nuclease domain comprises Fok1.
[0048] The present invention also provides that the one or more heterologous functional domains have one or more of the following activities: methylase activity, demethylase activity, translation activation activity, translation repression activity, transcription activation activity, transcription repression activity, transcription termination factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity. In certain embodiments of the present invention, the one or more heterologous functional domains may comprise an epitope tag or reporter. Non-limiting examples of epitope tags include histidine (His) tag, V5 tag, FLAG tag, influenza hemagglutinin (HA) tag, Myc tag, VSV-G tag, and thioredoxin (Trx) tag. Examples of reporters include, but are not limited to, glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins such as blue fluorescent protein (BFP).
[0049] At least one heterologous functional domain may be at or near the amino terminus of the effector protein, and / or wherein at least one heterologous functional domain is at or near the carboxy terminus of the effector protein. One or more heterologous functional domains may be fused to the effector protein. One or more heterologous functional domains may be tethered to the effector protein. One or more heterologous functional domains may be linked to the effector protein via a linker moiety.
[0050] The present invention also provides a method for the prevention and treatment of bacteria of the genera Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaete, and the like. erochaeta, Lactobacillus, Eubacterium, Corynebacter, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Clostridium iaridium, Leptotrichia, Francisella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus, Leptospira Also provided are effector proteins, including effector proteins derived from organisms of the genera including: Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium or Acidaminococcus.The effector protein can comprise a chimeric effector protein comprising a first fragment derived from a first effector protein ortholog and a second fragment derived from a second effector protein ortholog, where the first and second effector protein orthologs are different.At least one of the first and second effector protein orthologs is from a species of the genera Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, or the like. ospirillum, Sphaerochaeta, Lactobacillus, Eubacterium, Corynebacter, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Clostridium, Clostridiaridium, Leptotrichia, Francisella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus The effector protein may be derived from an organism including the genera Letospira, Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium, or Acidaminococcus.
[0051] In certain embodiments, the effector protein, in particular a type VI locus effector protein, more particularly C2c2p, may originate from, be isolated from or be derived from a bacterial species belonging to the taxonomic groups alpha-proteobacteria, Bacilli, Clostridia, Fusobacteria and Bacteroidetes. In certain embodiments, the effector protein, in particular a type VI locus effector protein, more particularly C2c2p, may originate from, be isolated from or be derived from a bacterial species belonging to a genus selected from the group consisting of Lachnospiraceae, Clostridium, Carnobacterium, Paludibacter, Listeria, Leptotrichia and Rhodobacter.In particular embodiments, the effector protein, in particular a type VI locus effector protein, more particularly C2c2p, is selected from the group consisting of Lachnospiraceae bacterium MA2020, Lachnospiraceae bacterium NK4A179, Clostridium aminophilum (e.g., DSM 10710), Lachnospiraceae bacterium NK4A144, Carnobacterium gallinarum (e.g., DSM 4847 strain MT44), Paludibacter propionicigenes (e.g., WB4), Listeria seeligeri (e.g., serotype 1 / 2b strain SLCC3954), Listeria weihenstephanensis (e.g., Lactobacillus casei), Lactobacillus casei ... weihenstephanensis (e.g., FSL R9-0317 c4), Listeria newyorkensis (e.g., strain FSL M6-0635; also “LbFSL”), Leptotrichia wadei (e.g., F0279; also “Lw” or “Lw2”), Leptotrichia buccalis (e.g., DSM 1135), Leptotrichia sp. oral bacterial taxon 225 (e.g., strain F0581), Leptotrichia sp. oral bacterial taxon 879 (e.g., strain F0557), Leptotrichia shahii (e.g., DSM 19757), Rhodobacter capsulatus (e.g., SB 1003, R121, or DE442), or may originate from, be isolated from, or be derived from a bacterial species selected from the group consisting of: Rhodobacter capsulatus (e.g., SB 1003, R121, or DE442).In certain preferred embodiments, the C2c2 effector protein is selected from the group consisting of Listeriaceae bacteria (e.g., FSL M6-0635; also "LbFSL"), Lachnospiraceae bacteria MA2020, Lachnospiraceae bacteria NK4A179, Clostridium aminophilum (e.g., DSM 10710), Carnobacterium gallinarum (e.g., DSM 4847), Paludibacter propionicigenes (e.g., WB4), Listeria seeligeri (e.g., serotype 1 / 2b strain SLCC3954), Listeria weihenstephanensis (e.g., FSL R9-0317 c4), Leptotrichia wadei (e.g., F0279; also "Lw" or "Lw2"), Leptotrichia shahii (e.g., DSM 19757), Rhodobacter capsulatus (e.g., SB 1003, R121, or DE442); preferably those derived from the Listeriaceae bacterium FSL M6-0635 (i.e., Listeria newyorkensis FSL M6-0635; "LbFSL" in Figures 4 to 7) or Leptotrichia wadei F0279 (also "Lw" or "Lw2").
[0052] In certain embodiments, a Type VI locus as contemplated herein may encode Cas1, Cas2, and C2c2p effector proteins.
[0053] In certain embodiments, the effector protein, particularly a type VI locus effector protein, more particularly a C2c2p such as a native C2c2p, may be about 1000 to about 1500 amino acids in length, e.g., about 1100 to about 1400 amino acids in length, e.g., about 1000 to about 1100, about 1100 to about 1200 amino acids in length, or about 1200 to about 1300 amino acids in length, or about 1300 to about 1400 amino acids in length, or about 1400 to about 1500 amino acids in length, e.g., about 1000, about 1100, about 1200, about 1300, about 1400, or about 1500 amino acids in length.
[0054] In certain embodiments, effector proteins, particularly type VI locus effector proteins, more particularly C2c2p, contain at least one and preferably at least two, e.g., more preferably, exactly two, conserved RxxxxH motifs. Catalytic RxxxxH motifs are unique to HEPN (Higher Eukaryotes and Prokaryotes Nucleotide-binding) domains. Thus, in certain embodiments, effector proteins, particularly type VI locus effector proteins, more particularly C2c2p, contain at least one and preferably at least two, e.g., more preferably, exactly two, HEPN domains. In certain embodiments, the HEPN domain may have RNase activity. In other embodiments, the HEPN domain may have DNase activity.
[0055] In particular embodiments, a Type VI locus as contemplated herein may comprise CRISPR repeats that are 30-40 bp in length, more typically 35-39 bp in length, e.g., 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 bp in length. In particular embodiments, the direct repeats are at least 25 nt in length.
[0056] In certain embodiments, a protospacer adjacent motif (PAM) or PAM-like motif directs binding of an effector protein complex as disclosed herein to a target locus of interest. In some embodiments, the PAM may be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM may be a 3' PAM (i.e., located downstream of the 5' end of the protospacer). The term "PAM" may be used synonymously with the term "PFS" or "protospacer adjacent site" or "protospacer adjacent sequence."
[0057] In preferred embodiments, the effector protein, particularly a type VI locus effector protein, more particularly C2c2p, can recognize a 3' PAM. In certain embodiments, the effector protein, particularly a type VI locus effector protein, more particularly C2c2p, can recognize a 3' PAM that is 5' H (wherein H is A, C, or U). In certain embodiments, the effector protein may be Leptotrichia shahii C2c2p, more preferably Leptotrichia shahii DSM 19757 C2c2, and the 5' PAM is 5' H. In certain embodiments, the effector protein may be Leptotrichia wadei F0279 (Lw2) C2c2, and the 5' PAM is H (wherein H is C, U, or A).
[0058] In certain embodiments, the CRISPR enzyme can be engineered to contain one or more mutations that reduce or eliminate nuclease activity. Mutations can also be made to adjacent residues, such as amino acids near those listed above that are involved in nuclease activity. In some embodiments, only one HEPN domain is inactivated, while in other embodiments, a second HEPN domain is inactivated.
[0059] In certain embodiments of the invention, the guide RNA or mature crRNA comprises, consists essentially of, or consists of a direct repeat sequence and a guide sequence or spacer sequence. In certain embodiments, the guide RNA or mature crRNA comprises, consists essentially of, or consists of a direct repeat sequence linked to a guide sequence or spacer sequence. In certain embodiments, the guide RNA or mature crRNA comprises a 19-nt partial direct repeat followed by an 18-, 19-, 20-, 21-, 22-, 23-, 24-, 25-, or more-nt guide sequence, e.g., 18-25, 19-25, 20-25, 21-25, 22-25, or 23-25-nt guide sequence or spacer sequence. In certain embodiments, the effector protein is a C2c2 effector protein, and a guide sequence of at least 16 nt is required to achieve detectable DNA cleavage and a guide sequence of at least 17 nt is required to achieve efficient DNA cleavage in vitro. In particular embodiments, the effector protein is a C2c2 protein, and at least 19 nt of the guide sequence is required to achieve detectable RNA cleavage. In certain embodiments, the direct repeat sequence is located upstream (i.e., 5') of the guide sequence or spacer sequence. In preferred embodiments, the seed sequence of the C2c2 guide RNA (i.e., the critical sequence essential for recognition and / or hybridization with a sequence at the target locus) is located within approximately the first 5 nt on the 5' end of the guide sequence or spacer sequence.
[0060] In preferred embodiments of the present invention, the mature crRNA comprises a stem-loop or optimized stem-loop structure or optimized secondary structure. In preferred embodiments, the mature crRNA comprises a stem-loop or optimized stem-loop structure in a direct repeat sequence, where the stem-loop or optimized stem-loop structure is important for cleavage activity. In certain embodiments, the mature crRNA preferably comprises a single stem-loop. In certain embodiments, the direct repeat sequence preferably comprises a single stem-loop. In certain embodiments, the cleavage activity of the effector protein complex is modified by introducing mutations that affect the stem-loop RNA duplex structure. In preferred embodiments, mutations that maintain the stem-loop RNA duplex may be introduced, thereby maintaining the cleavage activity of the effector protein complex. In other preferred embodiments, mutations that disrupt the stem-loop RNA duplex structure may be introduced, thereby completely abolishing the cleavage activity of the effector protein complex.
[0061] In specific embodiments, the C2c2 protein is an Lsh C2c2 effector protein, and the mature crRNA comprises a stem-loop or optimized stem-loop structure. In specific embodiments, the direct repeat of the crRNA comprises at least 25 nucleotides, including the stem-loop. In specific embodiments, the stem allows for individual base swaps, but activity is disrupted by most secondary structure changes or truncations of the crRNA. Examples of disruptive mutations include swapping three or more stem nucleotides, adding unpaired nucleotides in the stem, shortening the stem (by removing one of the paired nucleotides), or extending the stem (by adding a pair of paired nucleotides). However, the crRNA may also be capable of 5' and / or 3' extension to include non-functional RNA sequences as envisioned for specific applications described herein.
[0062] The present invention also provides nucleotide sequences encoding effector proteins that are codon-optimized for expression in eukaryotic organisms or cells in any of the methods or compositions described herein. In certain embodiments of the invention, the nucleotide sequence encoding the codon-optimized effector protein encodes any C2c2 discussed herein and is codon-optimized for operation in a eukaryotic cell or organism, such as a cell or organism listed elsewhere herein, including, without limitation, a yeast cell, or a mammalian cell or organism, such as a mouse cell, a rat cell, and a human cell, or a non-human eukaryotic organism, such as a plant.
[0063] In certain embodiments of the present invention, at least one nuclear localization signal (NLS) is added to a nucleic acid sequence encoding a C2c2 effector protein. In preferred embodiments, at least one or more C- or N-terminal NLSs are added (thus, one or more nucleic acid molecules encoding a C2c2 effector protein can contain coding for one or more NLSs, and thus, the expressed product will have one or more NLSs attached or connected). In certain embodiments of the present invention, at least one nuclear export signal (NES) is added to a nucleic acid sequence encoding a C2c2 effector protein. In preferred embodiments, at least one or more C- or N-terminal NESs are added (thus, one or more nucleic acid molecules encoding a C2c2 effector protein can contain coding for one or more NESs, and thus, the expressed product will have one or more NESs attached or connected). In preferred embodiments, C- and / or N-terminal NLSs or NESs are added for optimal expression and nuclear targeting in eukaryotic cells, preferably human cells. In a preferred embodiment, the codon-optimized effector protein is C2c2, and the spacer length of the guide RNA is 15 to 35 nt. In a specific embodiment, the spacer length of the guide RNA is at least 16 nucleotides, such as at least 17 nucleotides, preferably at least 18 nt, such as preferably at least 19 nt, at least 20 nt, at least 21 nt, or at least 22 nt. In a specific embodiment, the spacer length is 15 to 17 nt, 17 to 20 nt, 20 to 24 nt, such as 20, 21, 22, 23, or 24 nt, 23 to 25 nt, such as 23, 24, or 25 nt, 24 to 27 nt, 27 to 30 nt, 30 to 35 nt, or 35 nt or more. In a specific embodiment of the present invention, the codon-optimized effector protein is C2c2, and the direct repeat length of the guide RNA is at least 16 nucleotides. In certain embodiments, the codon-optimized effector protein is C2c2, and the direct repeat length of the guide RNA is 16-20 nt, for example, 16, 17, 18, 19, or 20 nucleotides.In a particularly preferred embodiment, the direct repeat length of the guide RNA is 19 nucleotides.
[0064] The present invention also encompasses methods for delivering multiple nucleic acid components, each specific for a different target locus of interest, thereby modifying multiple target loci of interest. The nucleic acid components of the complex may comprise one or more protein-binding RNA aptamers. The one or more aptamers may be capable of binding to a bacteriophage coat protein. The bacteriophage coat protein may be selected from the group including Qβ, F2, GA, fr, JP501, MS2, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s, and PRR1. In a preferred embodiment, the bacteriophage coat protein is MS2. The present invention also provides nucleic acid components of the complex that are 30 or more, 40 or more, or 50 or more nucleotides in length.
[0065] Accordingly, it is an object of the present invention not to include within its scope any previously known product, process for making a product, or method for using a product, to which the applicants reserve their rights and which hereby disclose a disclaimer of any previously known product, process, or method. It is further noted that the present invention is not intended to include within its scope any product, process for making a product, or method for using a product that does not meet the description and enablement requirements of the United States Patent and Trademark Office (USPTO) (35 U.S.C. § 112, first paragraph) or the European Patent Office (EPO) (Article 83 EPC), to which the applicants reserve their rights and which hereby disclose a disclaimer of any previously described product, process for making a product, or method for using a product. Compliance with Article 53(c) EPC and Rule 28(b) and (c) EPC may be advantageous in the practice of the present invention. Nothing herein should be construed as a promise.
[0066] It is noted that in this disclosure, and particularly in the claims and / or paragraphs, terms such as "comprises," "comprised," "comprising," and the like may have the meaning ascribed to them in U.S. patent law; e.g., they may mean "includes," "included," "including," and the like; and that terms such as "consisting essentially of" and "consists essentially of" have the meaning ascribed to them in U.S. patent law, e.g., they may allow for elements not expressly recited, but exclude elements found in the prior art or that affect a basic or novel characteristic of the invention.
[0067] In a further aspect, the present invention provides a eukaryotic cell comprising a nucleotide sequence encoding a CRISPR system described herein that ensures modification of a target locus of interest, wherein the target locus of interest is modified by any of the methods described herein. A further aspect provides a cell line of said cell. Another aspect provides a multicellular organism comprising one or more of said cells.
[0068] In certain embodiments, modification of a target locus of interest may result in a eukaryotic cell comprising an altered (protein) expression of at least one gene product; a eukaryotic cell comprising an altered (protein) expression of at least one gene product, wherein the (protein) expression of at least one gene product is increased; a eukaryotic cell comprising an altered (protein) expression of at least one gene product, wherein the (protein) expression of at least one gene product is decreased; or a eukaryotic cell comprising an edited transcriptome.
[0069] In certain embodiments, the eukaryotic cell may be a mammalian cell or a human cell.
[0070] In further embodiments, non-naturally occurring or engineered compositions, vector systems, or delivery systems as described herein may be used for RNA sequence-specific interference, RNA sequence-specific modification of expression (including isoform-specific expression), stability, localization, functionality (e.g., ribosomal RNA or miRNA), etc.; or multiplexing such processes.
[0071] In further embodiments, the non-naturally occurring or engineered compositions, vector systems, or delivery systems described herein can be used for RNA detection and / or quantification in a sample, such as a biological sample. In certain embodiments, RNA detection is in a cell. In some embodiments, the present invention provides a method for detecting a target RNA in a sample, the method comprising: (a) incubating the sample with i) a type VI CRISPR-Cas effector protein capable of cleaving RNA, ii) a guide RNA capable of hybridizing to the target RNA, and iii) an RNA-based cleavage-inducible reporter capable of being non-specifically and detectably cleaved by the effector protein; and (b) detecting the target RNA based on a signal generated by cleavage of the RNA-based cleavage-inducible reporter.
[0072] In some embodiments, the type VI CRISPR-Cas effector protein is a C2c2 effector protein. In some embodiments, the RNA-based cleavage-inducible reporter construct comprises a fluorescent dye and a quencher. In certain embodiments, the sample comprises a cell-free biological sample. In other embodiments, the sample comprises a cellular sample, such as, without limitation, a plant cell or an animal cell. In some embodiments of the present invention, the target RNA comprises pathogen RNA, including, but not limited to, target RNA from a virus, bacterium, fungus, or parasite. In some embodiments, the guide RNA is designed to detect a target RNA or splice variant of an RNA transcript that contains a single nucleotide polymorphism. In some embodiments, the guide RNA comprises one or more mismatched nucleotides with the target RNA. In certain embodiments, the guide RNA hybridizes to a target molecule that is diagnostic for a disease state, such as, but not limited to, cancer or an immune disorder.
[0073] The present invention provides a ribonucleic acid (RNA) detection system comprising: a) a type VI CRISPR-Cas effector protein capable of cleaving RNA; b) a guide RNA capable of binding to a target RNA; and c) an RNA-based cleavage-inducible reporter capable of being non-specifically and detectably cleaved by the effector protein. The present invention also provides a kit for RNA detection comprising: a) a type VI CRISPR-Cas effector protein capable of cleaving RNA; and b) an RNA-based cleavage-inducible reporter capable of being non-specifically and detectably cleaved by the effector protein. In certain embodiments, the RNA-based cleavage-inducible reporter construct comprises a fluorescent dye and a quencher.
[0074] In further embodiments, the non-naturally occurring or engineered compositions, vector systems, or delivery systems as described herein may be used in the generation of disease models and / or screening systems.
[0075] In further embodiments, the non-naturally occurring or engineered compositions, vector systems, or delivery systems as described herein may be used for site-specific transcriptome editing or purturbation; nucleic acid sequence-specific interference; or multiplex genome engineering.
[0076] Also provided are gene products from cells, cell lines, or organisms as described herein. In certain embodiments, the amount of expression of the gene product may be greater or less than the amount of the gene product from a cell that does not have the altered expression or edited genome. In certain embodiments, the gene product may be altered compared to the gene product from a cell that does not have the altered expression or edited genome.
[0077] These and other embodiments are disclosed or are apparent from and encompassed by the following detailed description.
[0078] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which: [Brief explanation of the drawings]
[0079] [Figure 1A]Figure 1A-1D. Cellular localization of C2c2 orthologs. (A) HEK293 cells were transfected with various C2c2 orthologs fused to mCherry with or without a nuclear localization signal (NLS) or nuclear export signal (NES). (B) NES fusions of Leptotricia wadei F0279 C2c2 (B1) and Lachnospiraceae bacterium NK4A179 C2c2 (B2) are localized in the cytoplasm. (C) NLS fusions of Leptotrichia wadei F0279 C2c2(C1), Lachnospiraceae bacterium NK4A179 C2c2(C2), and Leptotrichia shahii C2c2(C3) are located in the nucleus and variably in the nucleolus. (C) Leptotrichia wadei F0279 C2c2(D1), Lachnospiraceae bacterium NK4A179 C2c2(D2), and Leptotrichia shahii C2c2(D3) not fused to an NLS or NES are variably located in the nucleus and cytoplasm. (D) Schematic of the four mammalian LwaCas13a constructs evaluated (left) and imaging showing the localization and expression of each of the designs (right). [Figure 1B]Figure 1A-1D. Cellular localization of C2c2 orthologs. (A) HEK293 cells were transfected with various C2c2 orthologs fused to mCherry with or without a nuclear localization signal (NLS) or nuclear export signal (NES). (B) NES fusions of Leptotricia wadei F0279 C2c2 (B1) and Lachnospiraceae bacterium NK4A179 C2c2 (B2) are localized in the cytoplasm. (C) NLS fusions of Leptotrichia wadei F0279 C2c2(C1), Lachnospiraceae bacterium NK4A179 C2c2(C2), and Leptotrichia shahii C2c2(C3) are located in the nucleus and variably in the nucleolus. (C) Leptotrichia wadei F0279 C2c2(D1), Lachnospiraceae bacterium NK4A179 C2c2(D2), and Leptotrichia shahii C2c2(D3) not fused to an NLS or NES are variably located in the nucleus and cytoplasm. (D) Schematic of the four mammalian LwaCas13a constructs evaluated (left) and imaging showing the localization and expression of each of the designs (right). [Figure 1C]Figure 1A-1D. Cellular localization of C2c2 orthologs. (A) HEK293 cells were transfected with various C2c2 orthologs fused to mCherry with or without a nuclear localization signal (NLS) or nuclear export signal (NES). (B) NES fusions of Leptotricia wadei F0279 C2c2 (B1) and Lachnospiraceae bacterium NK4A179 C2c2 (B2) are localized in the cytoplasm. (C) NLS fusions of Leptotrichia wadei F0279 C2c2(C1), Lachnospiraceae bacterium NK4A179 C2c2(C2), and Leptotrichia shahii C2c2(C3) are located in the nucleus and variably in the nucleolus. (C) Leptotrichia wadei F0279 C2c2(D1), Lachnospiraceae bacterium NK4A179 C2c2(D2), and Leptotrichia shahii C2c2(D3) not fused to an NLS or NES are variably located in the nucleus and cytoplasm. (D) Schematic of the four mammalian LwaCas13a constructs evaluated (left) and imaging showing the localization and expression of each of the designs (right). [Figure 1D]Figure 1A-1D. Cellular localization of C2c2 orthologs. (A) HEK293 cells were transfected with various C2c2 orthologs fused to mCherry with or without a nuclear localization signal (NLS) or nuclear export signal (NES). (B) NES fusions of Leptotricia wadei F0279 C2c2 (B1) and Lachnospiraceae bacterium NK4A179 C2c2 (B2) are localized in the cytoplasm. (C) NLS fusions of Leptotrichia wadei F0279 C2c2(C1), Lachnospiraceae bacterium NK4A179 C2c2(C2), and Leptotrichia shahii C2c2(C3) are located in the nucleus and variably in the nucleolus. (C) Leptotrichia wadei F0279 C2c2(D1), Lachnospiraceae bacterium NK4A179 C2c2(D2), and Leptotrichia shahii C2c2(D3) not fused to an NLS or NES are variably located in the nucleus and cytoplasm. (D) Schematic of the four mammalian LwaCas13a constructs evaluated (left) and imaging showing the localization and expression of each of the designs (right). [Figure 2A-2B]Figures 2A-2C. Assays for evaluating C2c2 knockdown efficiency in mammalian cells. (A) Schematic diagram of the dual-luciferase reporter scheme used to determine the knockdown efficiency of C2c2. Cells were transfected with various plasmids: plasmids encoding Gaussia luciferase (Gluc) and Cypridina luciferase (Cluc); plasmids encoding C2c2 (with or without an NLS or NES); and a plasmid encoding a gRNA targeting Gaussia luciferase. The knockdown efficiency of Gluc was determined based on the ratio of Gluc to Cluc. (B) Location of various gRNAs targeting Gaussia luciferase mRNA. (C) Assays for evaluating the knockdown efficiency of C2c2 in mammalian cells. Cells were transfected with various plasmids: plasmids encoding EGFP; plasmids encoding C2c2 (with or without an NLS or NES); and a plasmid encoding a gRNA targeting EGFP. The knockdown efficiency of EGFP was determined by flow cytometry. [Figure 2C] Figures 2A-2C. Assays for evaluating C2c2 knockdown efficiency in mammalian cells. (A) Schematic diagram of the dual-luciferase reporter scheme used to determine the knockdown efficiency of C2c2. Cells were transfected with various plasmids: plasmids encoding Gaussia luciferase (Gluc) and Cypridina luciferase (Cluc); plasmids encoding C2c2 (with or without an NLS or NES); and a plasmid encoding a gRNA targeting Gaussia luciferase. The knockdown efficiency of Gluc was determined based on the ratio of Gluc to Cluc. (B) Location of various gRNAs targeting Gaussia luciferase mRNA. (C) Assays for evaluating the knockdown efficiency of C2c2 in mammalian cells. Cells were transfected with various plasmids: plasmids encoding EGFP; plasmids encoding C2c2 (with or without an NLS or NES); and a plasmid encoding a gRNA targeting EGFP. The knockdown efficiency of EGFP was determined by flow cytometry. [Figure 3A-3B]Figures 3A-3Q. Characterization of the Type VI CRISPR Cas13a family for RNA cleavage activity. (A) Schematic of the PFS characterization screen for Cas13a orthologs. (B) Schematic of the constructs for expression of Cas13a protein and crRNA. Length and residue numbers shown are for LwaCas13a. (C) Quantification of Cas13a in vivo activity. (D) Ratio of in vivo activity from (C). (E, F) PFS motif determination for LshC2c2 and LwaC2c2 as determined by next-generation sequencing of enriched plasmids in a viable E. coli population. (G) Distribution of PFS enrichments for LshCas13a and LwaCas13a in targeting and non-targeting samples. (H) In vivo PFS screening indicates that LwaCas13a has minimal PFS preference. (I) Boxplot showing the distribution of normalized PFS counts for targeting and non-targeting bioreplicates (n=2) of LshC2c2 and LwC2c2. The boxes extend from the first to third quartiles, and the whiskers represent 1.5x the interquartile range. The mean is indicated by a red horizontal bar. (J) Comparison of LshC2c2 and LwC2c2 RNA cleavage activity using purified proteins. LwaCas13a has higher RNase activity than LshCas13a. 5'-labeled RNA targets were incubated with C2c2 and the corresponding crRVA for 30 minutes, and the products were analyzed by gel electrophoresis. (K) LwaCas13a can process CRISPR array transcripts from the L. wadeii CRISPR locus. The two-spacer Lwa array was incubated with LwC2c2 for 30 minutes, and the reactions were then separated by gel electrophoresis. (L) Schematic of the mammalian LwaCas13a constructs evaluated and imaging showing the localization and expression of each design. Scale bar, 10 μm. (M) Schematic of the mammalian reporter system used to evaluate knockdown by luciferase protein readout. (N) Knockdown of Gaussia luciferase (Gluc) using engineered variants of LwaCas13a. Guide and shRNA sequences are shown at the top.(O) Knockdown of three different endogenous transcripts by LwaCas13a compared to the corresponding RNAi constructs. (P) Schematic of LwaCas13a knockdown of transcripts in rice (Oryza sativa) protoplasts. (Q) LwaCas13a knockdown of three transcripts in rice (Oryza sativa) protoplasts using three targeting and non-targeting guides for each transcript. All values are mean ± SEM with n=3 unless otherwise noted. [Figure 3C-3D]Figures 3A-3Q. Characterization of the Type VI CRISPR Cas13a family for RNA cleavage activity. (A) Schematic of the PFS characterization screen for Cas13a orthologs. (B) Schematic of the constructs for expression of Cas13a protein and crRNA. Length and residue numbers shown are for LwaCas13a. (C) Quantification of Cas13a in vivo activity. (D) Ratio of in vivo activity from (C). (E, F) PFS motif determination for LshC2c2 and LwaC2c2 as determined by next-generation sequencing of enriched plasmids in a viable E. coli population. (G) Distribution of PFS enrichments for LshCas13a and LwaCas13a in targeting and non-targeting samples. (H) In vivo PFS screening indicates that LwaCas13a has minimal PFS preference. (I) Boxplot showing the distribution of normalized PFS counts for targeting and non-targeting bioreplicates (n=2) of LshC2c2 and LwC2c2. The boxes extend from the first to third quartiles, and the whiskers represent 1.5x the interquartile range. The mean is indicated by a red horizontal bar. (J) Comparison of LshC2c2 and LwC2c2 RNA cleavage activity using purified proteins. LwaCas13a has higher RNase activity than LshCas13a. 5'-labeled RNA targets were incubated with C2c2 and the corresponding crRVA for 30 minutes, and the products were analyzed by gel electrophoresis. (K) LwaCas13a can process CRISPR array transcripts from the L. wadeii CRISPR locus. The two-spacer Lwa array was incubated with LwC2c2 for 30 minutes, and the reactions were then separated by gel electrophoresis. (L) Schematic of the mammalian LwaCas13a constructs evaluated and imaging showing the localization and expression of each design. Scale bar, 10 μm. (M) Schematic of the mammalian reporter system used to evaluate knockdown by luciferase protein readout. (N) Knockdown of Gaussia luciferase (Gluc) using engineered variants of LwaCas13a. Guide and shRNA sequences are shown at the top.(O) Knockdown of three different endogenous transcripts by LwaCas13a compared to the corresponding RNAi constructs. (P) Schematic of LwaCas13a knockdown of transcripts in rice (Oryza sativa) protoplasts. (Q) LwaCas13a knockdown of three transcripts in rice (Oryza sativa) protoplasts using three targeting and non-targeting guides for each transcript. All values are mean ± SEM with n=3 unless otherwise noted. [Figures 3E-3I]Figures 3A-3Q. Characterization of the Type VI CRISPR Cas13a family for RNA cleavage activity. (A) Schematic of the PFS characterization screen for Cas13a orthologs. (B) Schematic of the constructs for expression of Cas13a protein and crRNA. Length and residue numbers shown are for LwaCas13a. (C) Quantification of Cas13a in vivo activity. (D) Ratio of in vivo activity from (C). (E, F) PFS motif determination for LshC2c2 and LwaC2c2 as determined by next-generation sequencing of enriched plasmids in a viable E. coli population. (G) Distribution of PFS enrichments for LshCas13a and LwaCas13a in targeting and non-targeting samples. (H) In vivo PFS screening indicates that LwaCas13a has minimal PFS preference. (I) Boxplot showing the distribution of normalized PFS counts for targeting and non-targeting bioreplicates (n=2) of LshC2c2 and LwC2c2. The boxes extend from the first to third quartiles, and the whiskers represent 1.5x the interquartile range. The mean is indicated by a red horizontal bar. (J) Comparison of LshC2c2 and LwC2c2 RNA cleavage activity using purified proteins. LwaCas13a has higher RNase activity than LshCas13a. 5'-labeled RNA targets were incubated with C2c2 and the corresponding crRVA for 30 minutes, and the products were analyzed by gel electrophoresis. (K) LwaCas13a can process CRISPR array transcripts from the L. wadeii CRISPR locus. The two-spacer Lwa array was incubated with LwC2c2 for 30 minutes, and the reactions were then separated by gel electrophoresis. (L) Schematic of the mammalian LwaCas13a constructs evaluated and imaging showing the localization and expression of each design. Scale bar, 10 μm. (M) Schematic of the mammalian reporter system used to evaluate knockdown by luciferase protein readout. (N) Knockdown of Gaussia luciferase (Gluc) using engineered variants of LwaCas13a. Guide and shRNA sequences are shown at the top.(O) Knockdown of three different endogenous transcripts by LwaCas13a compared to the corresponding RNAi constructs. (P) Schematic of LwaCas13a knockdown of transcripts in rice (Oryza sativa) protoplasts. (Q) LwaCas13a knockdown of three transcripts in rice (Oryza sativa) protoplasts using three targeting and non-targeting guides for each transcript. All values are mean ± SEM with n=3 unless otherwise noted. [Figures 3J-3K]Figures 3A-3Q. Characterization of the Type VI CRISPR Cas13a family for RNA cleavage activity. (A) Schematic of the PFS characterization screen for Cas13a orthologs. (B) Schematic of the constructs for expression of Cas13a protein and crRNA. Length and residue numbers shown are for LwaCas13a. (C) Quantification of Cas13a in vivo activity. (D) Ratio of in vivo activity from (C). (E, F) PFS motif determination for LshC2c2 and LwaC2c2 as determined by next-generation sequencing of enriched plasmids in a viable E. coli population. (G) Distribution of PFS enrichments for LshCas13a and LwaCas13a in targeting and non-targeting samples. (H) In vivo PFS screening indicates that LwaCas13a has minimal PFS preference. (I) Boxplot showing the distribution of normalized PFS counts for targeting and non-targeting bioreplicates (n=2) of LshC2c2 and LwC2c2. The boxes extend from the first to third quartiles, and the whiskers represent 1.5x the interquartile range. The mean is indicated by a red horizontal bar. (J) Comparison of LshC2c2 and LwC2c2 RNA cleavage activity using purified proteins. LwaCas13a has higher RNase activity than LshCas13a. 5'-labeled RNA targets were incubated with C2c2 and the corresponding crRVA for 30 minutes, and the products were analyzed by gel electrophoresis. (K) LwaCas13a can process CRISPR array transcripts from the L. wadeii CRISPR locus. The two-spacer Lwa array was incubated with LwC2c2 for 30 minutes, and the reactions were then separated by gel electrophoresis. (L) Schematic of the mammalian LwaCas13a constructs evaluated and imaging showing the localization and expression of each design. Scale bar, 10 μm. (M) Schematic of the mammalian reporter system used to evaluate knockdown by luciferase protein readout. (N) Knockdown of Gaussia luciferase (Gluc) using engineered variants of LwaCas13a. Guide and shRNA sequences are shown at the top.(O) Knockdown of three different endogenous transcripts by LwaCas13a compared to the corresponding RNAi constructs. (P) Schematic of LwaCas13a knockdown of transcripts in rice (Oryza sativa) protoplasts. (Q) LwaCas13a knockdown of three transcripts in rice (Oryza sativa) protoplasts using three targeting and non-targeting guides for each transcript. All values are mean ± SEM with n=3 unless otherwise noted. [Figure 3L-3N]Figures 3A-3Q. Characterization of the Type VI CRISPR Cas13a family for RNA cleavage activity. (A) Schematic of the PFS characterization screen for Cas13a orthologs. (B) Schematic of the constructs for expression of Cas13a protein and crRNA. Length and residue numbers shown are for LwaCas13a. (C) Quantification of Cas13a in vivo activity. (D) Ratio of in vivo activity from (C). (E, F) PFS motif determination for LshC2c2 and LwaC2c2 as determined by next-generation sequencing of enriched plasmids in a viable E. coli population. (G) Distribution of PFS enrichments for LshCas13a and LwaCas13a in targeting and non-targeting samples. (H) In vivo PFS screening indicates that LwaCas13a has minimal PFS preference. (I) Boxplot showing the distribution of normalized PFS counts for targeting and non-targeting bioreplicates (n=2) of LshC2c2 and LwC2c2. The boxes extend from the first to third quartiles, and the whiskers represent 1.5x the interquartile range. The mean is indicated by a red horizontal bar. (J) Comparison of LshC2c2 and LwC2c2 RNA cleavage activity using purified proteins. LwaCas13a has higher RNase activity than LshCas13a. 5'-labeled RNA targets were incubated with C2c2 and the corresponding crRVA for 30 minutes, and the products were analyzed by gel electrophoresis. (K) LwaCas13a can process CRISPR array transcripts from the L. wadeii CRISPR locus. The two-spacer Lwa array was incubated with LwC2c2 for 30 minutes, and the reactions were then separated by gel electrophoresis. (L) Schematic of the mammalian LwaCas13a constructs evaluated and imaging showing the localization and expression of each design. Scale bar, 10 μm. (M) Schematic of the mammalian reporter system used to evaluate knockdown by luciferase protein readout. (N) Knockdown of Gaussia luciferase (Gluc) using engineered variants of LwaCas13a. Guide and shRNA sequences are shown at the top.(O) Knockdown of three different endogenous transcripts by LwaCas13a compared to the corresponding RNAi constructs. (P) Schematic of LwaCas13a knockdown of transcripts in rice (Oryza sativa) protoplasts. (Q) LwaCas13a knockdown of three transcripts in rice (Oryza sativa) protoplasts using three targeting and non-targeting guides for each transcript. All values are mean ± SEM with n=3 unless otherwise noted. [Figures 3O-3Q]Figures 3A-3Q. Characterization of the Type VI CRISPR Cas13a family for RNA cleavage activity. (A) Schematic of the PFS characterization screen for Cas13a orthologs. (B) Schematic of the constructs for expression of Cas13a protein and crRNA. Length and residue numbers shown are for LwaCas13a. (C) Quantification of Cas13a in vivo activity. (D) Ratio of in vivo activity from (C). (E, F) PFS motif determination for LshC2c2 and LwaC2c2 as determined by next-generation sequencing of enriched plasmids in a viable E. coli population. (G) Distribution of PFS enrichments for LshCas13a and LwaCas13a in targeting and non-targeting samples. (H) In vivo PFS screening indicates that LwaCas13a has minimal PFS preference. (I) Boxplot showing the distribution of normalized PFS counts for targeting and non-targeting bioreplicates (n=2) of LshC2c2 and LwC2c2. The boxes extend from the first to third quartiles, and the whiskers represent 1.5x the interquartile range. The mean is indicated by a red horizontal bar. (J) Comparison of LshC2c2 and LwC2c2 RNA cleavage activity using purified proteins. LwaCas13a has higher RNase activity than LshCas13a. 5'-labeled RNA targets were incubated with C2c2 and the corresponding crRVA for 30 minutes, and the products were analyzed by gel electrophoresis. (K) LwaCas13a can process CRISPR array transcripts from the L. wadeii CRISPR locus. The two-spacer Lwa array was incubated with LwC2c2 for 30 minutes, and the reactions were then separated by gel electrophoresis. (L) Schematic of the mammalian LwaCas13a constructs evaluated and imaging showing the localization and expression of each design. Scale bar, 10 μm. (M) Schematic of the mammalian reporter system used to evaluate knockdown by luciferase protein readout. (N) Knockdown of Gaussia luciferase (Gluc) using engineered variants of LwaCas13a. Guide and shRNA sequences are shown at the top.(O) Knockdown of three different endogenous transcripts by LwaCas13a compared to the corresponding RNAi constructs. (P) Schematic of LwaCas13a knockdown of transcripts in rice (Oryza sativa) protoplasts. (Q) LwaCas13a knockdown of three transcripts in rice (Oryza sativa) protoplasts using three targeting and non-targeting guides for each transcript. All values are mean ± SEM with n=3 unless otherwise noted. [Figure 4A-4B] Figure 4A-4B. Normalized luciferase protein expression by various gRNAs against Gluc and C2c2 orthologs fused to NLS, NES, or untagged. (A) The C2c2 ortholog is Leptotrichia wadei F0279 (Lw2). The spacer sequences used in these experiments are: (Guide 1) ATCAGGGCAAACAGAACTTTGACTCCCA; (Guide 2) AGATCCGTGGTCGCGAAGTTGCTGGCCA; (Guide 3) TCGCCTTCGTAGGTGTGGCAGCGTCCTG; and (Guide NP) TAGATTGCTGTTCTACCAAGTAATCCAT. (B) The C2c2 ortholog is Listeria newyorkensis FSL M6-0635 (LbFSL). The spacer sequences used in these experiments were (guide 1) TCGCCTTCGTAGGTGTGGCAGCGTCCTG and (guide NT) tagattgctgttctaccaagtaatccat. [Figure 5A-5B]Figure 5A-B. Normalized protein expression of GFP by various gRNAs against EGPF and by C2c2 orthologs fused to NLSs, NESs, or untagged. (A) The C2c2 ortholog is Leptotrichia wadei F0279 (Lw2). The spacer sequences used in these experiments were: (Guide 1) tgaacagctcctcgcccttgctcaccat; (Guide 2) tcagcttgccgtaggtggcatcgccctc; (Guide 3) gggtagcggctgaagcactgcacgccgt; (Guide 4) ggtcttgtagttgccgtcgtccttgaag; (Guide 5) tactccagcttgtgccccaggatgttgc; (Guide 6) cacgctgccgtcctcgatgttgtggcgg; (Guide 7) tctttgctcagggcggactgggtgctca; (Guide 8) gacttgtacagctcgtccatgccgagag; and (Guide NT) tagattgctgttctaccaagtaatccat. (B) The C2c2 ortholog is Listeria newyorkensis FSL M6-0635 (LbFSL). The spacer sequences used in these experiments are: (Guide 2) tcagcttgccgtaggtggcatcgccctc; (Guide 3) gggtagcggctgaagcactgcacgccgt; (Guide 4) ggtcttgtagttgccgtcgtccttgaag; (Guide 5) tactccagcttgtgccccaggatgttgc; (Guide 6) cacgctgccgtcctcgatgttgtggcgg; (Guide 7) tctttgctcagggcggactgggtgctca; (Guide 8) gacttgtacagctcgtccatgccgagag; and (Guide NT) tagattgctgttctaccaagtaatccat. [Figures 6A-6B]Figure 6A-6E. Engineering and optimization of LwaCas13a for mammalian knockdown. (A) Knockdown of Gluc transcripts by LwCas13a using various guides transfected into A375 cells. (B) Knockdown of Gaussia luciferase (Gluc) using engineered variants of LwaCas13a. (C) Knockdown of Gaussia luciferase (Gluc) using LwaCas13a and various lengths of the Gluc crRNA 1 spacer. (D) Knockdown of KRAS transcripts by LwaCas13a using various guides. (E) Knockdown of PPIB transcripts by LwaCas13a using various guides. [Figure 6C-6D] Figure 6A-6E. Engineering and optimization of LwaCas13a for mammalian knockdown. (A) Knockdown of Gluc transcripts by LwCas13a using various guides transfected into A375 cells. (B) Knockdown of Gaussia luciferase (Gluc) using engineered variants of LwaCas13a. (C) Knockdown of Gaussia luciferase (Gluc) using LwaCas13a and various lengths of the Gluc crRNA 1 spacer. (D) Knockdown of KRAS transcripts by LwaCas13a using various guides. (E) Knockdown of PPIB transcripts by LwaCas13a using various guides. [Figure 6E] Figure 6A-6E. Engineering and optimization of LwaCas13a for mammalian knockdown. (A) Knockdown of Gluc transcripts by LwCas13a using various guides transfected into A375 cells. (B) Knockdown of Gaussia luciferase (Gluc) using engineered variants of LwaCas13a. (C) Knockdown of Gaussia luciferase (Gluc) using LwaCas13a and various lengths of the Gluc crRNA 1 spacer. (D) Knockdown of KRAS transcripts by LwaCas13a using various guides. (E) Knockdown of PPIB transcripts by LwaCas13a using various guides. [Figure 7]Figure 7. Normalized protein expression of target genes by C2c2 Leptotrichia wadei F0279 fused with gRNA and NES for each target gene. The gRNAs for each target gene were ctgctgccacagaccgagaggcttaaaa (CTNNB1); tccttgattacacgatggaatttgctgt (PPIB); tcaaggtggggtcacaggagaagccaaa (mAPK14); atgataatgcaatagcaggacaggatga (CXCR4); gcgtgagccaccgcgcctggccggctgt (TINCR); ccagctgcagatgctgcagtttttggc g(PCAT1);ctggaaatggaagatgccggcatagcca(CAPN1);gatgacacctcacacggaccacccctag(LETMD1);taatactgctccagatatgggtgggcca(MAP K14); catgaagaccgagttatagaatactata (RB1); ggtgaaatattctccatccagtggtttc (TP53); and aatttctcgaactaatgtatagaaggca (KRAS). [Figure 8] Figure 8. Normalized luciferase protein expression by various gRNAs against Gluc and by C2c2 orthologs fused to NLSs, NESs, or untagged. The C2c2 orthologs are Leptotrichia shahii (Lsh), Leptotrichia wadei F0279 (Lw2), Clostridium aminophilum (Ca), Listeria newyorkensis FSL M6-0635 (LbFSL), Lachnospiraceae bacterium NK4A179 (LbNk), and Lachnospiraceae bacterium MA2020 (LbM). [Figure 9]Figure 9. Lw2 functions in multiple cell lines. Normalized luciferase activity for 293FT cells is shown. [Figures 10A-10B] Figure 10A-10B. Relative luciferase protein expression by gRNA against Gluc and by catalytically inactive C2c2 ortholog. (A) dC2c2 from Leptotrichia wadei (LwC2c2) was fused to EIF4E, EIF4E and NES, or untagged. (B) dC2c2 from Listeria newyorkensis FSL M6-0635 (LbFSL) was fused to NLS and EIF4E or untagged. [Figures 11A-11B] Figure 11A-11C. When cells were treated with NaAsO2, the cellular localization of Leptotrichia wadei C2c2, which targets β-actin, was localized to stress granules; [Figure 12] Figure 12. RNA knockdown in mammalian cells. C2c2 outperforms shRNA at two target sites. [Figure 13] Figure 13. Increasing crRNA transfection dose leads to increased knockdown. [Figure 14] Figure 14. Lw2C2c2 with an NES (NES-Lw2C2c2) efficiently cleaves tRNA and U6 knockdown RNA. [Figures 15A-15B] Figure 15A-B. (A) Protein transfection amount saturates knockdown. (B) Knockdown of Gluc transcripts by Gluc crRNA 1 and varying amounts of transfected LwaCas13a plasmid. [Figure 16] Figure 16. U6 drive DR-spacer-DR-spacer targeting construct [Figure 17] Figure 17. C2c2 outperforms optimized shRNAs against corresponding targets on endogenous genes. [Figure 18]Figure 18. dLw2C2c2-EIF4E fusions can upregulate the translation of three genes; protein levels as measured by band intensity on Western blots. [Figure 19] Figure 19. Knockdown of individual transcripts. Lw2 shows lower variability and higher specificity. [Figure 20] Figure 20. RNA knockdown in mammalian cells. [Figure 21] Figure 21. A superfolder derivative of GFP (sfGFP) improves imaging. Comparing C2c2-mCherry and sfGFP-C2c2 fusion proteins in HEK293FT cells and mouse embryonic stem cells (mESCs). [Figure 22A] Figure 22A. msfGFP-C2c2 enhances knockdown. Luciferase knockdown by various fusion proteins of C2c2 with mCherry or msfGFP, including an NLS or NES, is shown. Figures 22B-22D illustrate a protein tagging system for regulating transcription by C2c2, showing the transcription initiation factor-linked scFv element (SunTag) that binds to a short peptide sequence contained in the modified C2c2. Figures 22C-22D show the transcriptional efficacy of this system. [Figure 22B] Figure 22A. msfGFP-C2c2 enhances knockdown. Luciferase knockdown by various fusion proteins of C2c2 with mCherry or msfGFP, including an NLS or NES, is shown. Figures 22B-22D illustrate a protein tagging system for regulating transcription by C2c2, showing the transcription initiation factor-linked scFv element (SunTag) that binds to a short peptide sequence contained in the modified C2c2. Figures 22C-22D show the transcriptional efficacy of this system. [Figures 22C-22D]Figure 22A. msfGFP-C2c2 enhances knockdown. Luciferase knockdown by various fusion proteins of C2c2 with mCherry or msfGFP, including an NLS or NES, is shown. Figures 22B-22D illustrate a protein tagging system for regulating transcription by C2c2, showing the transcription initiation factor-linked scFv element (SunTag) that binds to a short peptide sequence contained in the modified C2c2. Figures 22C-22D show the transcriptional efficacy of this system. [Figures 23A-23B] Figure 23A-B. Splitting of a reporter protein (GFP, Venus, Cre, etc.). Each part is fused to C2c2 (see slide for a schematic including a caspase, for example). By designing two guides that target a transcript close to each other, the split protein is reconstituted in the presence of that transcript. [Figure 24A] Figures 24A-C. Rapid RNA detection by C2c2 collateral RNase activity. (A) Left panel: Activation of C2c2 collateral non-specific RNase activity by target RNA complementary to the guide sequence of C2c2 crRNA results in reporter cleavage. Right panel: In the absence of target RNA, collateral non-specific RNase activity is not induced, and therefore there is no reporter cleavage. (B) Schematic of the assay to detect the activity of the collateral effect by LwaCas13a. (C) Gel electrophoresis of collateral RNA targets after incubation with LwaCas13a-crRNA complexes in the presence or absence of target RNA. [Figures 24B-24C]Figures 24A-C. Rapid RNA detection by C2c2 collateral RNase activity. (A) Left panel: Activation of C2c2 collateral non-specific RNase activity by target RNA complementary to the guide sequence of C2c2 crRNA results in reporter cleavage. Right panel: In the absence of target RNA, collateral non-specific RNase activity is not induced, and therefore there is no reporter cleavage. (B) Schematic of the assay to detect the activity of the collateral effect by LwaCas13a. (C) Gel electrophoresis of collateral RNA targets after incubation with LwaCas13a-crRNA complexes in the presence or absence of target RNA. [Figure 25] Figure 25. Lentivirus detection via the collateral effect using various concentrations of C2c2. [Figure 26] Figure 26. Changes in lentivirus detection over time due to collateral effects. [Figure 27] Figure 27. Detection of rare RNA species by C2c2 collateral RNase activity. Nonspecific off-target RNase activity increases with increasing target concentration and target cleavage. [Figure 28A] Figures 28A-B. Alignment of Lw2C2c2 with LbuC2c2. [Figure 28B] Figures 28A-B. Alignment of Lw2C2c2 with LbuC2c2. [Figures 29A-29C]Figures 29A-C. (A) βLwC2c2 targeted luciferase mRNA knockdown with a single base pair mismatch in the spacer sequence. (B) LwC2c2 targeted CXCR4 mRNA knockdown with a single base pair mismatch in the spacer sequence. (C) LwC2c2 targeted luciferase mRNA knockdown with single and consecutive double base pair mismatches in the spacer sequence. Figures 29D-A-C. (A) Specificity of luciferase reporter mRNA knockdown. C2c2 is more specific than RNAi for luciferase reporter mRNA knockdown. Left: Expression levels as log2 (transcripts per million (TPM)) values of all genes detected in RNA-seq libraries of non-targeting transfected controls (x-axis of all graphs) compared to luciferase knockdown conditions for shRNA. Right: Expression levels, expressed as log2 (transcripts per million (TPM)), of all detected genes in RNA-seq libraries of non-targeting transfected controls (x-axis of all graphs) compared to luciferase knockdown conditions for C2c2. The target luciferase transcript targeted for knockdown is indicated by a red dot. Averages from n=3 biological replicates are shown. (B) Specificity of endogenous KRAS mRNA knockdown. C2c2 is more specific than RNAi for knockdown of endogenous KRAS. Target transcripts are indicated by red dots. (C) Specificity of endogenous PPIB mRNA knockdown. C2c2 is more specific than RNAi for knockdown of endogenous PPIB. Target transcripts are indicated by red dots. Figures 29E-G. (E) Differential gene expression analysis of six RNA-seq libraries (each with three biological replicates) comparing LwaCas13a knockdown with shRNA knockdown in three different genes. A gene is considered significantly differentially expressed if it has an average fold change >2 or <0.75 compared to the non-targeting control and a false discovery rate (FDR) <0.10. (F) Mean knockdown levels quantified for target genes from RNA-seq libraries.(G) Left: Luciferase knockdown in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection prior to measuring cell growth. Center: Cell viability in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection. Right: GFP fluorescence in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection. All values are mean ± SEM for n=3. [Figure 29D-A-29D-B]Figures 29A-C. (A) βLwC2c2 targeted luciferase mRNA knockdown with a single base pair mismatch in the spacer sequence. (B) LwC2c2 targeted CXCR4 mRNA knockdown with a single base pair mismatch in the spacer sequence. (C) LwC2c2 targeted luciferase mRNA knockdown with single and consecutive double base pair mismatches in the spacer sequence. Figures 29D-A-C. (A) Specificity of luciferase reporter mRNA knockdown. C2c2 is more specific than RNAi for luciferase reporter mRNA knockdown. Left: Expression levels as log2 (transcripts per million (TPM)) values of all genes detected in RNA-seq libraries of non-targeting transfected controls (x-axis of all graphs) compared to luciferase knockdown conditions for shRNA. Right: Expression levels, expressed as log2 (transcripts per million (TPM)), of all detected genes in RNA-seq libraries of non-targeting transfected controls (x-axis of all graphs) compared to luciferase knockdown conditions for C2c2. The target luciferase transcript targeted for knockdown is indicated by a red dot. Averages from n=3 biological replicates are shown. (B) Specificity of endogenous KRAS mRNA knockdown. C2c2 is more specific than RNAi for knockdown of endogenous KRAS. Target transcripts are indicated by red dots. (C) Specificity of endogenous PPIB mRNA knockdown. C2c2 is more specific than RNAi for knockdown of endogenous PPIB. Target transcripts are indicated by red dots. Figures 29E-G. (E) Differential gene expression analysis of six RNA-seq libraries (each with three biological replicates) comparing LwaCas13a knockdown with shRNA knockdown in three different genes. A gene is considered significantly differentially expressed if it has an average fold change >2 or <0.75 compared to the non-targeting control and a false discovery rate (FDR) <0.10. (F) Mean knockdown levels quantified for target genes from RNA-seq libraries.(G) Left: Luciferase knockdown in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection prior to measuring cell growth. Center: Cell viability in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection. Right: GFP fluorescence in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection. All values are mean ± SEM for n=3. [Figures 29E-29G]Figures 29A-C. (A) βLwC2c2 targeted luciferase mRNA knockdown with a single base pair mismatch in the spacer sequence. (B) LwC2c2 targeted CXCR4 mRNA knockdown with a single base pair mismatch in the spacer sequence. (C) LwC2c2 targeted luciferase mRNA knockdown with single and consecutive double base pair mismatches in the spacer sequence. Figures 29D-A-C. (A) Specificity of luciferase reporter mRNA knockdown. C2c2 is more specific than RNAi for luciferase reporter mRNA knockdown. Left: Expression levels as log2 (transcripts per million (TPM)) values of all genes detected in RNA-seq libraries of non-targeting transfected controls (x-axis of all graphs) compared to luciferase knockdown conditions for shRNA. Right: Expression levels, expressed as log2 (transcripts per million (TPM)), of all detected genes in RNA-seq libraries of non-targeting transfected controls (x-axis of all graphs) compared to luciferase knockdown conditions for C2c2. The target luciferase transcript targeted for knockdown is indicated by a red dot. Averages from n=3 biological replicates are shown. (B) Specificity of endogenous KRAS mRNA knockdown. C2c2 is more specific than RNAi for knockdown of endogenous KRAS. Target transcripts are indicated by red dots. (C) Specificity of endogenous PPIB mRNA knockdown. C2c2 is more specific than RNAi for knockdown of endogenous PPIB. Target transcripts are indicated by red dots. Figures 29E-G. (E) Differential gene expression analysis of six RNA-seq libraries (each with three biological replicates) comparing LwaCas13a knockdown with shRNA knockdown in three different genes. A gene is considered significantly differentially expressed if it has an average fold change >2 or <0.75 compared to the non-targeting control and a false discovery rate (FDR) <0.10. (F) Mean knockdown levels quantified for target genes from RNA-seq libraries.(G) Left: Luciferase knockdown in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection prior to measuring cell growth. Center: Cell viability in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection. Right: GFP fluorescence in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection. All values are mean ± SEM for n=3. [Figure 30A-30B] Figures 30A-30E. Comparison of RNA knockdown levels and specificity. C2c2 demonstrates increased specificity while maintaining similar levels of knockdown as RNAi. (A): Number of significantly differentially regulated genes in each RNA-seq library analyzed. Differential gene expression analysis of six RNA-seq libraries (each with three biological replicates) comparing LwaCas13a knockdown with shRNA knockdown in three different genes. A gene was considered significantly differentially expressed if it had an average fold change >2 or <0.75 compared to the non-targeting control and a false discovery rate (FDR) <0.10. (B): Mean knockdown levels quantified for target genes from the RNA-seq libraries. Normalized expression is calculated as the TPM of the target gene in the targeting condition divided by the TPM of the target gene in the non-targeting condition. Plots represent cumulative data from the RNA-seq plots in Figure 28. Off-targets are determined as genes that are significantly unregulated (>2-fold) or downregulated (<0.75-fold). (C) Luciferase knockdown in cells transfected with LwaCas13a containing non-selectable and blastcidin-selectable versions of C2c2 for 72 hours. (D) Cell viability in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection. (E) GFP fluorescence in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection. [Figures 30C-30E]Figures 30A-30E. Comparison of RNA knockdown levels and specificity. C2c2 demonstrates increased specificity while maintaining similar levels of knockdown as RNAi. (A): Number of significantly differentially regulated genes in each RNA-seq library analyzed. Differential gene expression analysis of six RNA-seq libraries (each with three biological replicates) comparing LwaCas13a knockdown with shRNA knockdown in three different genes. A gene was considered significantly differentially expressed if it had an average fold change >2 or <0.75 compared to the non-targeting control and a false discovery rate (FDR) <0.10. (B): Mean knockdown levels quantified for target genes from the RNA-seq libraries. Normalized expression is calculated as the TPM of the target gene in the targeting condition divided by the TPM of the target gene in the non-targeting condition. Plots represent cumulative data from the RNA-seq plots in Figure 28. Off-targets are determined as genes that are significantly unregulated (>2-fold) or downregulated (<0.75-fold). (C) Luciferase knockdown in cells transfected with LwaCas13a containing non-selectable and blastcidin-selectable versions of C2c2 for 72 hours. (D) Cell viability in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection. (E) GFP fluorescence in cells transfected with LwaCas13a for 72 hours with and without antibiotic selection. [Figure 31] Figure 31. C2c2 has multiplex knockdown potential. crRNAs were designed to target PPIB, CXCR4, KRAS, TINCR, and PCAT. The top panel shows the normalized expression levels of color-coded mRNAs in the same order from left to right as the multiplexed crRNA transcripts shown. The bottom panel compares single-administered guides with multiplexed and pooled guides. [Figures 32A-32D]Figures 32A-D. Tiling guides along the length of gLuc or cLuc reveals retargeting capabilities and vulnerable regions on the transcript. (A) Schematic of LwaCas13a arrayed screening. (B) Knockdown efficiency of Gaussia luciferase mRNA by LwC2c2 with 186 guides evenly tiled across the length of the transcript. (C) Arrayed knockdown screen of Cypridina luciferase by LwC2c2 with guides evenly tiled across the length of the transcript. (D) Validation of the top three crRNAs from the arrayed knockdown screen by shRNA comparison. [Figure 33A] Figures 33A-B. RNA immunoprecipitation (RIP) shows enrichment of target binding. dC2c2-msfGFP containing an HA tag resulted in precipitation of protein-RNA complexes. Bound RNA was quantified by qPCR, and enrichment was determined relative to input RNA (no immunoprecipitation sample) and immunoprecipitation with a negative control antibody. Panel A shows enrichment of the β-actin RNA target. Panel B shows enrichment of the luciferase target. [Figure 33B] Figures 33A-B. RNA immunoprecipitation (RIP) shows enrichment of target binding. dC2c2-msfGFP containing an HA tag resulted in precipitation of protein-RNA complexes. Bound RNA was quantified by qPCR, and enrichment was determined relative to input RNA (no immunoprecipitation sample) and immunoprecipitation with a negative control antibody. Panel A shows enrichment of the β-actin RNA target. Panel B shows enrichment of the luciferase target. [Figure 34] Figure 34. C2c2 imaging of β-actin in 3T3L1 fibroblasts. Under targeting conditions (top frame), the C2c2 complex binds to actin mRNA and leaves the nucleus, revealing cytoplasmic β-actin mRNA. Under non-targeting conditions, NLS-tagged C2c2 is observed in the nucleus. The left and right columns correspond to different fields of view. [Figure 35]Figure 35. C2c2 imaging of β-actin in HEK293FT cells. Under targeting conditions (top frame), the C2c2 complex binds to actin mRNA and leaves the nucleus, revealing cytoplasmic β-actin mRNA. Under non-targeting conditions, NLS-tagged C2c2 is observed in the nucleus. The left and right columns correspond to different fields of view. [Figure 36] Figure 36. In vitro characterization of RNA cleavage kinetics of Lw2Cas13a. Denaturing gel after 0.5 h of RNA cleavage of 5'-end-labeled Target 1 using serially diluted LwaCas13a-crRNA complexes in half-logarithmic steps. [Figure 37] Figure 37. In vitro characterization of Lw2Cas13a RNA cleavage kinetics. Denutaring gel time series of Lw2Cas13a ssRNA cleavage using 5'-end labeled Target 1. [Figure 38] Figure 38. Lw2C2c2 and crRNA mediate RNA-guided ssRNA cleavage. Denaturing gel showing crRNA-mediated ssRNA cleavage by Lw2 C2c2 after 1 hour of incubation. The ssRNA target is 5'-labeled with IRDye 800. Cleavage requires the presence of crRNA and is abolished by the addition of EDTA. [Figure 39A-39B] Figures 39A-B. Lw2C2c2 and crRNA mediate RNA-guided ssRNA cleavage. (A) Schematic of the ssRNA substrate targeted by the crRNA. The protospacer region is highlighted in blue, and the PFS is indicated by a magenta bar. (B) Denaturing gel showing the requirement for a PFS after 3 hours of incubation. Four ssRNA substrates, identical except for the PFS (indicated by a magenta X in the schematic), were used in the in vitro cleavage reaction. ssRNA cleavage activity depends on the nucleotide immediately 3' to the target site. [Figure 40] Figure 40. Direct repeat length affects the RNA-guided RNase activity of Lw2C2c2. Denaturing gel showing crRNA-guided cleavage of ssRNA 1 as a function of direct repeat length after 3 hours of incubation. [Figure 41]Figure 41. Spacer length affects the RNA-guided RNase activity of Lw2C2c2. Denaturing gel showing crRNA-guided cleavage of ssRNA 1 as a function of spacer length after 3 hours of incubation. [Figures 42A-42C] Figures 42A-42C. Lw2C2c2 cleavage sites are determined by the secondary structure and sequence of the target RNA. (A) Denaturing gel showing ssRNA 4 and ssRNA 5 after incubation with LwCas13a and crRNA 1. (B) ssRNA 4 (blue) and (C) ssRNA 5 (green) share the same protospacer but are flanked by different sequences. Despite the identical protospacer, the different flanking sequences result in distinct cleavage patterns. [Figures 43A-43C] Figures 43A-C. (A) Schematic of ssRNA 4 modified with homopolymeric stretches in the highlighted loop (red) for each of the four possible nucleotides. The crRNA spacer sequence is highlighted in blue. (B) Denaturing gel showing Lw2C2c2-crRNA-mediated cleavage after 3 hours of incubation for each of the four possible homopolymeric targets. Lw2C2c2 cleaves C and U. (C) LwaCas13a can process pre-crRNA from the L. wadeii CRISPR-Cas locus. [Figures 44A-44C]Figure 44A-K. Catalytically inactive LwaCas13a (dCas13a) is capable of binding to transcripts in mammalian cells. (A) Schematic of the dCas13a-GFP construct used for imaging and assessing dCas13a binding. (B) Schematic of RNA immunoprecipitation to quantify dCas13a binding. (C) dCas13a targeting gLuc transcripts is significantly enriched compared to non-targeting controls. Quantification of dC2c2-msfGFP binding to Gaussian luciferase mRNA by RIP normalized to either control antibody or input lysate. Values are normalized to the non-targeting guide. (n=3, *, p<0.05; ****, p<0.0001 by t-test). (D) Schematic of the dCas13a-GFP-KRAB construct used for negative feedback imaging. In the absence of target and guide, the reporter protein inhibits its own transcription. (E) Gluc, Cluc, PPIB, and KRAS knockdown partially correlates with target accessibility as measured by predicted folding of the transcript. (F) Comparison between the localization of dCas13-GFP and dCas13a-GFP-KRAB constructs for imaging b-actin. (G) Representative images of dCas13a-GFP-KRAB imaging with multiple guides targeting b-actin in HEK293 cells. (H) Representative images of dCas13a-GFP-KRAB imaging with multiple guides targeting ACTB. (I) Quantification of cytoplasmic translocation of dCas13a-GFP-KRAB as measured by the nuclear-to-total-cell signal ratio. (J) Representative fixed immunofluorescence image of 293FT cells treated with 400 μM sodium arsenite. Stress granules are demonstrated by staining for the marker G3BP1. Scale bar, 5 μm. Scale bar, 5 μm. (K) Colocalization of G3BP1 and dCas13a-GFP-KRAB quantified per cell by Pearson correlation. All values are mean ± SEM for n=3. ****p<0.0001; ***p<0.001; **p<0.01; *p<0.05. ns=not significant.A one-tailed Student's t-test was used for comparisons in (a), and a two-tailed Student's t-test was used for comparisons in (I) and (K). [Fig. 44D-44E]Figure 44A-K. Catalytically inactive LwaCas13a (dCas13a) is capable of binding to transcripts in mammalian cells. (A) Schematic of the dCas13a-GFP construct used for imaging and assessing dCas13a binding. (B) Schematic of RNA immunoprecipitation to quantify dCas13a binding. (C) dCas13a targeting gLuc transcripts is significantly enriched compared to non-targeting controls. Quantification of dC2c2-msfGFP binding to Gaussian luciferase mRNA by RIP normalized to either control antibody or input lysate. Values are normalized to the non-targeting guide. (n=3, *, p<0.05; ****, p<0.0001 by t-test). (D) Schematic of the dCas13a-GFP-KRAB construct used for negative feedback imaging. In the absence of target and guide, the reporter protein inhibits its own transcription. (E) Gluc, Cluc, PPIB, and KRAS knockdown partially correlates with target accessibility as measured by predicted folding of the transcript. (F) Comparison between the localization of dCas13-GFP and dCas13a-GFP-KRAB constructs for imaging b-actin. (G) Representative images of dCas13a-GFP-KRAB imaging with multiple guides targeting b-actin in HEK293 cells. (H) Representative images of dCas13a-GFP-KRAB imaging with multiple guides targeting ACTB. (I) Quantification of cytoplasmic translocation of dCas13a-GFP-KRAB as measured by the nuclear-to-total-cell signal ratio. (J) Representative fixed immunofluorescence image of 293FT cells treated with 400 μM sodium arsenite. Stress granules are demonstrated by staining for the marker G3BP1. Scale bar, 5 μm. Scale bar, 5 μm. (K) Colocalization of G3BP1 and dCas13a-GFP-KRAB quantified per cell by Pearson correlation. All values are mean ± SEM for n=3. ****p<0.0001; ***p<0.001; **p<0.01; *p<0.05. ns=not significant.A one-tailed Student's t-test was used for comparisons in (a), and a two-tailed Student's t-test was used for comparisons in (I) and (K). [Fig. 44F-44G]Figure 44A-K. Catalytically inactive LwaCas13a (dCas13a) is capable of binding to transcripts in mammalian cells. (A) Schematic of the dCas13a-GFP construct used for imaging and assessing dCas13a binding. (B) Schematic of RNA immunoprecipitation to quantify dCas13a binding. (C) dCas13a targeting gLuc transcripts is significantly enriched compared to non-targeting controls. Quantification of dC2c2-msfGFP binding to Gaussian luciferase mRNA by RIP normalized to either control antibody or input lysate. Values are normalized to the non-targeting guide. (n=3, *, p<0.05; ****, p<0.0001 by t-test). (D) Schematic of the dCas13a-GFP-KRAB construct used for negative feedback imaging. In the absence of target and guide, the reporter protein inhibits its own transcription. (E) Gluc, Cluc, PPIB, and KRAS knockdown partially correlates with target accessibility as measured by predicted folding of the transcript. (F) Comparison between the localization of dCas13-GFP and dCas13a-GFP-KRAB constructs for imaging b-actin. (G) Representative images of dCas13a-GFP-KRAB imaging with multiple guides targeting b-actin in HEK293 cells. (H) Representative images of dCas13a-GFP-KRAB imaging with multiple guides targeting ACTB. (I) Quantification of cytoplasmic translocation of dCas13a-GFP-KRAB as measured by the nuclear-to-total-cell signal ratio. (J) Representative fixed immunofluorescence image of 293FT cells treated with 400 μM sodium arsenite. Stress granules are demonstrated by staining for the marker G3BP1. Scale bar, 5 μm. Scale bar, 5 μm. (K) Colocalization of G3BP1 and dCas13a-GFP-KRAB quantified per cell by Pearson correlation. All values are mean ± SEM for n=3. ****p<0.0001; ***p<0.001; **p<0.01; *p<0.05. ns=not significant.A one-tailed Student's t-test was used for comparisons in (a), and a two-tailed Student's t-test was used for comparisons in (I) and (K). [Fig. 44H-44K]Figure 44A-K. Catalytically inactive LwaCas13a (dCas13a) is capable of binding to transcripts in mammalian cells. (A) Schematic of the dCas13a-GFP construct used for imaging and assessing dCas13a binding. (B) Schematic of RNA immunoprecipitation to quantify dCas13a binding. (C) dCas13a targeting gLuc transcripts is significantly enriched compared to non-targeting controls. Quantification of dC2c2-msfGFP binding to Gaussian luciferase mRNA by RIP normalized to either control antibody or input lysate. Values are normalized to the non-targeting guide. (n=3, *, p<0.05; ****, p<0.0001 by t-test). (D) Schematic of the dCas13a-GFP-KRAB construct used for negative feedback imaging. In the absence of target and guide, the reporter protein inhibits its own transcription. (E) Gluc, Cluc, PPIB, and KRAS knockdown partially correlates with target accessibility as measured by predicted folding of the transcript. (F) Comparison between the localization of dCas13-GFP and dCas13a-GFP-KRAB constructs for imaging b-actin. (G) Representative images of dCas13a-GFP-KRAB imaging with multiple guides targeting b-actin in HEK293 cells. (H) Representative images of dCas13a-GFP-KRAB imaging with multiple guides targeting ACTB. (I) Quantification of cytoplasmic translocation of dCas13a-GFP-KRAB as measured by the nuclear-to-total-cell signal ratio. (J) Representative fixed immunofluorescence image of 293FT cells treated with 400 μM sodium arsenite. Stress granules are demonstrated by staining for the marker G3BP1. Scale bar, 5 μm. Scale bar, 5 μm. (K) Colocalization of G3BP1 and dCas13a-GFP-KRAB quantified per cell by Pearson correlation. All values are mean ± SEM for n=3. ****p<0.0001; ***p<0.001; **p<0.01; *p<0.05. ns=not significant.A one-tailed Student's t-test was used for comparisons in (a), and a two-tailed Student's t-test was used for comparisons in (I) and (K). [Figure 45] Figure 45. Detection of β-actin using C2c2-GFP imaging protein as shown in Figure 43. β-actin was imaged using dC2c2-eGFP-ZF-KRAB-NLS with a targeting guide (left panel) or a non-targeting guide (right panel). [Figure 46] Figure 46. Detection of stress granules using GFP-tagged G3BP1. [Figure 47] Figure 47. Detection of actin (acting) in stress granules and stress granule substructures ("cores"). Actin mRNA was imaged using C2c2-eGFP-ZF-KRAB with and without targeting constructs and with or without NaAsO2 treatment to stabilize the core substructure. Stress granules were labeled using anti-G3BP1. [Figure 48A-48B] Figures 48A-48F: dCas13a can image stress granule formation in live cells. (A) Schematic of imaging β-actin mRNA localization to stress granules upon sodium arsenite treatment using a negative feedback dCas13a-msfGFP-KRAB construct. (B) Representative fixed immunofluorescence images of 293FT cells treated with 400 μM sodium arsenite. dCas13a-msfGFP-KRAB transfected with a β-actin mRNA targeting guide localizes to stress granules. Representative images of fixed HEK293 cells immunostained with an antibody against the G3BP1 marker for stress granules are shown. (C) Quantification of stress granule localization by Pearson correlation analysis of approximately 20 cells per condition. (D) Quantification of stress granule localization by Manders colocalization analysis of approximately 20 cells per condition. (E) Representative images from live cell analysis of stress granule formation in response to 400 μM sodium arsenite treatment. (F) Quantification of stress granule formation in response to sodium arsenite treatment. [Fig. 48C-48D]Figures 48A-48F: dCas13a can image stress granule formation in live cells. (A) Schematic of imaging β-actin mRNA localization to stress granules upon sodium arsenite treatment using a negative feedback dCas13a-msfGFP-KRAB construct. (B) Representative fixed immunofluorescence images of 293FT cells treated with 400 μM sodium arsenite. dCas13a-msfGFP-KRAB transfected with a β-actin mRNA targeting guide localizes to stress granules. Representative images of fixed HEK293 cells immunostained with an antibody against the G3BP1 marker for stress granules are shown. (C) Quantification of stress granule localization by Pearson correlation analysis of approximately 20 cells per condition. (D) Quantification of stress granule localization by Manders colocalization analysis of approximately 20 cells per condition. (E) Representative images from live cell analysis of stress granule formation in response to 400 μM sodium arsenite treatment. (F) Quantification of stress granule formation in response to sodium arsenite treatment. [Fig. 48E-48F] Figures 48A-48F: dCas13a can image stress granule formation in live cells. (A) Schematic of imaging β-actin mRNA localization to stress granules upon sodium arsenite treatment using a negative feedback dCas13a-msfGFP-KRAB construct. (B) Representative fixed immunofluorescence images of 293FT cells treated with 400 μM sodium arsenite. dCas13a-msfGFP-KRAB transfected with a β-actin mRNA targeting guide localizes to stress granules. Representative images of fixed HEK293 cells immunostained with an antibody against the G3BP1 marker for stress granules are shown. (C) Quantification of stress granule localization by Pearson correlation analysis of approximately 20 cells per condition. (D) Quantification of stress granule localization by Manders colocalization analysis of approximately 20 cells per condition. (E) Representative images from live cell analysis of stress granule formation in response to 400 μM sodium arsenite treatment. (F) Quantification of stress granule formation in response to sodium arsenite treatment. [Figure 49A-49B]Figures 49A-49O. LwaCas13a can be reprogrammed to target endogenous mammalian coding and non-coding RNA targets. (A) Arrayed knockdown screen of 93 guides evenly tiled across the KRAS transcript. (B) Arrayed knockdown screen of 93 guides evenly tiled across the PPIB transcript. (C) Schematic of LwaCas13a arrayed screening. (D) Arrayed knockdown screen of 186 guides evenly tiled across the Gluc transcript. (E) Arrayed knockdown screen of 93 guides evenly tiled across the Cypridinia luciferase (Cluc) transcript. (F) Arrayed knockdown screen of 93 guides evenly tiled across the KRAS transcript. (G) Arrayed knockdown screen of 93 guides evenly tiled across the PPIB transcript. (H) Validation of the top three guides from the endogenous arrayed knockdown screen by shRNA comparison. All values are means ± SEM, n=3. ***p<0.001; **p<0.01. Comparisons were performed using a two-tailed Student's t-test. (I) An arrayed knockdown screen of 93 guides evenly tiled across the MALAT1 transcript. (J) Validation of the top three guides from the endogenous arrayed MALAT1 knockdown screen by shRNA comparison. (K) Multiplexed delivery of five guides in a CRISPR array to five different endogenous genes under the control of a single promoter enables robust knockdown. (L) Multiplexed delivery of three guides to three different endogenous genes or constructs replacing each guide with a non-targeting sequence demonstrates specific knockdown of the targeted gene. All values are means ± SEM, n=3. (M) Knockdown of three different endogenous transcripts by LwaCas13a compared to the corresponding RNAi construct. (N) LwaCas13a is capable of knockdown of the nuclear lncRNA transcript MALAT1. (O) Multiplex delivery of three guides to three different endogenous genes or constructs replacing each of the crRNAs with a non-targeting sequence demonstrates specific knockdown of the targeted genes. [Fig. 49C-49F] Figures 49A-49O. LwaCas13a can be reprogrammed to target endogenous mammalian coding and non-coding RNA targets. (A) Arrayed knockdown screen of 93 guides evenly tiled across the KRAS transcript. (B) Arrayed knockdown screen of 93 guides evenly tiled across the PPIB transcript. (C) Schematic of LwaCas13a arrayed screening. (D) Arrayed knockdown screen of 186 guides evenly tiled across the Gluc transcript. (E) Arrayed knockdown screen of 93 guides evenly tiled across the Cypridinia luciferase (Cluc) transcript. (F) Arrayed knockdown screen of 93 guides evenly tiled across the KRAS transcript. (G) Arrayed knockdown screen of 93 guides evenly tiled across the PPIB transcript. (H) Validation of the top three guides from the endogenous arrayed knockdown screen by shRNA comparison. All values are means ± SEM, n=3. ***p<0.001; **p<0.01. Comparisons were performed using a two-tailed Student's t-test. (I) An arrayed knockdown screen of 93 guides evenly tiled across the MALAT1 transcript. (J) Validation of the top three guides from the endogenous arrayed MALAT1 knockdown screen by shRNA comparison. (K) Multiplexed delivery of five guides in a CRISPR array to five different endogenous genes under the control of a single promoter enables robust knockdown. (L) Multiplexed delivery of three guides to three different endogenous genes or constructs replacing each guide with a non-targeting sequence demonstrates specific knockdown of the targeted gene. All values are means ± SEM, n=3. (M) Knockdown of three different endogenous transcripts by LwaCas13a compared to the corresponding RNAi construct. (N) LwaCas13a is capable of knockdown of the nuclear lncRNA transcript MALAT1. (O) Multiplex delivery of three guides to three different endogenous genes or constructs replacing each of the crRNAs with a non-targeting sequence demonstrates specific knockdown of the targeted genes. [Figure 49G-49L]Figures 49A-49O. LwaCas13a can be reprogrammed to target endogenous mammalian coding and non-coding RNA targets. (A) Arrayed knockdown screen of 93 guides evenly tiled across the KRAS transcript. (B) Arrayed knockdown screen of 93 guides evenly tiled across the PPIB transcript. (C) Schematic of LwaCas13a arrayed screening. (D) Arrayed knockdown screen of 186 guides evenly tiled across the Gluc transcript. (E) Arrayed knockdown screen of 93 guides evenly tiled across the Cypridinia luciferase (Cluc) transcript. (F) Arrayed knockdown screen of 93 guides evenly tiled across the KRAS transcript. (G) Arrayed knockdown screen of 93 guides evenly tiled across the PPIB transcript. (H) Validation of the top three guides from the endogenous arrayed knockdown screen by shRNA comparison. All values are means ± SEM, n=3. ***p<0.001; **p<0.01. Comparisons were performed using a two-tailed Student's t-test. (I) An arrayed knockdown screen of 93 guides evenly tiled across the MALAT1 transcript. (J) Validation of the top three guides from the endogenous arrayed MALAT1 knockdown screen by shRNA comparison. (K) Multiplexed delivery of five guides in a CRISPR array to five different endogenous genes under the control of a single promoter enables robust knockdown. (L) Multiplexed delivery of three guides to three different endogenous genes or constructs replacing each guide with a non-targeting sequence demonstrates specific knockdown of the targeted gene. All values are means ± SEM, n=3. (M) Knockdown of three different endogenous transcripts by LwaCas13a compared to the corresponding RNAi construct. (N) LwaCas13a is capable of knockdown of the nuclear lncRNA transcript MALAT1. (O) Multiplex delivery of three guides to three different endogenous genes or constructs replacing each of the crRNAs with a non-targeting sequence demonstrates specific knockdown of the targeted genes. [Figure 49M-49O] Figures 49A-49O. LwaCas13a can be reprogrammed to target endogenous mammalian coding and non-coding RNA targets. (A) Arrayed knockdown screen of 93 guides evenly tiled across the KRAS transcript. (B) Arrayed knockdown screen of 93 guides evenly tiled across the PPIB transcript. (C) Schematic of LwaCas13a arrayed screening. (D) Arrayed knockdown screen of 186 guides evenly tiled across the Gluc transcript. (E) Arrayed knockdown screen of 93 guides evenly tiled across the Cypridinia luciferase (Cluc) transcript. (F) Arrayed knockdown screen of 93 guides evenly tiled across the KRAS transcript. (G) Arrayed knockdown screen of 93 guides evenly tiled across the PPIB transcript. (H) Validation of the top three guides from the endogenous arrayed knockdown screen by shRNA comparison. All values are means ± SEM, n=3. ***p<0.001; **p<0.01. Comparisons were performed using a two-tailed Student's t-test. (I) An arrayed knockdown screen of 93 guides evenly tiled across the MALAT1 transcript. (J) Validation of the top three guides from the endogenous arrayed MALAT1 knockdown screen by shRNA comparison. (K) Multiplexed delivery of five guides in a CRISPR array to five different endogenous genes under the control of a single promoter enables robust knockdown. (L) Multiplexed delivery of three guides to three different endogenous genes or constructs replacing each guide with a non-targeting sequence demonstrates specific knockdown of the targeted gene. All values are means ± SEM, n=3. (M) Knockdown of three different endogenous transcripts by LwaCas13a compared to the corresponding RNAi construct. (N) LwaCas13a is capable of knockdown of the nuclear lncRNA transcript MALAT1. (O) Multiplex delivery of three guides to three different endogenous genes or constructs replacing each of the crRNAs with a non-targeting sequence demonstrates specific knockdown of the targeted genes. [Figure 50] Figure 50. HEPN sequence motifs from 21 C2c2 orthologs. [Figure 51] Figure 51. HEPN sequence motifs from 33 C2c2 orthologs. [Figure 52] Figure 52. Exemplary linkage locations for effector domains. C2c2 proteins linked at the C- or N-terminus to heterologous functional domains (mCherry and msfGFP as shown) retain function as demonstrated by luciferase knockdown. [Figure 53A] Figures 53A-L. (A-K) Sequence alignment of C2c2 orthologs. (L) Sequence alignment of the HEPN domain. [Figure 53B] Figures 53A-L. (A-K) Sequence alignment of C2c2 orthologs. (L) Sequence alignment of the HEPN domain. [Figure 53C] Figures 53A-L. (A-K) Sequence alignment of C2c2 orthologs. (L) Sequence alignment of the HEPN domain. [Figure 53D] Figures 53A-L. (A-K) Sequence alignment of C2c2 orthologs. (L) Sequence alignment of the HEPN domain. [Figure 53E] Figures 53A-L. (A-K) Sequence alignment of C2c2 orthologs. (L) Sequence alignment of the HEPN domain. [Figure 53F] Figures 53A-L. (A-K) Sequence alignment of C2c2 orthologs. (L) Sequence alignment of the HEPN domain. [Figure 53G] Figures 53A-L. (A-K) Sequence alignment of C2c2 orthologs. (L) Sequence alignment of the HEPN domain. [Figure 53H] Figures 53A-L. (A-K) Sequence alignment of C2c2 orthologs. (L) Sequence alignment of the HEPN domain. [Figure 53I] Figures 53A-L. (A-K) Sequence alignment of C2c2 orthologs. (L) Sequence alignment of the HEPN domain. [Figure 53J] Figures 53A-L. (A-K) Sequence alignment of C2c2 orthologs. (L) Sequence alignment of the HEPN domain. [Figure 53K-53L] Figures 53A-L. (A-K) Sequence alignment of C2c2 orthologs. (L) Sequence alignment of the HEPN domain. [Figure 54] Figure 54. Tree alignment of C2c2 orthologs. [Figure 55A] Figures 55A-B. Tree alignment of C2c2 and Cas13b orthologs. [Figure 55B] Figures 55A-B. Tree alignment of C2c2 and Cas13b orthologs. [Figures 56A-56D] Figure 56A-F: Engineering and optimization of LwaCas13a for mammalian knockdown. (A) Knockdown of Gluc transcripts with Gluc crRNA 1 and varying amounts of transfected LwaCas13a plasmid. (B) Knockdown of Gluc transcripts with LwaCas13a and varying amounts of transfected Gluc crRNA 1 and 2 plasmid. (C) Knockdown of Gluc transcripts using crRNA expressed from either the U6 promoter or the tRNAVal promoter. (D) Knockdown of KRAS transcripts using crRNA expressed from either the U6 promoter or the tRNAVal promoter. (E) Knockdown of KRAS transcripts using guides expressed from either the U6 promoter or the tRNAVal promoter. (F) Arrayed knockdown screen of 93 guides evenly tiled across the XIST transcript. [Fig. 56E-56F]Figure 56A-F: Engineering and optimization of LwaCas13a for mammalian knockdown. (A) Knockdown of Gluc transcripts with Gluc crRNA 1 and varying amounts of transfected LwaCas13a plasmid. (B) Knockdown of Gluc transcripts with LwaCas13a and varying amounts of transfected Gluc crRNA 1 and 2 plasmid. (C) Knockdown of Gluc transcripts using crRNA expressed from either the U6 promoter or the tRNAVal promoter. (D) Knockdown of KRAS transcripts using crRNA expressed from either the U6 promoter or the tRNAVal promoter. (E) Knockdown of KRAS transcripts using guides expressed from either the U6 promoter or the tRNAVal promoter. (F) Arrayed knockdown screen of 93 guides evenly tiled across the XIST transcript. [Figures 57A-57C] Figures 57A-57E: Evaluation of LwaCas13a PFS preference and comparison with LshCas13a. (A) Sequence comparison tree of 15 Cas13a orthologs evaluated in this study. (B) Number of LshCas13a and LwaCas13a PFS sequences above the depletion threshold for various depletion thresholds. (C) Distribution of LshCas13a and LwaCas13a PFS enrichment in targeted samples, normalized to non-targeted samples. (D) Sequence logos and counts of remaining PFS sequences after LshCas13a cleavage at various enrichment cutoff thresholds. (E) Sequence logos and counts of remaining PFS sequences after LwaCas13a cleavage at various enrichment cutoff thresholds. [Fig. 57D-57E]Figures 57A-57E: Evaluation of LwaCas13a PFS preference and comparison with LshCas13a. (A) Sequence comparison tree of 15 Cas13a orthologs evaluated in this study. (B) Number of LshCas13a and LwaCas13a PFS sequences above the depletion threshold for various depletion thresholds. (C) Distribution of LshCas13a and LwaCas13a PFS enrichment in targeted samples, normalized to non-targeted samples. (D) Sequence logos and counts of remaining PFS sequences after LshCas13a cleavage at various enrichment cutoff thresholds. (E) Sequence logos and counts of remaining PFS sequences after LwaCas13a cleavage at various enrichment cutoff thresholds. [Figure 58A] Figures 58A-58D: LwaCas13a targeting efficiency is affected by accessibility along the transcript. (A) Column 1: Top knockdown guides are plotted by position along the target transcript. For Gluc, the top 20% of guides are selected, and for Cluc, KRAS, and PPIB, the top 30% of guides are selected. Column 2: Histograms of pairwise distances (blue) between adjacent top guides for each transcript compared to a random null distribution (red). The inset shows the cumulative frequency curves for these histograms. A shift from the blue curve (actual measured distance) to the red curve (null distribution of distances) to the left indicates that the guides are closer to each other than expected by chance. (B) Gluc, Cluc, PPIB, and KRAS knockdown partially correlates with target accessibility as measured by the predicted folding of the transcript. (C) Kernel density estimation plots showing the correlation between target accessibility (probability of a region being base-paired) and target expression after knockdown with LwaCas13a. (D) Column 1: Correlation between target expression and target accessibility (probability of a region to base pair) measured for various window sizes (W) and various k-mer lengths. Column 2: p-values for the correlation between target expression and target accessibility (probability of a region to base pair) measured for various window sizes (W) and various k-mer lengths. The color scale is designed so that p-values > 0.05 are shaded red and p-values < 0.05 are shaded blue. [Fig. 58B-58C] Figures 58A-58D: LwaCas13a targeting efficiency is affected by accessibility along the transcript. (A) Column 1: Top knockdown guides are plotted by position along the target transcript. For Gluc, the top 20% of guides are selected, and for Cluc, KRAS, and PPIB, the top 30% of guides are selected. Column 2: Histograms of pairwise distances (blue) between adjacent top guides for each transcript compared to a random null distribution (red). The inset shows the cumulative frequency curves for these histograms. A shift from the blue curve (actual measured distance) to the red curve (null distribution of distances) to the left indicates that the guides are closer to each other than expected by chance. (B) Gluc, Cluc, PPIB, and KRAS knockdown partially correlates with target accessibility as measured by the predicted folding of the transcript. (C) Kernel density estimation plots showing the correlation between target accessibility (probability of a region being base-paired) and target expression after knockdown with LwaCas13a. (D) Column 1: Correlation between target expression and target accessibility (probability of a region to base pair) measured for various window sizes (W) and various k-mer lengths. Column 2: p-values for the correlation between target expression and target accessibility (probability of a region to base pair) measured for various window sizes (W) and various k-mer lengths. The color scale is designed so that p-values > 0.05 are shaded red and p-values < 0.05 are shaded blue. [Figure 58D]Figures 58A-58D: LwaCas13a targeting efficiency is affected by accessibility along the transcript. (A) Column 1: Top knockdown guides are plotted by position along the target transcript. For Gluc, the top 20% of guides are selected, and for Cluc, KRAS, and PPIB, the top 30% of guides are selected. Column 2: Histograms of pairwise distances (blue) between adjacent top guides for each transcript compared to a random null distribution (red). The inset shows the cumulative frequency curves for these histograms. A shift from the blue curve (actual measured distance) to the red curve (null distribution of distances) to the left indicates that the guides are closer to each other than expected by chance. (B) Gluc, Cluc, PPIB, and KRAS knockdown partially correlates with target accessibility as measured by the predicted folding of the transcript. (C) Kernel density estimation plots showing the correlation between target accessibility (probability of a region being base-paired) and target expression after knockdown with LwaCas13a. (D) Column 1: Correlation between target expression and target accessibility (probability of a region to base pair) measured for various window sizes (W) and various k-mer lengths. Column 2: p-values for the correlation between target expression and target accessibility (probability of a region to base pair) measured for various window sizes (W) and various k-mer lengths. The color scale is designed so that p-values > 0.05 are shaded red and p-values < 0.05 are shaded blue. [Figure 59A-59B]Figures 59A-59K: Detailed evaluation of LwaCas13a sensitivity to mismatches in the crRNA:target duplex with varying spacer lengths. (A) Knockdown of KRAS assessed with crRNAs containing a single mismatch at various positions in the spacer sequence. (B) Knockdown of PPIB assessed with crRNAs containing a single mismatch at various positions in the spacer sequence. (C) Knockdown of Gluc assessed with guides containing non-consecutive double mismatches at various positions in the spacer sequence. The wild-type sequence is shown at the top, and the identities of the mismatches are indicated below. (D) Collateral cleavage activity for ssRNA 1 and 2 for various spacer lengths. (n = 4 technical replicates; bars represent mean ± sem). (E) Specificity ratios of the crRNAs tested in (D). Specificity ratios are calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n = 4 technical replicates; bars represent mean ± sem). (F) Collateral cleavage activity of 28-nt spacer crRNAs containing synthetic mismatches tiled along the spacer against ssRNAs 1 and 2 (n = 4 technical replicates; bars represent mean ± sem). (G) Specificity ratio of the crRNAs tested in (F). The specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage (n = 4 technical replicates; bars represent mean ± sem). (H) Collateral cleavage activity of 23-nt spacer crRNAs containing synthetic mismatches tiled along the spacer against ssRNAs 1 and 2 (n = 4 technical replicates; bars represent mean ± sem). (I) Specificity ratio of the crRNAs tested in (H). The specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n=4 technical replicates; bars represent mean ± sem.) (J) Collateral cleavage activity against ssRNAs 1 and 2 for 20 nt spacer crRNAs containing synthetic mismatches tiled along the spacer.(n = 4 technical replicates; bars represent mean ± sem). (K) Specificity ratio of the crRNAs tested in (J). Specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n = 4 technical replicates; bars represent mean ± sem). [Figure 59C]Figures 59A-59K: Detailed evaluation of LwaCas13a sensitivity to mismatches in the crRNA:target duplex with varying spacer lengths. (A) Knockdown of KRAS assessed with crRNAs containing a single mismatch at various positions in the spacer sequence. (B) Knockdown of PPIB assessed with crRNAs containing a single mismatch at various positions in the spacer sequence. (C) Knockdown of Gluc assessed with guides containing non-consecutive double mismatches at various positions in the spacer sequence. The wild-type sequence is shown at the top, and the identities of the mismatches are indicated below. (D) Collateral cleavage activity for ssRNA 1 and 2 for various spacer lengths. (n = 4 technical replicates; bars represent mean ± sem). (E) Specificity ratios of the crRNAs tested in (D). Specificity ratios are calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n = 4 technical replicates; bars represent mean ± sem). (F) Collateral cleavage activity of 28-nt spacer crRNAs containing synthetic mismatches tiled along the spacer against ssRNAs 1 and 2 (n = 4 technical replicates; bars represent mean ± sem). (G) Specificity ratio of the crRNAs tested in (F). The specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage (n = 4 technical replicates; bars represent mean ± sem). (H) Collateral cleavage activity of 23-nt spacer crRNAs containing synthetic mismatches tiled along the spacer against ssRNAs 1 and 2 (n = 4 technical replicates; bars represent mean ± sem). (I) Specificity ratio of the crRNAs tested in (H). The specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n=4 technical replicates; bars represent mean ± sem.) (J) Collateral cleavage activity against ssRNAs 1 and 2 for 20 nt spacer crRNAs containing synthetic mismatches tiled along the spacer.(n = 4 technical replicates; bars represent mean ± sem). (K) Specificity ratio of the crRNAs tested in (J). Specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n = 4 technical replicates; bars represent mean ± sem). [Fig. 59D-59G]Figures 59A-59K: Detailed evaluation of LwaCas13a sensitivity to mismatches in the crRNA:target duplex with varying spacer lengths. (A) Knockdown of KRAS assessed with crRNAs containing a single mismatch at various positions in the spacer sequence. (B) Knockdown of PPIB assessed with crRNAs containing a single mismatch at various positions in the spacer sequence. (C) Knockdown of Gluc assessed with guides containing non-consecutive double mismatches at various positions in the spacer sequence. The wild-type sequence is shown at the top, and the identities of the mismatches are indicated below. (D) Collateral cleavage activity for ssRNA 1 and 2 for various spacer lengths. (n = 4 technical replicates; bars represent mean ± sem). (E) Specificity ratios of the crRNAs tested in (D). Specificity ratios are calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n = 4 technical replicates; bars represent mean ± sem). (F) Collateral cleavage activity of 28-nt spacer crRNAs containing synthetic mismatches tiled along the spacer against ssRNAs 1 and 2 (n = 4 technical replicates; bars represent mean ± sem). (G) Specificity ratio of the crRNAs tested in (F). The specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage (n = 4 technical replicates; bars represent mean ± sem). (H) Collateral cleavage activity of 23-nt spacer crRNAs containing synthetic mismatches tiled along the spacer against ssRNAs 1 and 2 (n = 4 technical replicates; bars represent mean ± sem). (I) Specificity ratio of the crRNAs tested in (H). The specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n=4 technical replicates; bars represent mean ± sem.) (J) Collateral cleavage activity against ssRNAs 1 and 2 for 20 nt spacer crRNAs containing synthetic mismatches tiled along the spacer.(n = 4 technical replicates; bars represent mean ± sem). (K) Specificity ratio of the crRNAs tested in (J). Specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n = 4 technical replicates; bars represent mean ± sem). [Fig. 59H-59K]Figures 59A-59K: Detailed evaluation of LwaCas13a sensitivity to mismatches in the crRNA:target duplex with varying spacer lengths. (A) Knockdown of KRAS assessed with crRNAs containing a single mismatch at various positions in the spacer sequence. (B) Knockdown of PPIB assessed with crRNAs containing a single mismatch at various positions in the spacer sequence. (C) Knockdown of Gluc assessed with guides containing non-consecutive double mismatches at various positions in the spacer sequence. The wild-type sequence is shown at the top, and the identities of the mismatches are indicated below. (D) Collateral cleavage activity for ssRNA 1 and 2 for various spacer lengths. (n = 4 technical replicates; bars represent mean ± sem). (E) Specificity ratios of the crRNAs tested in (D). Specificity ratios are calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n = 4 technical replicates; bars represent mean ± sem). (F) Collateral cleavage activity of 28-nt spacer crRNAs containing synthetic mismatches tiled along the spacer against ssRNAs 1 and 2 (n = 4 technical replicates; bars represent mean ± sem). (G) Specificity ratio of the crRNAs tested in (F). The specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage (n = 4 technical replicates; bars represent mean ± sem). (H) Collateral cleavage activity of 23-nt spacer crRNAs containing synthetic mismatches tiled along the spacer against ssRNAs 1 and 2 (n = 4 technical replicates; bars represent mean ± sem). (I) Specificity ratio of the crRNAs tested in (H). The specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n=4 technical replicates; bars represent mean ± sem.) (J) Collateral cleavage activity against ssRNAs 1 and 2 for 20 nt spacer crRNAs containing synthetic mismatches tiled along the spacer.(n = 4 technical replicates; bars represent mean ± sem). (K) Specificity ratio of the crRNAs tested in (J). Specificity ratio is calculated as the ratio of on-target RNA (ssRNA 1) collateral cleavage to off-target RNA (ssRNA 2) collateral cleavage. (n = 4 technical replicates; bars represent mean ± sem). [Figure 60A-60B]Figures 60A-60F: LwaCas13a is more specific for endogenous targets than shRNA knockdown, with little variation corresponding to biological noise. (A) Left: Expression levels as log2(transcripts per million (TPM)) of all genes detected in RNA-seq libraries of non-targeting shRNA-transfected controls (x-axis) compared with KRAS-targeting shRNA (y-axis). Mean values of three biological replicates are shown. KRAS transcript data points are highlighted in red. Right: Expression levels as log2(transcripts per million (TPM)) of all genes detected in RNA-seq libraries of non-targeting LwaCas13a-crRNA-transfected controls (x-axis) compared with KRAS-targeting LwaCas13a-crRNA (y-axis). Mean values of three biological replicates are shown. KRAS transcript data points are highlighted in red. (B) Left: Expression levels in log2 (transcripts per million (TPM)) of all genes detected in the RNA-seq library of the non-targeting shRNA-transfected control (x-axis) compared to the PPIB-targeting shRNA (y-axis). Average values of three biological replicates are shown. PPIB transcript data points are highlighted in red. Right: Expression levels in log2 (transcripts per million (TPM)) of all genes detected in the RNA-seq library of the non-targeting LwaCas13a-crRNA-transfected control (x-axis) compared to the PPIB-targeting LwaCas13a-crRNA (y-axis). Average values of three biological replicates are shown. PPIB transcript data points are highlighted in red. (C) Comparison of individual replicates of the non-targeting shRNA condition (column 1) and the Gluc-targeting shRNA condition (column 2). (D) Comparison of individual replicates between non-targeting crRNA conditions (column 1) and Gluc-targeting crRNA conditions (column 2). (E) Pairwise comparison of individual replicates between non-targeting shRNA conditions and Gluc-targeting shRNA conditions. (F) Pairwise comparison of individual replicates between non-targeting crRNA conditions and Gluc-targeting crRNA conditions. [Figure 60C-60D]Figures 60A-60F: LwaCas13a is more specific for endogenous targets than shRNA knockdown, with little variation corresponding to biological noise. (A) Left: Expression levels as log2(transcripts per million (TPM)) of all genes detected in RNA-seq libraries of non-targeting shRNA-transfected controls (x-axis) compared with KRAS-targeting shRNA (y-axis). Mean values of three biological replicates are shown. KRAS transcript data points are highlighted in red. Right: Expression levels as log2(transcripts per million (TPM)) of all genes detected in RNA-seq libraries of non-targeting LwaCas13a-crRNA-transfected controls (x-axis) compared with KRAS-targeting LwaCas13a-crRNA (y-axis). Mean values of three biological replicates are shown. KRAS transcript data points are highlighted in red. (B) Left: Expression levels in log2 (transcripts per million (TPM)) of all genes detected in the RNA-seq library of the non-targeting shRNA-transfected control (x-axis) compared to the PPIB-targeting shRNA (y-axis). Average values of three biological replicates are shown. PPIB transcript data points are highlighted in red. Right: Expression levels in log2 (transcripts per million (TPM)) of all genes detected in the RNA-seq library of the non-targeting LwaCas13a-crRNA-transfected control (x-axis) compared to the PPIB-targeting LwaCas13a-crRNA (y-axis). Average values of three biological replicates are shown. PPIB transcript data points are highlighted in red. (C) Comparison of individual replicates of the non-targeting shRNA condition (column 1) and the Gluc-targeting shRNA condition (column 2). (D) Comparison of individual replicates between non-targeting crRNA conditions (column 1) and Gluc-targeting crRNA conditions (column 2). (E) Pairwise comparison of individual replicates between non-targeting shRNA conditions and Gluc-targeting shRNA conditions. (F) Pairwise comparison of individual replicates between non-targeting crRNA conditions and Gluc-targeting crRNA conditions. [Figure 60E]Figures 60A-60F: LwaCas13a is more specific for endogenous targets than shRNA knockdown, with little variation corresponding to biological noise. (A) Left: Expression levels as log2(transcripts per million (TPM)) of all genes detected in RNA-seq libraries of non-targeting shRNA-transfected controls (x-axis) compared with KRAS-targeting shRNA (y-axis). Mean values of three biological replicates are shown. KRAS transcript data points are highlighted in red. Right: Expression levels as log2(transcripts per million (TPM)) of all genes detected in RNA-seq libraries of non-targeting LwaCas13a-crRNA-transfected controls (x-axis) compared with KRAS-targeting LwaCas13a-crRNA (y-axis). Mean values of three biological replicates are shown. KRAS transcript data points are highlighted in red. (B) Left: Expression levels in log2 (transcripts per million (TPM)) of all genes detected in the RNA-seq library of the non-targeting shRNA-transfected control (x-axis) compared to the PPIB-targeting shRNA (y-axis). Average values of three biological replicates are shown. PPIB transcript data points are highlighted in red. Right: Expression levels in log2 (transcripts per million (TPM)) of all genes detected in the RNA-seq library of the non-targeting LwaCas13a-crRNA-transfected control (x-axis) compared to the PPIB-targeting LwaCas13a-crRNA (y-axis). Average values of three biological replicates are shown. PPIB transcript data points are highlighted in red. (C) Comparison of individual replicates of the non-targeting shRNA condition (column 1) and the Gluc-targeting shRNA condition (column 2). (D) Comparison of individual replicates between non-targeting crRNA conditions (column 1) and Gluc-targeting crRNA conditions (column 2). (E) Pairwise comparison of individual replicates between non-targeting shRNA conditions and Gluc-targeting shRNA conditions. (F) Pairwise comparison of individual replicates between non-targeting crRNA conditions and Gluc-targeting crRNA conditions. [Figure 60F]Figures 60A-60F: LwaCas13a is more specific for endogenous targets than shRNA knockdown, with little variation corresponding to biological noise. (A) Left: Expression levels as log2(transcripts per million (TPM)) of all genes detected in RNA-seq libraries of non-targeting shRNA-transfected controls (x-axis) compared with KRAS-targeting shRNA (y-axis). Mean values of three biological replicates are shown. KRAS transcript data points are highlighted in red. Right: Expression levels as log2(transcripts per million (TPM)) of all genes detected in RNA-seq libraries of non-targeting LwaCas13a-crRNA-transfected controls (x-axis) compared with KRAS-targeting LwaCas13a-crRNA (y-axis). Mean values of three biological replicates are shown. KRAS transcript data points are highlighted in red. (B) Left: Expression levels in log2 (transcripts per million (TPM)) of all genes detected in the RNA-seq library of the non-targeting shRNA-transfected control (x-axis) compared to the PPIB-targeting shRNA (y-axis). Average values of three biological replicates are shown. PPIB transcript data points are highlighted in red. Right: Expression levels in log2 (transcripts per million (TPM)) of all genes detected in the RNA-seq library of the non-targeting LwaCas13a-crRNA-transfected control (x-axis) compared to the PPIB-targeting LwaCas13a-crRNA (y-axis). Average values of three biological replicates are shown. PPIB transcript data points are highlighted in red. (C) Comparison of individual replicates of the non-targeting shRNA condition (column 1) and the Gluc-targeting shRNA condition (column 2). (D) Comparison of individual replicates between non-targeting crRNA conditions (column 1) and Gluc-targeting crRNA conditions (column 2). (E) Pairwise comparison of individual replicates between non-targeting shRNA conditions and Gluc-targeting shRNA conditions. (F) Pairwise comparison of individual replicates between non-targeting crRNA conditions and Gluc-targeting crRNA conditions. [Figure 61A-61B]Figure 61A-C: Detailed analysis of LwaCas13a and RNAi knockdown variability (standard deviation) across all samples. (A) Heatmap of correlation (Kendall's tau) for the log2(transcripts per million (TPM+1)) values of all genes detected in the RNA-seq library between targeting and non-targeting replicates of shRNA or crRNA targeting either the luciferase reporter or endogenous gene. (B) Heatmap of correlation (Kendall's tau) for the log2(transcripts per million (TPM+1)) values of all genes detected in the RNA-seq library between all replicates and perturbations. (C) Standard deviation distribution of the log2(transcripts per million (TPM+1)) values of all genes detected in the RNA-seq library between targeting and non-targeting replicates of each gene targeted by either shRNA or crRNA. [Figure 61C] Figure 61A-C: Detailed analysis of LwaCas13a and RNAi knockdown variability (standard deviation) across all samples. (A) Heatmap of correlation (Kendall's tau) for the log2(transcripts per million (TPM+1)) values of all genes detected in the RNA-seq library between targeting and non-targeting replicates of shRNA or crRNA targeting either the luciferase reporter or endogenous gene. (B) Heatmap of correlation (Kendall's tau) for the log2(transcripts per million (TPM+1)) values of all genes detected in the RNA-seq library between all replicates and perturbations. (C) Standard deviation distribution of the log2(transcripts per million (TPM+1)) values of all genes detected in the RNA-seq library between targeting and non-targeting replicates of each gene targeted by either shRNA or crRNA. [Figure 62A-62B]Figure 62A-J: LwaCas13a knockdown is specific to the target transcript and has no activity against measured off-target transcripts. (A) Heat map of absolute Gluc signal for the first 96 spacers tiling to Gluc. (B) Heat map of absolute Cluc signal for the first 96 spacers tiling to Gluc. (C) Relationship between absolute Gluc signal and normalized luciferase for Gluc tiling guides. (D) Relationship between absolute Cluc signal and normalized luciferase for Gluc tiling guides. (E) Relationship between absolute Cluc signal and normalized luciferase for Cluc tiling guides. (F) Relationship between absolute Gluc signal and normalized luciferase for Cluc tiling guides. (G) Relationship between PPIB 2-Ct levels and PPIB knockdown for PPIB tiling guides. (H) Relationship between GAPDH 2-Ct levels and PPIB knockdown for PPIB tiling guides. (I) The relationship between KRAS 2-Ct levels and KRAS knockdown in PPIB KRAS guides. (J) The relationship between GAPDH 2-Ct levels and KRAS knockdown in PPIB KRAS guides. [Figure 62C-62F]Figure 62A-J: LwaCas13a knockdown is specific to the target transcript and has no activity against measured off-target transcripts. (A) Heat map of absolute Gluc signal for the first 96 spacers tiling to Gluc. (B) Heat map of absolute Cluc signal for the first 96 spacers tiling to Gluc. (C) Relationship between absolute Gluc signal and normalized luciferase for Gluc tiling guides. (D) Relationship between absolute Cluc signal and normalized luciferase for Gluc tiling guides. (E) Relationship between absolute Cluc signal and normalized luciferase for Cluc tiling guides. (F) Relationship between absolute Gluc signal and normalized luciferase for Cluc tiling guides. (G) Relationship between PPIB 2-Ct levels and PPIB knockdown for PPIB tiling guides. (H) Relationship between GAPDH 2-Ct levels and PPIB knockdown for PPIB tiling guides. (I) The relationship between KRAS 2-Ct levels and KRAS knockdown in PPIB KRAS guides. (J) The relationship between GAPDH 2-Ct levels and KRAS knockdown in PPIB KRAS guides. [Figure 62G-62J]Figure 62A-J: LwaCas13a knockdown is specific to the target transcript and has no activity against measured off-target transcripts. (A) Heat map of absolute Gluc signal for the first 96 spacers tiling to Gluc. (B) Heat map of absolute Cluc signal for the first 96 spacers tiling to Gluc. (C) Relationship between absolute Gluc signal and normalized luciferase for Gluc tiling guides. (D) Relationship between absolute Cluc signal and normalized luciferase for Gluc tiling guides. (E) Relationship between absolute Cluc signal and normalized luciferase for Cluc tiling guides. (F) Relationship between absolute Gluc signal and normalized luciferase for Cluc tiling guides. (G) Relationship between PPIB 2-Ct levels and PPIB knockdown for PPIB tiling guides. (H) Relationship between GAPDH 2-Ct levels and PPIB knockdown for PPIB tiling guides. (I) The relationship between KRAS 2-Ct levels and KRAS knockdown in PPIB KRAS guides. (J) The relationship between GAPDH 2-Ct levels and KRAS knockdown in PPIB KRAS guides. [Figures 63A-63C] Figure 63A-F: dCas13a represses reporter gene expression and binds to endogenous genes. (A) dCas13a tiled across the synthetic HBG1 intron separating Cluc and Gluc has the ability to repress Gluc translation at a specific distance from the translation start site. (B) RNA immunoprecipitation enrichment of β-actin mRNA targeted by dCas13a and two targeting and one non-targeting crRNAs. (C) Comparison between the localization of dCas13-GFP and dCas13a-GFP-KRAB constructs for imaging ACTB. (D) Further view of the dCas13a-NLS-msfGFP negative feedback construct delivered with a non-targeting guide. (E) Further view of the dCas13a-NLS-msfGFP negative feedback construct delivered with an ACTB guide. (F) Further view of the ACTB-guided delivered dCas13a-NLS-msfGFP negative feedback construct. [Figure 63D] Figure 63A-F: dCas13a represses reporter gene expression and binds to endogenous genes. (A) dCas13a tiled across the synthetic HBG1 intron separating Cluc and Gluc has the ability to repress Gluc translation at a specific distance from the translation start site. (B) RNA immunoprecipitation enrichment of β-actin mRNA targeted by dCas13a and two targeting and one non-targeting crRNAs. (C) Comparison between the localization of dCas13-GFP and dCas13a-GFP-KRAB constructs for imaging ACTB. (D) Further view of the dCas13a-NLS-msfGFP negative feedback construct delivered with a non-targeting guide. (E) Further view of the dCas13a-NLS-msfGFP negative feedback construct delivered with an ACTB guide. (F) Further view of the ACTB-guided delivered dCas13a-NLS-msfGFP negative feedback construct. [Figure 63E] Figure 63A-F: dCas13a represses reporter gene expression and binds to endogenous genes. (A) dCas13a tiled across the synthetic HBG1 intron separating Cluc and Gluc has the ability to repress Gluc translation at a specific distance from the translation start site. (B) RNA immunoprecipitation enrichment of β-actin mRNA targeted by dCas13a and two targeting and one non-targeting crRNAs. (C) Comparison between the localization of dCas13-GFP and dCas13a-GFP-KRAB constructs for imaging ACTB. (D) Further view of the dCas13a-NLS-msfGFP negative feedback construct delivered with a non-targeting guide. (E) Further view of the dCas13a-NLS-msfGFP negative feedback construct delivered with an ACTB guide. (F) Further view of the ACTB-guided delivered dCas13a-NLS-msfGFP negative feedback construct. [Figure 63F]Figure 63A-F: dCas13a represses reporter gene expression and binds to endogenous genes. (A) dCas13a tiled across the synthetic HBG1 intron separating Cluc and Gluc has the ability to repress Gluc translation at a specific distance from the translation start site. (B) RNA immunoprecipitation enrichment of β-actin mRNA targeted by dCas13a and two targeting and one non-targeting crRNAs. (C) Comparison between the localization of dCas13-GFP and dCas13a-GFP-KRAB constructs for imaging ACTB. (D) Further view of the dCas13a-NLS-msfGFP negative feedback construct delivered with a non-targeting guide. (E) Further view of the dCas13a-NLS-msfGFP negative feedback construct delivered with an ACTB guide. (F) Further view of the ACTB-guided delivered dCas13a-NLS-msfGFP negative feedback construct. [Figure 64A-64B] Figure 64A-64B. dCas13a-NF can image stress granule formation in live cells. (A) Representative RNA FISH images of ACTB transcripts in dCas13a-NF-expressing cells with corresponding ACTB-targeting and non-targeting guides. Cell outlines are indicated by dashed lines. (B) Overall signal overlap between ACTB RNA FISH signals and dCas13a-NF quantified by Manders overlap coefficient (left) and Pearson correlation (right). Correlation and signal overlap are calculated pixel by pixel per cell. All values are mean ± SEM with n=3. ****p<0.0001; ***p<0.001; **p<0.01. Two-tailed Student's t-test was used for comparison. [Figure 65A-65B]Figures 65A-C. Direct and collateral transcript expression knockdown. (A) Knockdown of luciferase by active vs. dead Cas13a. (B) Knockdown of endogenous gene expression by active vs. dead Cas13a. Knockdown by dead Cas13a may be due to translation blockage or transcript destabilization due to binding. (C) Lack of collateral activity by dead Cas13a in mammalian cells. No change in transcript distribution size is observed using dead Cas13a in (A) and guides 1 and 2 for luciferase compared to non-targeting controls. [Figure 66A] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66B] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66C] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66D] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66E] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66F] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66G] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66H] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66I] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66J] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66K] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66L] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66M] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66N] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 66O] Figures 66A-66O. Sequence alignment of the various C2C2 orthologs of Figure 53 with the indicated consensus sequences. [Figure 67A] Figures 67A-67C. Alignment of Leptotrichia wadei F0279 C2c2 ("Lew2C2c2") with Listeria newyorkensis FSL M6-0635 C2c2 ("LibC2c2"). [Figure 67B] Figures 67A-67C. Alignment of Leptotrichia wadei F0279 C2c2 ("Lew2C2c2") with Listeria newyorkensis FSL M6-0635 C2c2 ("LibC2c2"). [Figure 67C] Figures 67A-67C. Alignment of Leptotrichia wadei F0279 C2c2 ("Lew2C2c2") with Listeria newyorkensis FSL M6-0635 C2c2 ("LibC2c2"). [Figure 68]Figure 68. Alignment of C2c2 HEPN domains exhibiting RNase activity. The top alignment block contains selected HEPN domains described so far, and the bottom block contains catalytic motifs from C2c2 effector proteins. Below each domain structure is an alignment of conserved motifs in selected representatives of the respective protein family. Catalytic residues are shown in white letters on a black background; conserved hydrophobic residues are highlighted in yellow; conserved small residues are highlighted in green; in bridge-helix alignments, positively charged residues are in red. Secondary structure predictions are shown below the aligned sequences: H represents an α-helix, and E represents an extended conformation (β-strand). The least conserved spacers between alignment blocks are indicated by numbers. DETAILED DESCRIPTION OF THE INVENTION
[0080] The figures herein are for illustrative purposes only and are not necessarily drawn to scale.
[0081] Detailed Description of the Invention Generally, CRISPR-Cas or CRISPR system, as used in the aforementioned documents, such as WO 2014 / 093622 (PCT / US2013 / 074667), collectively refers to the transcripts and other elements involved in the expression of or directing the activity of CRISPR-associated ("Cas") genes, including sequences encoding Cas genes, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr mate sequences (including "direct repeats" in the context of endogenous CRISPR systems and partial direct repeats processed by tracrRNA), guide sequences (also referred to as "spacers" in the context of endogenous CRISPR systems), or "RNAs" as that term is used herein (e.g., RNAs that guide Cas, such as Cas9, e.g., CRISPR RNA and trans-activating (tracr) RNA or single guide RNA (sgRNA) (chimeric RNA)), or other sequences and transcripts from the CRISPR locus. Generally, CRISPR systems are characterized by elements that promote the formation of CRISPR complexes at the site of the target sequence (also referred to as a protospacer in the context of endogenous CRISPR systems). When the CRISPR protein is the C2c2 protein, tracrRNA is not required.
[0082] In the context of CRISPR complex formation, a "target sequence" refers to a sequence to which a guide sequence is designed to have complementarity, where hybridization between the target sequence and the guide sequence promotes CRISPR complex formation. The target sequence may comprise an RNA polynucleotide. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell. In some embodiments, direct repeats may be identified in silico by searching for repetitive motifs that meet some or all of the following criteria: 1. present in a 2 Kb window of genomic sequence flanking a Type II CRISPR locus; 2. spanning 20-50 bp; and 3. spaced 20-50 bp apart. In some embodiments, two of these criteria may be used, e.g., 1 and 2, 2 and 3, or 1 and 3. In some embodiments, all three criteria may be used.
[0083] Generally, a guide sequence is any polynucleotide sequence that has sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct the sequence-specific binding of a CRISPR complex to the target sequence.The term "targeting sequence" refers to a portion of a guide sequence that has sufficient complementarity with a target sequence.In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any algorithm suitable for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., the Burrows-Wheeler aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequence is about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length. In some embodiments, the guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12 nucleotides in length, or shorter. Preferably, the guide sequence is 10 to 30 nucleotides in length. The ability of a guide sequence to direct sequence-specific binding of a CRISPR complex to a target sequence can be assessed by any suitable assay.For example, sufficient components of a CRISPR system to form a CRISPR complex, including the guide sequence to be tested, may be provided to a host cell having a corresponding target sequence, such as by transfection of a vector encoding the components of the CRISPR sequence, followed by assessing preferential cleavage within the target sequence, such as by a Surveyor assay as described herein. Similarly, cleavage of a target polynucleotide sequence can be determined in vitro by providing the target sequence, the components of the CRISPR complex, including the guide sequence to be tested, and a control guide sequence that differs from the test guide sequence, and comparing the binding or cleavage rate at the target sequence between reactions with the test guide sequence and the control guide sequence. Other assays are possible and will occur to those skilled in the art.
[0084] In classical CRISPR-Cas systems, the degree of complementarity between a guide sequence and its corresponding target sequence may be about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or 100% or more; a guide or RNA or sgRNA may be about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 or more nucleotides in length; or a guide or RNA or sgRNA may be less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12 nucleotides in length, or shorter. However, one aspect of the present invention is to reduce off-target interactions, e.g., to reduce guide interaction with less complementary target sequences. Indeed, in examples, the present invention demonstrates mutations that result in CRISPR-Cas systems that can distinguish between target sequences and off-target sequences having 80% to greater than about 95% complementarity, e.g., 83% to 84%, or 88% to 89%, or 94% to 95% complementarity (e.g., distinguishing an 18-nucleotide target from an 18-nucleotide off-target with one, two, or three mismatches). Thus, in the context of the present invention, the degree of complementarity between a guide sequence and its corresponding target sequence is 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, 99.9%, or greater than 100%. An off-target is one that has less than 100% or 99.9% or 99.5% or 99% or 99% or 98.5% or 98% or 97.5% or 97% or 96.5% or 96% or 95.5% or 95% or 94.5% or 94% or 93% or 92% or 91% or 90% or 89% or 88% or 87% or 86% or 85% or 84% or 83% or 82% or 81% or 80% complementarity between the sequence and the guide, advantageously an off-target is one that has less than 100% or 99.9% or 99.5% or 99% or 99% or 98.5% or 98% or 97.5% or 97% or 96.5% or 96% or 95.5% or 95% or 94.5% complementarity between the sequence and the guide.
[0085] In certain embodiments, the introduction of one or more mismatches, such as one or two mismatches, between the spacer sequence and the target sequence can be used to adjust the cleavage efficiency, including the position of the mismatch along the spacer / target. For example, the closer to the center (i.e., not closer to the 3' or 5') the double mismatch is, the greater the impact on cleavage efficiency. Thus, the cleavage efficiency can be adjusted by selecting the mismatch position along the spacer. For example, if less than 100% cleavage of the target is desired (e.g., in a certain cell population), one or more, preferably two, mismatches between the spacer and the target sequence can be introduced into the spacer sequence. The closer to the center the mismatch position is along the spacer, the lower the cleavage rate.
[0086] The methods of the invention as described herein include inducing one or more nucleotide modifications in a eukaryotic cell as discussed herein (in vitro, i.e., in an isolated eukaryotic cell), comprising delivering a vector as discussed herein to the cell. The one or more mutations can include the introduction, deletion, or substitution of one or more nucleotides in each target sequence of the one or more cells via one or more guide RNA(s) or one or more sgRNA(s). The mutations can include the introduction, deletion, or substitution of 1 to 75 nucleotides in each target sequence of the one or more cells via one or more guide RNA(s). The mutation may include the introduction, deletion, or substitution of 1, 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or 75 nucleotides in each target sequence of the one or more cells via one or more guide(s). The mutation may include the introduction, deletion, or substitution of 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or 75 nucleotides in each target sequence of the one or more cells via one or more guide(s). The mutation may include the introduction, deletion, or substitution of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or 75 nucleotides in each target sequence of said cell(s) via one or more guide(s). The mutation may include the introduction, deletion, or substitution of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or 75 nucleotides in each target sequence of said cell(s) via one or more guide(s).The mutation may include the introduction, deletion, or substitution of 40, 45, 50, 75, 100, 200, 300, 400, or 500 nucleotides in each target sequence of said cell(s) via one or more guide(s) RNA(s).
[0087] To minimize toxicity and off-target effects, it may be important to control the concentration of delivered Cas mRNA or protein and guide RNA. The optimal concentration of Cas mRNA or protein and guide RNA can be determined by testing various concentrations in cell models or non-human eukaryotic animal models and analyzing the extent of modification at potential off-target genomic loci using deep sequencing.
[0088] Typically, in the context of endogenous CRISPR systems, formation of a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) results in cleavage at or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50 or more base pairs of the target sequence), which can depend, for example, on secondary structure, particularly in the case of RNA targets.
[0089] The nucleic acid molecule encoding Cas is advantageously a codon-optimized Cas. An example of a codon-optimized sequence is one optimized for expression in a eukaryote, such as a human (i.e., optimized for human expression), or another eukaryote, animal, or mammal as discussed herein; see, for example, the SaCas9 human codon-optimized sequence in WO 2014 / 093622 (PCT / US2013 / 074667). While this is preferred, it is understood that other examples are possible, and codon optimization for host species other than humans, or for specific organs, are known. In some embodiments, the enzyme coding sequence encoding Cas is codon-optimized for expression in a specific cell, such as a eukaryotic cell. The eukaryotic cell may be of or derived from a specific organism, such as a mammal, including but not limited to a human, or a non-human eukaryote or animal or mammal as discussed herein, such as a mouse, rat, rabbit, dog, livestock, or non-human mammal or primate. In some embodiments, methods for modifying human germline genetic identity and / or methods for modifying animal genetic identity that may cause suffering to humans or animals without any substantial medical benefit to them, as well as the animals resulting from such methods, may be excluded. Generally, codon optimization refers to a method of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) of the native sequence with a codon that is more frequently or most frequently used in the genes of that host cell while maintaining the native amino acid sequence. Different species exhibit specific biases for certain codons for specific amino acids. Codon bias (differences in codon usage between organisms) is often correlated with the efficiency of messenger RNA (mRNA) translation, which in turn is thought to depend, among other things, on the properties of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell generally reflects the codons most frequently used in peptide synthesis.Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the "Codon Usage Database" available at www.kazusa.orjp / codon / , and these tables can be adapted in a number of ways. See Nakamura, Y., et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000," Nucl. Acids Res. 28:292 (2000). Computer algorithms are also available for codon-optimizing a particular sequence for expression in a particular host cell, such as Gene Forge (Aptagen; Jacobus, PA). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more, or all codons) in the Cas-encoding sequence correspond to the most frequently used codon for a particular amino acid.
[0090] In certain embodiments, the methods described herein may include providing a Cas transgenic cell in which one or more nucleic acids encoding one or more guide RNAs operably linked within the cell to regulatory elements comprising promoters of one or more genes of interest are provided or introduced. As used herein, the term "Cas transgenic cell" refers to a cell, such as a eukaryotic cell, in which a Cas gene has been genomically integrated. The nature, type, or origin of the cell is not particularly limited according to the present invention. Additionally, the method of introducing a Cas transgene into a cell may vary and may be any method known in the art. In certain embodiments, the Cas transgenic cell is obtained by introducing a Cas transgene into an isolated cell. In certain other embodiments, the Cas transgenic cell is obtained by isolating cells from a Cas transgenic organism. By way of example and without limitation, the Cas transgenic cell as referred to herein may be derived from a Cas transgenic eukaryotic organism, such as a Cas knock-in eukaryotic organism. See International Publication No. WO 2014 / 093622 (PCT / US13 / 74667), incorporated herein by reference. The methods of U.S. Patent Application Publication Nos. 20120017290 and 20110265198, assigned to Sangamo BioSciences, Inc., relating to targeting the Rosa locus, can be modified to utilize the CRISPR Cas system of the present invention. The methods of U.S. Patent Application Publication No. 20130236946, assigned to Cellectis, relating to targeting the Rosa locus can also be modified to utilize the CRISPR Cas system of the present invention. As a further example, see Platt et al. (Cell;159(2):440-455(2014)), which describes Cas9 knock-in mice, incorporated herein by reference. The Cas transgene may further comprise a Lox-Stop-PolyA-Lox (LSL) cassette, which allows Cas expression to be inducible by Cre recombinase.Alternatively, Cas transgenic cells may be obtained by introducing a Cas transgene into isolated cells. Transgene delivery systems are well known in the art. For example, Cas transgenes may be delivered using vectors (e.g., AAV, adenovirus, lentivirus) and / or particle and / or nanoparticle delivery, for example, in eukaryotic cells, as also described elsewhere herein.
[0091] Those skilled in the art will understand that cells, such as Cas transgenic cells as referenced herein, may contain additional genomic alterations in addition to having an integrated Cas gene, or may contain mutations, such as one or more oncogenic mutations, that arise due to the sequence-specific action of Cas when complexed with an RNA capable of guiding Cas to a target locus, as described, for example and without limitation, in Platt et al. (2014), Chen et al., (2014), or Kumar et al. (2009).
[0092] In some embodiments, the Cas sequence is fused to one or more nuclear localization sequences (NLSs) or nuclear export sequences (NESs), such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs or NESs. In some embodiments, the Cas contains about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs or NESs at or near the amino terminus, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs or NESs at or near the carboxy terminus, or a combination thereof (e.g., zero or at least one NLS or NES at the amino terminus and zero or one or more NLSs or NESs at the carboxy terminus). When more than one NLS or NES is present, each may be selected independently of the others, and thus a single NLS or NES may be present in two or more copies and / or in combination with one or more other NLSs or NESs present in one or more copies. In preferred embodiments of the invention, the Cas comprises no more than six NLSs. In some embodiments, an NLS or NES is considered to be near the N- or C-terminus when the nearest amino acid of the NLS or NES is within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the N- or C-terminus.Non-limiting examples of NLSs include the NLS of the SV40 virus large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO: X); the nucleoplasmin bipartite NLS having the NLS of nucleoplasmin (e.g., the sequence KRPAATKKAGQAKKKK) (SEQ ID NO: X); the c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: X) or RQRRNELKRSP (SEQ ID NO: X); the hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: X); the IBB domain of importin-α having the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: X); the sequences VSRKRPRP (SEQ ID NO: X) and PPKKARED (SEQ ID NO: X) of the fibroid T protein; the sequence POPKKKPL (SEQ ID NO: X) of human p53; and the mouse c-abl Examples of NLS sequences include those derived from the sequence SALIKKKKKMAP (SEQ ID NO: X) of Cas IV; the sequences DRLRR (SEQ ID NO: X) and PKQKKRK (SEQ ID NO: X) of influenza virus NS1; the sequence RKLKKKIKKL (SEQ ID NO: X) of hepatitis virus delta antigen; the sequence REKKKFLKRR (SEQ ID NO: X) of mouse Mx1 protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: X) of human poly(ADP-ribose) polymerase; and the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: X) of steroid hormone receptor (human) glucocorticoid. Non-limiting examples of NESs include the NES sequence LYPERLRRILT (ctgtaccctgagcggctgcggcggatcctgacc). Generally, one or more NLSs or NESs are strong enough to drive the accumulation of detectable amounts of Cas in the nucleus or cytoplasm, respectively, of eukaryotic cells. In general, the strength of nuclear localization / export activity can be derived from the number of NLSs / NESs in the Cas, the specific NLSs or NESs used, or a combination of these factors. Detection of nuclear / cytoplasmic accumulation can be performed by any suitable technique. For example, a detectable marker can be fused to the Cas, thereby visualizing its location within the cell, for example, by combining it with a means to detect nuclear (e.g., a nuclear-specific stain such as DAPI) or cytoplasmic location.Cell nuclei may also be isolated from cells, and their contents may then be analyzed by any suitable protein detection method, such as immunohistochemistry, Western blot, or enzyme activity assay. Accumulation in the nucleus may also be determined indirectly, such as by assaying for the effect of CRISPR complex formation (e.g., assaying for DNA cleavage or mutation at the target sequence, or assaying for altered gene expression activity affected by CRISPR complex formation and / or Cas enzymatic activity), compared to a control not exposed to the Cas or complex, or to a control exposed to Cas lacking one or more NLSs or NESs. In certain embodiments, other localization tags may be fused to the Cas protein, such as, without limitation, to localize Cas to specific sites within the cell, such as organelles, e.g., mitochondria, plastids, chloroplasts, vesicles, Golgi, (nuclear or cell) membranes, ribosomes, nucleoli, ER, cytoskeleton, vacuoles, centrosomes, nucleosomes, granules, centrioles, etc.
[0093] In certain aspects, the present invention relates to vectors for delivering or introducing, for example, Cas and / or RNA capable of guiding Cas to a target locus (i.e., guide RNA) into cells and for propagating these components (e.g., in prokaryotic cells). As used herein, a "vector" is a tool that allows or facilitates the transfer of an entity from one environment to another. It is a replicon, e.g., a plasmid, phage, or cosmid, into which another DNA segment can be inserted, resulting in replication of the inserted segment. Generally, a vector is capable of replication when associated with appropriate control elements. In general, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules containing one or more free ends, or no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and various other polynucleotides known in the art. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, in which the vector contains virus-derived DNA or RNA sequences for packaging into a virus (e.g., retrovirus, replication-deficient retrovirus, adenovirus, replication-deficient adenovirus, and adeno-associated virus (AAV)). Viral vectors also include virus-carried polynucleotides for transfection into host cells. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operably linked. Such vectors are referred to herein as "expression vectors." Common expression vectors useful in recombinant DNA techniques are often in the form of plasmids.
[0094] A recombinant expression vector can contain a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, meaning that the recombinant expression vector contains one or more regulatory elements (which may be selected based on the host cell used for expression) operably linked to the nucleic acid sequence to be expressed. Within the scope of a recombinant expression vector, "operably linked" is intended to mean that the nucleotide sequence of interest is linked to one or more regulatory elements in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). For recombination and cloning methods, see U.S. Patent Application No. 10 / 815,730, published September 2, 2004 as U.S. Patent Application Publication No. 2004-0171156 A1, the contents of which are incorporated herein by reference in their entirety.
[0095] The one or more vectors may include one or more regulatory elements, such as one or more promoters. The one or more vectors may include a Cas coding sequence, and / or a single guide RNA (e.g., sgRNA) coding sequence, but may include at least 3, 8, 16, 32, 48, or 50 RNAs (e.g., sgRNAs), for example, 1-2, 1-3, 1-4, 1-5, 3-6, 3-7, 3-8, 3-9, 3-10, 3-8, 3-16, 3-30, 3-32, 3-48, 3-50 RNAs (e.g., sgRNAs). A single vector advantageously contains about 16 or fewer RNAs, each of which (e.g., sgRNA) may have a promoter; and if a single vector provides more than 16 RNAs, one or more promoters may drive expression of two or more of those RNAs, e.g., if there are 32 RNAs, each promoter may drive expression of two RNAs, and if there are 48 RNAs, each promoter may drive expression of three RNAs. With simple arithmetic, well-established cloning protocols and the teachings of this disclosure, one skilled in the art can easily practice the present invention with RNAs for a suitable exemplary vector, such as AAV, and a suitable promoter, such as the U6 promoter. For example, the packaging limit of AAV is about 4.7 kb. The length of a single U6-gRNA (plus restriction sites for cloning) is 361 bp. Thus, one skilled in the art can easily fit about 12 to 16, e.g., 13, U6-gRNA cassettes into a single vector. This can be assembled by any suitable means, such as the Golden Gate strategy used in TALE assembly (http: / / www.genome-engineering.org / taleffectors / ). One skilled in the art can also use a tandem guide strategy to increase the number of U6-gRNAs by about 1.5-fold, e.g., from 12-16, e.g., 13, to about 18-24, e.g., about 19 U6-gRNAs. Thus, one skilled in the art can easily arrive at about 18-24, e.g., about 19 promoter-RNAs, e.g., U6-gRNAs, in a single vector, e.g., an AAV vector.Another way to increase the number of promoters and RNAs in a vector is to use a single promoter (e.g., U6) to express an array of RNAs separated by cleavable sequences. Yet another way to increase the number of promoter-RNAs in a vector is to express an array of promoter-RNAs separated by cleavable sequences in the coding sequence or intron of a gene; in this case, it is advantageous to use a polymerase II promoter, which can increase expression and enable transcription of long RNAs in a tissue-specific manner (see, for example, http: / / nar.oxfordjournals.org / content / 34 / 7 / e53.short, http: / / www.nature.com / mt / journal / v16 / n9 / abs / mt2008144a.html). In an advantageous embodiment, AAV can package U6 tandem gRNAs targeting up to about 50 genes. Thus, from knowledge in the art and the teachings of this disclosure, one of skill in the art can readily make and use one or more vectors, e.g., a single vector, that express multiple RNAs or guides under the control of, or operably or functionally linked to, one or more promoters (especially with respect to the number of RNAs or guides discussed herein) without any undue experimentation.
[0096] The coding sequence(s) of one or more guide RNAs and / or Cas coding sequences may be functionally or operably linked to one or more regulatory elements, thereby driving expression. The one or more promoters may be constitutive and / or conditional and / or inducible and / or tissue-specific promoters. The promoter may be selected from the group consisting of RNA polymerase, pol I, pol II, pol III, T7, U6, H1, retroviral Rous sarcoma virus (RSV) LTR promoter, cytomegalovirus (CMV) promoter, SV40 promoter, dihydrofolate reductase promoter, β-actin promoter, phosphoglycerol kinase (PGK) promoter, and EF1α promoter. A preferred promoter is U6.
[0097] Aspects of the present invention relate to the identification and engineering of novel effector proteins associated with Class 2 CRISPR-Cas systems. In preferred embodiments, the effector proteins comprise single-subunit effector modules. In further embodiments, the effector proteins function in prokaryotic or eukaryotic cells in in vitro, in vivo, or ex vivo applications. Certain aspects of the present invention encompass computational methods and algorithms for predicting novel Class 2 CRISPR-Cas systems and identifying components therein.
[0098] In one embodiment, a computational method for identifying novel class 2 CRISPR-Cas loci includes the following steps: detecting all contigs encoding the Cas1 protein; identifying all predicted protein-coding genes within a 20 kB region of the Cas1 gene, more specifically within a 20 kb region from the beginning of the Cas1 gene and within a 20 kb region from the end of the Cas1 gene; comparing the identified genes with a Cas protein-specific profile to predict CRISPR arrays; selecting partial and / or unclassified candidate CRISPR-Cas loci containing proteins greater than 500 amino acids (>500 aa); and analyzing the selected candidates using PSI-BLAST and HHPred, thereby isolating and identifying novel class 2 CRISPR-Cas loci. In addition to the above steps, further analysis of the candidates can be performed by searching metagenomics databases for additional homologs.
[0099] In one embodiment, the step of detecting all contigs encoding Cas1 proteins is performed by GenemarkS, a gene prediction program as further described in "GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions," John Besemer, Alexandre Lomsadze and Mark Borodovsky, Nucleic Acids Research (2001) 29, pp 2607-2618, incorporated herein by reference.
[0100] In one embodiment, the step of identifying all predicted protein-coding genes is performed by comparing the identified genes to the Cas protein-specific profile and annotating them according to the NCBI Conserved Domain Database (CDD), a protein annotation resource consisting of a collection of fully annotated multiple sequence alignment models of ancient domains and full-length proteins. These are available as position-specific score matrices (PSSMs) for rapid identification of conserved domains in protein sequences via RPS-BLAST. The CDD content includes NCBI-curated domains that explicitly define domain boundaries using three-dimensional structural information and provide insight into sequence / structure / function relationships, as well as domain models imported from several external database sources (Pfam, SMART, COG, PRK, TIGRFAM). In a further embodiment, CRISPR arrays were predicted using the PILER-CR program, a publicly available domain software for finding CRISPR repeats, as described in "PILER-CR: fast and accurate identification of CRISPR repeats," Edgar, RC, BMC Bioinformatics, Jan 20;8:18 (2007), incorporated herein by reference.
[0101] In a further embodiment, a case-by-case analysis is performed using PSI-BLAST (Position-Specific Iterative Basic Local Alignment Search Tool). PSI-BLAST derives a position-specific scoring matrix (PSSM) or profile from a multiple sequence alignment of sequences found above a given score threshold using protein-protein BLAST. This PSSM is used to further search the database for new matches and is subsequently updated to iterate with these newly found sequences. Thus, PSI-BLAST provides a means to detect distant relationships between proteins.
[0102] In another embodiment, case-by-case analysis is performed using HHpred, a sequence database search and structure prediction method that is as easy to use as BLAST or PSI-BLAST, yet much more sensitive in finding distant homologs. In fact, HHpred's sensitivity rivals that of the most powerful structure prediction servers currently available. HHpred is the first server based on pairwise comparison of profile hidden Markov models (HMMs). While most traditional sequence search methods search sequence databases such as UniProt or NR, HHpred searches alignment databases such as Pfam or SMART. This significantly simplifies the list of hits to a few sequence families rather than a cluttered pile of single sequences. All major publicly available profile and alignment databases are accessible in HHpred. HHpred accepts a single query sequence or multiple alignments as input. HHpred returns search results in an easy-to-read format similar to PSI-BLAST within just a few minutes. Search options include local or global alignments and secondary structure similarity scoring. HHpred can generate pairwise query-template sequence alignments, synthetic query-template multiple alignments (eg, for transitive searches), as well as three-dimensional structural models calculated by the MODELLER software from HHpred alignments.
[0103] The term "nucleic acid targeting system," in which the nucleic acid is DNA or RNA and in some embodiments can also refer to a DNA-RNA hybrid or derivative thereof, collectively refers to the transcripts and other elements involved in the expression of or directing the activity of a DNA- or RNA-targeting CRISPR-associated ("Cas") gene, which may include sequences encoding a DNA- or RNA-targeting Cas protein and a DNA- or RNA-targeting guide RNA, which may include a CRISPR RNA (crRNA) sequence and (in some, but not all systems) a trans-activating CRISPR-Cas system RNA (tracrRNA) sequence, or other sequences and transcripts from a DNA- or RNA-targeting CRISPR locus. Generally, RNA targeting systems are characterized by elements that promote the formation of a DNA- or RNA-targeting complex at the site of the target DNA or RNA sequence. In the context of forming a DNA or RNA targeting complex, "target sequence" refers to a DNA or RNA sequence to which a DNA or RNA targeting guide RNA is designed to have complementarity, where hybridization between the target sequence and the RNA targeting guide RNA promotes the formation of an RNA targeting complex. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell.
[0104] In one embodiment of the present invention, the novel RNA targeting system, also referred to as the RNA- or RNA-targeting CRISPR / Cas or CRISPR-Cas based RNA targeting system of the present application, is based on the identified type VI Cas protein, which does not require the creation of a customized protein to target a specific RNA sequence, but rather a single enzyme can be programmed to recognize a specific RNA target by an RNA molecule, which in turn can be used to recruit an enzyme to a specific RNA target.
[0105] In one embodiment of the present invention, the present novel DNA targeting system, also referred to as a DNA- or DNA-targeting CRISPR / Cas or CRISPR-Cas based RNA targeting system, is based on the identified type VI Cas protein, which does not require the creation of a customized protein to target a specific RNA sequence, but rather a single enzyme can be programmed to recognize a specific DNA target by an RNA molecule, which in turn can be used to recruit the enzyme to a specific DNA target.
[0106] The nucleic acid targeting systems, vector systems, vectors and compositions described herein can be used in a variety of nucleic acid targeting applications, altering or modifying the synthesis of gene products such as proteins, nucleic acid cleavage, nucleic acid editing, nucleic acid splicing; target nucleic acid transport, target nucleic acid tracking, target nucleic acid isolation, target nucleic acid visualization, etc.
[0107] As used herein, Cas protein or CRISPR enzyme refers to any of the proteins represented in the novel class of CRISPR-Cas systems.
[0108] C2c2 nuclease The class 2 type VI effector protein C2c2 is an RNA-guided RNase that can be efficiently programmed to degrade ssRNA.The C2c2 effector proteins of the present invention include, without limitation, the following 21 ortholog species (including multiple CRISPR loci): Leptotrichia shahii; Leptotrichia wadei (Lw2); Listeria seeligeri; Lachnospiraceae bacterium MA2020; Lachnospiraceae bacterium NK4A179; [Clostridium] aminophilum DSM 10710; Carnobacterium gallinarum DSM 4847; Carnobacterium gallinarum DSM 4847; 4847 (second CRISPR locus); Paludibacter propionicigenes WB4; Listeria weihenstephanensis FSL R9-0317; Listeriaceae FSL M6-0635; Leptotrichia wadei F0279; Rhodobacter capsulatus SB 1003; Rhodobacter capsulatus R121; Rhodobacter capsulatus DE442; Leptotrichia buccalis buccalis C-1013-b; Herbinix hemicellulosilytica; [Eubacterium] rectale; Eubacteriaceae bacterium CHKCI004; Blautia sp. Marseille-P2398; and Leptotrichia sp. oral bacterial taxon 879 strain F0557.Twelve additional non-limiting examples include the Lachnospiraceae bacterium NK4A144; Chloroflexus aggregans; Demequina aurantiaca; Thalassospira sp. TSL5-1; Pseudobutyrivibrio sp. OR37; Butyrivibrio sp. YAB3001; Blautia sp. Marseille-P2398; Leptotrichia sp. Marseille-P3007; Bacteroides ihuae; Porphyromonadaceae bacterium KH3CP3RA; and Listeria riparia. riparia); and Insolitispirillum peregrinum.
[0109] C2c2 achieves RNA cleavage through conserved basic residues within its two HEPN domains. Mutations in the HEPN domain, such as substitutions of predicted HEPN domain catalytic residues (e.g., alanine), can be used to convert C2c2 into an inactive programmable RNA-binding protein (dC2c2, similar to dCas9).
[0110] According to the present invention, consensus sequences can be generated from multiple C2c2 orthologs, which can aid in locating conserved amino acid residues and motifs, including but not limited to, catalytic residues and the HEPN motif, of C2c2 orthologs that mediate C2c2 function. One such consensus sequence generated from the 33 orthologs described above using Geneious alignment is as follows: [ka]
[0111] Figures 49 and 50 provide the HEPN sequence motifs identified from the above orthologs, based on the first 21 and all 33 orthologs, respectively. Non-limiting examples of amino acid residues that can be mutated to generate catalytically ineffective C2c2 mutants based on the above consensus include D372, R377, Q / H382, and F383 or the corresponding amino acids in orthologs of HEPN1, and K893, N894, R898, N899, H903, F904, Y906, Y927, D928, K930, K932 or the corresponding amino acids in orthologs of HEPN2.
[0112] In another non-limiting example, a sequence alignment tool that aids in generating consensus sequences and identifying conserved residues is the MUSCLE alignment tool (www.ebi.ac.uk / Tools / msa / muscle / ). For example, using MUSCLE, the following amino acid positions can be identified in Leptotrichia wadei C2c2 that are conserved among C2c2 orthologs: K2; K5; V6; E301; L331; I335; N341; G351; K352; E375; L392; L396; D403; F446; I466; I470; R474 (HEPN); H475; H479 (HEPN), E508; P556; L561; I595; Y596; F600; Y669; I6 73;F681;L685;Y761;L676;L779;Y782;L836;D847;Y863;L869;I872;K879;I933;L954;I958;R96 1;Y965;E970;R971;D972;R1046(HEPN), H1051(HEPN), Y1075;D1076;K1078;K1080;I1083;I1090.
[0113] Figures 52A-K show alignments of C2c2 orthologs, and Figure 52L shows an exemplary sequence alignment of the HEPN domain and highly conserved residues.
[0114] C2c2 HEPN can also target DNA, or potentially DNA and / or RNA. Based on the ability of the HEPN domain of C2c2 to bind to, and in its wild-type form cleave, at least RNA, it is preferred that the C2c2 effector protein possess RNase function. Additionally or alternatively, it may possess DNase function.
[0115] Thus, in some embodiments, the effector protein may be an RNA-binding protein, such as a dead Cas-type effector protein, which may optionally be functionalized with, for example, a transcription activator or repressor domain, an NLS, or other functional domain, as described herein. In some embodiments, the effector protein may be an RNA-binding protein that cleaves a single strand of RNA. If the bound RNA is ssRNA, the ssRNA is completely cleaved. In some embodiments, the effector protein may be an RNA-binding protein that cleaves a double strand of RNA, for example, if it contains two RNase domains. If the bound RNA is dsRNA, the dsRNA is completely cleaved.
[0116] RNase function in CRISPR systems is well known; for example, mRNA targeting has been reported for certain type III CRISPR-Cas systems (Hale et al., 2014, Genes Dev, vol. 28, 2432-2443; Hale et al., 2009, Cell, vol. 139, 945-956; Peng et al., 2015, Nucleic Acids Research, vol. 43, 406-417), providing significant advantages. In the Staphylococcus epidermis type III-A system, transcription across the target leads to cleavage of the target DNA and its transcript, which is mediated by a separate active site within the Cas10-Csm ribonucleoprotein effector complex (Samai et al., 2015, Cell, vol. 151, 1164-1174). Thus, there is provided a CRISPR-Cas system, composition or method for targeting RNA with the present effector protein.
[0117] The target RNA, i.e., the RNA of interest, is the RNA to be targeted by the present invention, leading to the recruitment and binding of an effector protein to a desired target site on the target RNA. The target RNA may be any suitable form of RNA, which in some embodiments may include mRNA. In other embodiments, the target RNA may include tRNA or rRNA. In other embodiments, the target RNA may include miRNA. In other embodiments, the target RNA may include siRNA.
[0118] Interfering RNA (RNAi) and microRNA (miRNA) In other embodiments, the target RNA can include both eukaryotic and prokaryotic interfering RNA, i.e., RNA involved in the RNA interference pathway, such as shRNA, siRNA, etc. In other embodiments, the target RNA can include microRNA (miRNA). Controlling interfering RNA or miRNA can help reduce the off-target effects (OTE) seen in methods by shortening the lifespan of interfering RNA or miRNA in vivo or in vitro.
[0119] In certain embodiments, the target is not the miRNA itself, but rather the miRNA binding site of the miRNA target.
[0120] In certain embodiments, the miRNA may be sequestered (including, for example, being relocated within the cell). In certain embodiments, the miRNA may be truncated, such as, without limitation, a hairpin.
[0121] In certain embodiments, miRNA processing (including, for example, turnover) is increased or decreased.
[0122] If the effector protein and suitable guide are selectively expressed (e.g., under the control of a spatially or temporally suitable promoter, such as a tissue- or cell cycle-specific promoter and / or enhancer), this could be used to "protect" a cell or system (in vivo or in vitro) from RNAi in that cell. This could be useful in adjacent tissues or cells where RNAi is not required, or for purposes of comparing cells or tissues where the effector protein and suitable guide are expressed and not expressed (i.e., where RNAi is not and is controlled, respectively). The effector protein may be used to regulate or bind molecules that contain or consist of RNA, such as ribozymes, ribosomes, or riboswitches. In embodiments of the invention, the RNA guide recruits the effector protein to such molecules, thereby enabling the effector protein to bind to them.
[0123] The protein system of the present invention can be applied in the area of RNAi technology, including therapeutics, assays, and other applications, without undue experimentation from this disclosure, as this application provides a basis for informed engineering of this system (see, e.g., Guidi et al., PLoS Negl Trop Dis 9(5):e0003801. doi:10.1371 / journal.pntd; Crotty et al., "In vivo RNAi screens: concepts and applications." Shane Crotty...2015 Elsevier Ltd. Published by Elsevier Inc., Pesticide Biochemistry and Physiology (Impact Factor: 2.01). 01 / 2015; 120. DOI:10.1016 / j.pestbp.2015.01.002 and Makkonen et al., Viruses 2015,7(4),2099-2125; see doi:10.3390 / v7042099).
[0124] ribosomal RNA (rRNA) For example, azalide antibiotics such as azithromycin are well known. They target and destroy the 50S ribosomal subunit. In some embodiments, the effector protein, together with a suitable guide RNA targeting the 50S ribosomal subunit, can be recruited to and bind to the 50S ribosomal subunit. Thus, the effector protein is provided in combination with a suitable guide directed to a ribosome (particularly the 50S ribosomal subunit). The use of this effector protein in combination with a suitable guide directed to a ribosome (particularly the 50S ribosomal subunit) can include use as an antibiotic. In particular, the use as an antibiotic is similar to the function of azalide antibiotics such as azithromycin. In some embodiments, prokaryotic ribosomal subunits, such as the prokaryotic 70S subunit, the above-mentioned 50S subunit, 30S subunit, and 16S and 5S subunits, can be targeted. In other embodiments, eukaryotic ribosomal subunits may be targeted, such as the eukaryotic 80S subunit, 60S subunit, 40S subunit, and the 28S, 18S, 5.8S and 5S subunits.
[0125] In some embodiments, the effector protein may be an RNA binding protein, optionally functionalized as described herein. In some embodiments, the effector protein may be an RNA binding protein that cleaves a single strand of RNA. In either case, but particularly when the RNA binding protein cleaves a single strand of RNA, ribosome function may be regulated, and in particular, may be reduced or destroyed. This may apply to any ribosomal RNA and any ribosomal subunit, and the sequence of rRNA is well known.
[0126] Thus, control of ribosomal activity is contemplated through the use of the present effector proteins in conjunction with a suitable guide to a ribosomal target. This may be by cleaving or binding to the ribosome. In particular, reduction of ribosomal activity is contemplated. This may be useful in in vivo or in vitro assays of ribosomal function and also as a means of controlling ribosomal activity-based therapies in vivo or in vitro. Furthermore, control (i.e., reduction) of protein synthesis in in vivo or in vitro systems is contemplated, including use as antibiotics and in research and diagnostic uses.
[0127] Riboswitches Riboswitches (also known as aptazymes) are regulatory segments of messenger RNA molecules that bind to small molecules. This typically results in altered production of the protein encoded by the mRNA. Thus, control of riboswitch activity through the use of the instant effector proteins in conjunction with suitable guides to the riboswitch target is therefore contemplated. This may be by cleaving the riboswitch or by binding to it. In particular, reduction of riboswitch activity is contemplated. This may be useful in in vivo or in vitro assays of riboswitch function and also as a means to control therapeutics based on riboswitch activity in vivo or in vitro. Furthermore, control (i.e., reduction) of protein synthesis in in vivo or in vitro systems is contemplated. This control, insofar as rRNA is concerned, may include uses as antibiotics and in research and diagnostic uses.
[0128] Ribozymes Ribozymes are RNA molecules with catalytic properties similar to enzymes (which are, of course, proteins). Because ribozymes, both naturally occurring and engineered, contain or consist of RNA, they can similarly be targets for the present RNA-binding effector proteins. In some embodiments, the effector protein can be an RNA-binding protein that cleaves the ribozyme, thereby disabling it. Thus, control of ribozyme activity is contemplated through the use of the present effector proteins in conjunction with a suitable guide to the ribozyme target. This may be by cleaving or binding to the ribozyme. In particular, reduction of ribozyme activity is contemplated. This may be useful in in vivo or in vitro assays of ribozyme function, and also as a means of controlling ribozyme activity-based therapies in vivo or in vitro.
[0129] Gene expression including RNA processing Effector proteins, together with a suitable guide, may also be used to target gene expression, including by controlling RNA processing. Control of RNA processing may include RNA splicing, including alternative splicing by targeting RNApol; viral replication, including plant viroids (particularly those of satellite viruses, bacteriophages, and retroviruses, such as HBV, HBV, and HIV, as well as other viruses listed herein); and RNA processing reactions such as tRNA biosynthesis. Effector proteins and a suitable guide may also be used to control RNA activation (RNAa). RNAa leads to enhanced gene expression, so gene expression control may be achieved by disrupting or reducing RNAa, thereby reducing the enhanced gene expression. This will be discussed in more detail below.
[0130] RNAi screen RNAi screens allow for the investigation of biological pathways and the identification of components by identifying gene products whose knockdown is associated with phenotypic changes. Control can also be exerted during such screens by using effector proteins and suitable guides to remove or reduce the activity of RNAi in the screen, thereby restoring (by removing or reducing interference / repression) the activity of the (previously interfered) gene product.
[0131] Satellite RNA (satRNA) and satellite viruses may also be treated.
[0132] Regulation herein in relation to RNase activity generally means reduction, negative disruption or knockdown or knockout.
[0133] In vivo RNA applications Inhibition of gene expression The target-specific RNases provided herein are capable of highly specific cleavage of target RNAs. Interference at the RNA level can be regulated both spatially and temporally and in a non-invasive manner since the genome is not altered.
[0134] It has been demonstrated that several diseases can be treated by mRNA targeting, and although most of these studies involve the administration of siRNA, it is clear that the RNA targeting effector proteins provided herein can be applied in the same way.
[0135] Examples of mRNA targets (and corresponding disease treatments) are VEGF, VEGF-R1 and RTP801 (for the treatment of AMD and / or DME), caspase 2 (for the treatment of Na+), ADRB2 (for the treatment of intraocular pressure), TRPVI (for the treatment of dry eye syndrome, Syk kinase (for the treatment of asthma), Apo B (for the treatment of hypercholesterolemia or hypobetalipoproteinemia), PLK1, KSP and VEGF (for the treatment of solid tumors), Ber-Abl (for the treatment of CML) (Burnett and Rossi Chem Biol. 2012, 19(1):60-71)). Similarly, RNA targeting has been demonstrated to be effective in treating RNA virus-mediated diseases, such as HIV (targeting HIV Tet and Rev), RSV (targeting RSV nucleocapsid), and HCV (targeting miR-122) (Burnett and Rossi Chem Biol. 2012, 19(1):60-71).
[0136] It is further envisioned that the RNA targeting effector proteins of the present invention can be used for mutation-specific or allele-specific knockdown. Guide RNAs can be designed to specifically target the sequence of transcribed mRNA containing a mutation or an allele-specific sequence. Such specific knockdown is particularly suitable for therapeutic applications involving disorders associated with mutant or allele-specific gene products. For example, most cases of familial hypobetalipoproteinemia (FHBL) are caused by mutations in the ApoB gene. This gene encodes two versions of apolipoprotein B protein: a short version (ApoB-48) and a longer version (ApoB-100). Some ApoB gene mutations that lead to FHBL cause both versions of ApoB to be abnormally short. Specific targeting and knockdown of mutated ApoB mRNA transcripts with the RNA targeting effector proteins of the present invention may be beneficial in treating FHBL. As another example, Huntington's disease (HD) is caused by an expansion of a CAG triplet repeat in the gene encoding huntingtin, resulting in an abnormal protein. Specific targeting and knockdown of mutated or allele-specific mRNA transcripts encoding huntingtin protein with the RNA targeting effector proteins of the present invention may be beneficial in the treatment of HD.
[0137] In this regard, and more generally, for various applications as described herein, it is noted that the use of split versions of RNA-targeting effector proteins may be envisioned. Indeed, this may not only allow for increased specificity but may also be advantageous for delivery. C2c2 is split in the sense that the two parts of the C2c2 enzyme essentially form a functional C2c2. Ideally, the split should always be such that one or more catalytic domains remain unaffected. This C2c2 may function as a nuclease, or it may be a dead C2c2, an RNA-binding protein that essentially has little or no catalytic activity, typically due to one or more mutations in its catalytic domain.
[0138] Each half of split C2c2 may be fused to a dimerization partner. For example, and not by way of limitation, a rapamycin-sensitive dimerization domain can be used to create a chemically inducible split C2c2 for temporal control of C2c2 activity. Thus, C2c2 can be made chemically inducible by splitting it into two fragments, and the rapamycin-sensitive dimerization domain can be used to reassemble C2c2 in a controlled manner. The two parts of split C2c2 can be considered the N'-terminal and C'-terminal parts of split C2c2. This fusion is typically at the split point of C2c2. In other words, the C'-terminus of the N'-terminal part of split C2c2 is fused to one of the dimer halves, while the N'-terminus of the C'-terminal part is fused to the other dimer half.
[0139] C2c2 does not need to be split in the sense that the cleavage point is newly created. The split point is typically designed in silico and cloned into the construct. Together, the two parts of the split C2c2, the N'-terminal part and the C'-terminal part, form a complete C2c2 that preferably contains at least 70% or more of the wild-type amino acids (or the nucleotides encoding them), preferably at least 80% or more, preferably at least 90% or more, preferably at least 95% or more, and most preferably at least 99% or more of the wild-type amino acids (or the nucleotides encoding them). Some trimming may be possible, and mutants are envisioned. Non-functional domains may be completely removed. The important point is that the two parts can be brought together and the desired C2c2 function is restored or reverted. The dimer may be a homodimer or a heterodimer.
[0140] In certain embodiments, C2c2 effectors as described herein may be used for mutation-specific or allele-specific targeting, such as mutation-specific or allele-specific knockdown.
[0141] The RNA targeting effector protein can further be fused to another functional RNase domain, such as nonspecific RNase or Argonaute 2, which acts synergistically to increase RNase activity or ensure further degradation of the message.
[0142] Regulation of gene expression by modulating RNA function Aside from the direct effect on gene expression through mRNA cleavage, RNA targeting can also be used to affect specific aspects of RNA processing in cells, potentially enabling more finely tuned gene expression. Generally, regulation can be mediated by interfering with protein binding to RNA, for example, by blocking protein binding or recruiting RNA-binding proteins. Regulation can be achieved at various levels, including mRNA splicing, transport, localization, translation, and turnover. Similarly, in therapeutic contexts, the use of RNA-specific targeting molecules can be envisioned to address (pathogenic) dysfunctions at each of these levels. In these embodiments, it is often preferred that the RNA targeting protein is a "dead" C2c2, such as the mutant c2c2 described herein, that has lost the ability to cleave RNA targets but still retains its ability to bind to them.
[0143] A) Alternative splicing Many human genes express multiple mRNAs as a result of alternative splicing. A variety of diseases have been shown to be associated with aberrant splicing, leading to loss or gain of function of the expressed gene. Some of these diseases are caused by mutations that result in splicing defects, but many others are not. One therapeutic option is to directly target the splicing machinery. The RNA targeting effector proteins described herein can be used, for example, to block or promote slicing, to influence exon inclusion or exclusion and expression of specific isoforms, and / or to stimulate expression of alternative protein products. Such applications are described in further detail below.
[0144] When an RNA targeting effector protein binds to a target RNA, it can sterically block the access of splicing factors to the RNA sequence. An RNA targeting effector protein that targets a splice site can block splicing at that site and optionally redirect splicing to an adjacent site. For example, binding of an RNA targeting effector protein that binds to a 5' splice site can block the recruitment of the U1 component of the spliceosome, favoring skipping of that exon. Alternatively, an RNA targeting effector protein that targets a splicing enhancer or silencer can prevent the binding of trans-acting regulatory splicing factors at the target site, effectively blocking or promoting splicing. Furthermore, exon exclusion can be achieved by recruiting ILF2 / 3 to the vicinity of an exon in a pre-mRNA using an RNA targeting effector protein as described herein. As yet another example, a glycine-rich domain can be added for hnRNP A1 recruitment and exon exclusion (Del Gatto-Konczak et al. Mol Cell Biol. 1999 Jan;19(1):251-60).
[0145] In certain embodiments, by appropriate selection of the gRNA, specific splice variants can be targeted while other splice variants are not targeted.
[0146] In some cases, RNA targeting effector proteins can be used to promote slicing (e.g., in cases where splicing is deficient). For example, RNA targeting effector proteins can be associated with effectors capable of stabilizing splicing-regulatory stem-loops for further splicing. Linking RNA targeting effector proteins to consensus binding site sequences for specific splicing factors can recruit the proteins to target DNA.
[0147] Examples of diseases associated with aberrant splicing include, but are not limited to, paraneoplastic opsoclonus-myoclonus ataxia (POMA), which is caused by loss of the Nova protein, which regulates the splicing of proteins that function at synapses, and cystic fibrosis, which is caused by a splicing defect in the cystic fibrosis transmembrane conductance regulator, resulting in the production of nonfunctional chloride channels. In other diseases, aberrant RNA splicing results in a gain of function, such as myotonic dystrophy, which is caused by a CUG triplet repeat expansion (50 to >1500 repeats) in the 3' UTR of mRNA, resulting in a splicing defect.
[0148] RNA-targeting effector proteins can be used to exclude exons by recruiting splicing factors (such as U1) to the 5' splice site to promote excision of the intron surrounding the desired exon. Such recruitment can be mediated by fusion with an arginine / serine-rich domain that functions as a splicing activator (Gravely BR and Maniatis T, Mol Cell. 1998(5):765-71).
[0149] It is envisioned that RNA targeting effector proteins can be used to block the splicing machinery at a desired gene locus, thereby preventing exon recognition and the expression of an alternative protein product. An example of a treatable disorder is Duchenne muscular dystrophy (DMD), which is caused by mutations in the gene encoding the dystrophin protein. Nearly all DMD mutations lead to frameshifts, resulting in impaired dystrophin translation. By combining RNA targeting effector proteins with splice junctions or exon splicing enhancers (ESEs), thereby preventing exon recognition, translation of a partially functional protein can be achieved. This converts the lethal Duchenne phenotype to the less severe Becker phenotype.
[0150] B) RNA modification RNA editing is a natural process in which small modifications to RNA increase the diversity of gene products of a given sequence. Typically, this modification involves the conversion of adenosine (A) to inosine (I), resulting in an RNA sequence that differs from that encoded by the genome. RNA modification is generally achieved by ADAR enzymes, whereby the pre-RNA target forms an incomplete double-stranded RNA through base pairing between the exon containing the edited adenosine and an intronic non-coding element. A classic example of AI editing is the glutamate receptor GluR-B mRNA, where this change results in an alteration of the channel's conductance properties (Higuchi M, et al. Cell. 1993;75:1361-70).
[0151] According to the present invention, enzymatic approaches are used to induce transitions (A⇔G or C⇔U changes) or transversions (any purine to any pyrimidine, or vice versa) in RNA bases of a given transcript. Transitions can be induced directly using adenine deaminase (ADAR1 / 2, APOBEC) or cytosine deaminase (AID), which convert A to I or C to U, respectively. Transitions can be induced indirectly by localizing reactive oxygen species damage to the target base, resulting in a chemical modification of the affected base, such as the conversion of guanine to oxoguanine. Oxoguanine is recognized as T and therefore base pairs with adenine, affecting translation. Proteins that can be recruited for ROS-mediated base damage include APEX and mini-SOG. In both approaches, these effectors can be fused to catalytically inactive C2c2 and recruited to sites on the transcript where these types of mutations are desired.
[0152] In humans, heterozygous functional null mutations in the ADAR1 gene lead to the skin disease human pigmented genodermatopathy (Miyamura Y, et al. Am J Hum Genet. 2003; 73: 693-9). It is envisioned that the RNA targeting effector proteins of the present invention can be used to correct dysfunctional RNA modifications.
[0153] Furthermore, it is envisioned that RNA adenosine methylase (N(6)-methyladenosine) can be fused to the RNA targeting effector proteins of the invention to target transcripts of interest. This methylase causes reversible methylation, has a regulatory role, and can affect gene expression and cell fate decisions by modulating multiple RNA-related cellular pathways (Fu et al Nat Rev Genet. 2014;15(5):293-306).
[0154] C) Polyadenylation Polyadenylation of mRNA is important for the nuclear transport, translation efficiency, and stability of mRNA, all of which, as well as the polyadenylation process, depend on specific RBPs. Many eukaryotic mRNAs receive a 3' poly(A) tail of approximately 200 nucleotides after transcription. Polyadenylation involves various RNA-binding protein complexes that stimulate the activity of poly(A) polymerase (Minvielle-Sebastia L et al. Curr Opin Cell Biol. 1999;11:352-7). It is envisioned that the RNA targeting effector proteins provided herein can be used to interfere with or promote the interaction between RNA-binding proteins and RNA.
[0155] An example of a disease that has been linked to defective proteins involved in polyadenylation is oculopharyngeal muscular dystrophy (OPMD) (Brais B, et al. Nat Genet. 1998;18:164-7).
[0156] D) RNA nuclear export After pre-mRNA processing, mRNA is transported from the nucleus to the cytoplasm by a cellular mechanism that involves the generation of a carrier complex, which then translocates through nuclear pores and releases the mRNA in the cytoplasm, where it is subsequently recycled.
[0157] Overexpression of proteins that play a role in the nuclear export of RNA (such as TAP) has been shown to increase the nuclear export of transcripts that are normally inefficiently exported in Xenopus laevis (Katahira J, et al. EMBO J. 1999;18:2593-609).
[0158] E)mRNA localization mRNA localization ensures spatially regulated protein production. The localization of transcripts to specific regions of the cell can be ensured by localization elements. In a detailed embodiment, it is envisioned that the effector proteins described herein can be used to target the localization elements to the RNA of interest. The effector protein can be designed to bind to the target transcript and shuttle it to a location within the cell determined by its peptide signal tag. For example, more specifically, the RNA localization can be altered using an RNA targeting effector protein fused to one or more nuclear localization signals (NLS) and / or one or more nuclear export signals (NES).
[0159] Further examples of localization signals include the zipcode binding protein (ZBP1), which ensures β-actin localization to the cytoplasm in some asymmetric cell types, the KDEL retention sequence (localization to the endoplasmic reticulum), the nuclear export signal (localization to the cytoplasm), the mitochondrial targeting signal (localization to mitochondria), the peroxisomal targeting signal (localization to peroxisomes), and the m6A tag / YTHDF2 (localization to p-bodies). Another envisioned approach is the fusion of an RNA-targeting effector protein with a protein of known localization (e.g., membrane, synapse).
[0160] Alternatively, the effector protein of the present invention may be used, for example, for location-dependent knockdown. By fusing the effector protein with an appropriate localization signal, the effector can be targeted to a specific intracellular compartment. Only the target RNA in this compartment will be effectively targeted, while targets in other identical but different intracellular compartments will not be targeted, thereby achieving location-dependent knockdown.
[0161] F) Translation The RNA targeting effector proteins described herein can be used to enhance or suppress translation. It is anticipated that upregulation of translation is a highly robust method for controlling cellular circuits. Furthermore, for functional studies, protein translation screens may be advantageous compared to transcriptional upregulation screens, which have the drawback that upregulation of transcripts does not lead to increased protein production.
[0162] It is envisioned that the RNA targeting effector proteins described herein can be used to bring translation initiation factors such as EIF4G into proximity with the 5' untranslated repeat (5'UTR) of a messenger RNA of interest to drive translation (as described for non-reprogrammable RNA-binding proteins in De Gregorio et al. EMBO J. 1999;18(17):4865-74). As another example, the cytoplasmic poly(A) polymerase GLD2 can be recruited to a target mRNA by an RNA targeting effector protein. This can allow for directed polyadenylation of the target mRNA and thereby stimulate translation.
[0163] Similarly, the RNA targeting effector proteins envisioned herein can be used to block translation repressors of mRNA, such as ZBP1 (Huttelmaier S, et al. Nature. 2005; 438: 512-5). Binding of target RNA to the translation start site can directly affect translation.
[0164] Additionally, fusing an RNA targeting effector protein to a protein that stabilizes the mRNA, for example by preventing its degradation, such as an RNase inhibitor, can increase protein production from the transcript of interest.
[0165] It is envisioned that the RNA targeting effector proteins described herein may be used to bind to the 5UTR region of an RNA transcript and inhibit translation by preventing ribosome formation and translation initiation.
[0166] Furthermore, RNA-targeting effector proteins can be used to recruit Caf1, a component of the CCR4-NOT deadenylase complex, to target mRNAs, resulting in deadenylation of the target transcript and inhibition of protein translation.
[0167] For example, the RNA-targeting effector proteins of the present invention can be used to increase or decrease the translation of therapeutically relevant proteins. Examples of therapeutic applications in which RNA-targeting effector proteins can be used to down- or up-regulate translation are amyotrophic lateral sclerosis (ALS) and cardiovascular disorders. Reduced levels of the glial glutamate transporter EAAT2 have been reported in the ALS motor cortex and spinal cord, and multiple abnormal EAAT2 mRNA transcripts have also been reported in ALS brain tissue. Loss of EAAT2 protein and function is believed to be a major cause of excitotoxicity in ALS. Restoring EAAT2 protein levels and function can provide therapeutic benefits. Therefore, RNA-targeting effector proteins can be beneficially used to up-regulate EAAT2 protein expression, for example, by blocking translational repressors or stabilizing mRNA, as described above. Apolipoprotein A1 is a major protein component of high-density lipoprotein (HDL), and ApoA1 and HDL are generally considered to be anti-atherogenic. It is envisioned that RNA targeting effector proteins may be beneficially used to upregulate expression of ApoA1, for example, by blocking translational repressors or stabilizing mRNA as described above.
[0168] G) mRNA turnover Translation is closely linked to mRNA turnover and regulated mRNA stability. Specific proteins have been described to be involved in transcript stability (e.g., ELAV / Hu proteins in neurons, Keene JD, 1999, Proc Natl Acad Sci U S A. 96:5-7) and tristetraprolin (TTP). These proteins stabilize target mRNAs by protecting the message from degradation in the cytoplasm (Peng SS et al., 1988, EMBO J. 17:3461-70).
[0169] It is conceivable that the RNA targeting effector proteins of the present invention can be used to interfere with or promote the activity of proteins that stabilize mRNA transcripts, such that mRNA turnover is affected. For example, the RNA targeting effector proteins can be used to recruit human TTP to target RNAs, enabling adenylate-uridylate-rich element (AU-rich element)-mediated translational repression and target degradation. AU-rich elements are found in the 3'UTRs of many mRNAs encoding protooncogenes, nuclear transcription factors, and cytokines, and promote RNA stability. As another example, the RNA targeting effector protein can be fused to another mRNA stabilizing protein, HuR (Hinman MN and Lou H, Cell Mol Life Sci 2008;65:3168-81), and recruited to target transcripts to extend their lifespan or stabilize short-lived mRNAs.
[0170] It is further contemplated that the RNA targeting effector proteins described herein can be used to promote degradation of target transcripts, for example, by recruiting m6A methyltransferase to target transcripts and localizing the transcripts to P-bodies, thereby resulting in target degradation.
[0171] As yet another example, an RNA-targeting effector protein as described herein can be fused to the non-specific endonuclease domain PilT N-terminus (PIN), allowing it to be recruited to and degraded by the target transcript.
[0172] Patients with paraneoplastic neuropathy (PND)-associated encephalomyelitis and neuropathy develop autoantibodies against Hu proteins in tumors outside the central nervous system (Szabo A et al. 1991, Cell.; 67:325-33), which then cross the blood-brain barrier. It is envisioned that the RNA targeting effector proteins of the present invention may be used to interfere with the binding of autoantibodies to mRNA transcripts.
[0173] Patients with dystrophy type 1 (DM1), caused by a (CUG)n expansion in the 3'UTR of the myotonic dystrophy protein kinase (DMPK) gene, are characterized by the accumulation of such transcripts in the nucleus. It is anticipated that the RNA-targeting effector protein of the present invention fused to an endonuclease that targets the (CUG)n repeat can inhibit the accumulation of such abnormal transcripts.
[0174] H) Interaction with multifunctional proteins Some RNA-binding proteins bind to multiple sites on many RNAs and function in a variety of processes.For example, hnRNP A1 protein has been shown to bind to exon splicing silencer sequences to antagonize splicing factors, associate with telomere ends (thereby stimulating telomere activity), and bind to miRNA to promote Drosha-mediated processing, thereby affecting maturation.It is assumed that the RNA-binding effector protein of the present invention can interfere with the binding of RNA-binding proteins at one or more positions.
[0175] I) RNA folding RNA adopts a defined structure to exert its biological activity. Conformational transitions between alternative tertiary structures are crucial for many RNA-mediated processes. However, RNA folding can be associated with several problems. For example, RNA may tend to fold into and remain in an inappropriate alternative conformation, and / or the correct tertiary structure may not be sufficiently thermodynamically favorable compared to the alternative structures. The RNA targeting effector proteins of the present invention, particularly cleavage-deficient or dead RNA targeting proteins, may be used to direct the folding of (m)RNA and / or ensure its correct tertiary structure.
[0176] Use of RNA-targeting effector proteins in modulating cellular states In certain embodiments, C2c2 complexed with crRNA is activated upon binding to the target RNA and subsequently cleaves any nearby ssRNA targets (i.e., a "collateral" or "bystander" effect). Once primed by its cognate target, C2c2 can cleave other (non-complementary) RNA molecules. Such indiscriminate RNA cleavage can potentially cause cytotoxicity or otherwise affect cellular physiology or state.
[0177] Thus, in certain embodiments, a non-naturally occurring or engineered composition, vector system, or delivery system as described herein is used or for use in inducing cellular dormancy. In certain embodiments, a non-naturally occurring or engineered composition, vector system, or delivery system as described herein is used or for use in inducing cell cycle arrest. In certain embodiments, a non-naturally occurring or engineered composition, vector system, or delivery system as described herein is used or for use in reducing cell growth and / or cell proliferation. In certain embodiments, a non-naturally occurring or engineered composition, vector system, or delivery system as described herein is used or for use in inducing cellular anergy. In certain embodiments, a non-naturally occurring or engineered composition, vector system, or delivery system as described herein is used or for use in inducing cellular apoptosis. In certain embodiments, a non-naturally occurring or engineered composition, vector system, or delivery system as described herein is used or for use in inducing cellular necrosis. In certain embodiments, a non-naturally occurring or engineered composition, vector system, or delivery system as described herein is used or for use in inducing cell death. In certain embodiments, a non-naturally occurring or engineered composition, vector system, or delivery system as described herein is used or for use in inducing programmed cell death.
[0178] In certain embodiments, the present invention relates to a method for inducing cellular dormancy, comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein. In certain embodiments, the present invention relates to a method for inducing cell cycle arrest, comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein. In certain embodiments, the present invention relates to a method for reducing cell growth and / or cell proliferation, comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein. In certain embodiments, the present invention relates to a method for inducing cellular anergy, comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein. In certain embodiments, the present invention relates to a method for inducing cellular apoptosis, comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein. In certain embodiments, the present invention relates to a method for inducing cellular necrosis, comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein. In certain embodiments, the present invention relates to methods of inducing cell death comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein.In certain embodiments, the present invention relates to methods of inducing programmed cell death comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein.
[0179] The methods and uses as described herein may be therapeutic or prophylactic and may target specific cells, cell (sub)populations, or cell / tissue types. In particular, the methods and uses as described herein may be therapeutic or prophylactic and may target specific cells, cell (sub)populations, or cell / tissue types that express one or more target sequences, such as one or more specific target RNAs (e.g., ssRNAs). Without limitation, target cells may be, for example, cancer cells that express specific transcripts, for example, neurons of a given class, (immune) cells that cause, for example, autoimmunity, or cells infected with a specific (e.g., viral) pathogen, etc.
[0180] Thus, in certain embodiments, the present invention relates to a method of treating a pathological condition characterized by the presence of undersirable cells (host cells), comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein. In certain embodiments, the present invention relates to the use of a non-naturally occurring or engineered composition, vector system, or delivery system as described herein to treat a pathological condition characterized by the presence of undersirable cells (host cells). In certain embodiments, the present invention relates to a non-naturally occurring or engineered composition, vector system, or delivery system as described herein for use in treating a pathological condition characterized by the presence of undersirable cells (host cells). It should be understood that preferably, the CRISPR-Cas system targets a target specific to the undersirable cells. In certain embodiments, the present invention relates to the use of a non-naturally occurring or engineered composition, vector system, or delivery system as described herein to treat, prevent, or alleviate cancer. In certain embodiments, the present invention relates to a non-naturally occurring or engineered composition, vector system, or delivery system as described herein for use in the treatment, prevention, or alleviation of cancer. In certain embodiments, the present invention relates to a method of treating, preventing, or alleviating cancer, comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein. It should be understood that preferably, the CRISPR-Cas system targets a target specific to cancer cells. In certain embodiments, the present invention relates to the use of a non-naturally occurring or engineered composition, vector system, or delivery system as described herein to treat, prevent, or alleviate infection of cells by a pathogen.In certain embodiments, the present invention relates to a non-naturally occurring or engineered composition, vector system, or delivery system as described herein for use in treating, preventing, or alleviating infection of a cell by a pathogen. In certain embodiments, the present invention relates to a method of treating, preventing, or alleviating infection of a cell by a pathogen, comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein. It should be understood that preferably, the CRISPR-Cas system targets a target specific to a cell infected by the pathogen (e.g., a target derived from the pathogen). In certain embodiments, the present invention relates to the use of a non-naturally occurring or engineered composition, vector system, or delivery system as described herein for treating, preventing, or alleviating an autoimmune disorder. In certain embodiments, the present invention relates to a non-naturally occurring or engineered composition, vector system, or delivery system as described herein for use in treating, preventing, or alleviating an autoimmune disorder. In certain embodiments, the present invention relates to a method of treating, preventing, or alleviating an autoimmune disorder comprising introducing or inducing a non-naturally occurring or engineered composition, vector system, or delivery system as described herein. It should be understood that preferably, the CRISPR-Cas system targets specific targets (e.g., specific immune cells) to cells involved in the autoimmune disorder.
[0181] Use of RNA-targeting effector proteins in RNA or protein detection It is further envisioned that the RNA targeting effector proteins may be used for the detection of nucleic acids or proteins in biological samples, which may be cellular or cell-free.
[0182] In Northern blot assays, Northern blotting involves the size separation of RNA samples using electrophoresis. RNA targeting effector proteins can be used to specifically bind and detect target RNA sequences.
[0183] RNA targeting effector proteins can also be fused to fluorescent proteins (such as GFP) and used to track RNA localization in live cells. More specifically, the RNA targeting effector protein can be inactivated at the point where it no longer cleaves RNA. In detailed embodiments, to ensure more precise visualization, it is contemplated that split RNA targeting effector proteins may be used, whereby the signal depends on the binding of both subproteins. Alternatively, split fluorescent proteins can be used that are reconstituted upon binding of multiple RNA targeting effector protein complexes to the target transcript. It is further contemplated that transcripts may be targeted at multiple binding sites along the mRNA, thereby amplifying the true signal and enabling localized discrimination. As yet another alternative, fluorescent proteins may be reconstituted from split inteins.
[0184] RNA targeting effector proteins are suitable for use in determining, for example, the localization of RNA or specific splice variants, mRNA transcript levels, transcript up- or down-regulation, and disease-specific diagnosis. RNA targeting effector proteins can be used to visualize RNA in (live) cells, for example, using fluorescence microscopy or flow cytometry, such as fluorescence-activated cell sorting (FACS), which allows high-throughput screening of cells and the recovery of live cells after cell sorting. Furthermore, the expression levels of various transcripts can be simultaneously assessed under stress, such as the inhibition of cancer growth using molecular inhibitors or hypoxic conditions for cells. Another application could be tracking the localization of transcripts to synaptic junctions during neuronal stimulation using two-photon microscopy.
[0185] In a particular embodiment, the components or complexes of the invention as described herein can be used in multiplexed error-robust fluorescence in situ hybridization (MERFISH; Chen et al. Science; 2015; 348(6233)), e.g. with (fluorescently) labeled C2c2 effectors.
[0186] In vitro APEX labeling Cellular processes depend on a network of molecular interactions between proteins, RNA, and DNA. Accurate detection of protein-DNA and protein-RNA interactions is important for understanding such processes. In vitro proximity labeling techniques utilize affinity tags in combination with, for example, photoactivatable probes to label polypeptides and RNAs in vitro near a protein or RNA of interest. After UV irradiation, the photoactivatable groups react with proteins and other molecules in close proximity to the tagged molecule, thereby labeling them. The labeled interacting molecules can then be recovered and identified. The RNA targeting effector proteins of the present invention can be used, for example, to target probes to selected RNA sequences.
[0187] These applications may also be applicable in animal models for disease-related applications or in vivo imaging of difficult-to-culture cell types.
[0188] The present invention provides agents and methods for diagnosing and monitoring health conditions through noninvasive sampling of cell-free RNA, including risk assessment and guidance for RNA-targeted therapy, and is useful in situations where rapid administration of treatment is critical to treatment outcome. In one embodiment, the present invention provides cancer detection methods and agents related to circulating tumor RNA, including monitoring for recurrence and / or the occurrence of common drug resistance mutations. In another embodiment, the present invention provides detection methods and agents for detecting and / or identifying bacterial species directly from blood or serum, thereby monitoring, for example, disease progression and sepsis. In one embodiment of the present invention, C2c2 proteins and derivatives are used to diagnose and distinguish common illnesses such as rhinovirus infections or upper respiratory tract infections from more serious infections such as bronchitis.
[0189] The present invention provides methods and agents for rapid genotyping towards emergency pharmacogenomics, including guiding the administration of anticoagulants based on, for example, VKORC1, CYP2C9, and CYP2C19 genotyping in the treatment of myocardial infarction or stroke.
[0190] The present invention provides agents and methods for monitoring bacterial contamination of food products at any point along the food production and distribution chain. In another embodiment, the present invention provides quality control and monitoring, for example, by determining the identity and purity of food ingredients. In one non-limiting example, the present invention may be used to identify or verify food ingredients, such as animal meats and seafood species.
[0191] In another embodiment, the present invention is used in forensic determination, for example, crime scene samples containing blood or other bodily fluids. In one embodiment of the present invention, the present invention is used to identify nucleic acid samples from fingerprints.
[0192] Use of RNA targeting effector proteins in RNA origami / in vitro assembly lines - Combinatrix RNA origami refers to nanoscale folding of structures using RNA as an integrated template to create two-dimensional or three-dimensional structures. The folding structure is encoded by RNA, and therefore the shape of the resulting RNA is determined by the synthesized RNA sequence (Geary, et al. 2014. Science, 345 (6198), pp. 799-804). RNA origami can serve as a scaffold for arranging other components, such as proteins, into a complex. Using the RNA targeting effector protein of the present invention, for example, a protein of interest can be targeted to the RNA origami using a suitable guide RNA.
[0193] Use of RNA targeting effector proteins in RNA isolation or purification, enrichment or depletion It is further envisioned that RNA-targeting effector proteins, when complexed with RNA, can be used to isolate and / or purify RNA. For example, the RNA-targeting effector protein can be fused to an affinity tag that can be used to isolate and / or purify the RNA-RNA-targeting effector protein complex. Such applications are useful, for example, in analyzing gene expression profiles in cells. In particular embodiments, it is envisioned that RNA-targeting effector proteins can be used to target specific non-coding RNAs (ncRNAs), thereby blocking their activity and providing useful functional probes. In certain embodiments, effector proteins as described herein can be used to specifically enrich specific RNAs (including, but not limited to, increasing stability) or specifically deplete specific RNAs (e.g., without limitation, specific splice variants, isoforms, etc.).
[0194] Examination of lincRNA function and other nuclear RNAs Current RNA knockdown strategies, such as siRNA, have the disadvantage that the protein machinery is cytoplasmic, so they are mostly limited to targeting cytoplasmic transcripts. The advantage of the RNA-targeting effector protein of the present invention, an exogenous system that is not essential for cellular function, is that it can be used in any compartment of the cell. By fusing an NLS signal to the RNA-targeting effector protein, it can be directed to the nucleus, enabling targeting of nuclear RNA. For example, probing the function of lincRNAs is envisioned. Long intergenic non-coding RNAs (lincRNAs) are a highly extensively investigated research area. Many lincRNAs have as yet unexplored functions, which could potentially be studied using the RNA-targeting effector proteins of the present invention.
[0195] Identification of RNA-binding proteins Identifying proteins that bind to specific RNAs can be useful for understanding the role of many RNAs. For example, many lincRNAs are involved in transcriptional and epigenetic regulation of transcription. Understanding which proteins bind to a given lincRNA can help elucidate the components of a given regulatory pathway. The RNA targeting effector proteins of the present invention can be designed to recruit biotin ligase to specific transcripts, thereby locally labeling the bound protein with biotin. The protein can then be pulled down and analyzed by mass spectrometry to identify it.
[0196] Assembly of the complex onto RNA and substrate shuttling Furthermore, the RNA targeting effector proteins of the present invention can be used to assemble complexes on RNA. This can be achieved by functionalizing the RNA targeting effector protein with multiple related proteins (e.g., components of a specific synthetic pathway). Alternatively, multiple RNA targeting effector proteins can be functionalized with such various related proteins to target the same or adjacent target RNAs. A useful application of the assembly of complexes on RNA is, for example, promoting substrate shuttling between proteins.
[0197] synthetic biology The development of biological systems has broad utility, including clinical applications. It is envisioned that the programmable RNA-targeting effector proteins of the present invention can be used to fuse toxic domains to split proteins for targeted cell death using, for example, cancer-associated RNAs as target transcripts. Furthermore, in synthetic biology systems, pathways involving protein-protein interactions can be influenced by fusion complexes with appropriate effectors, such as kinases or other enzymes.
[0198] Protein splicing: inteins Protein splicing is a post-translational process in which an intervening polypeptide, called an intein, catalyzes its excision from adjacent polypeptides, called exteins, and subsequent ligation of the exteins. The assembly of two or more RNA-targeting effector proteins as described herein onto a target transcript can be used to direct the release of a split intein (Topilina and Mills Mob DNA. 2014 Feb 4;5(1):5), thereby enabling direct accounting of the presence of mRNA transcripts and subsequent release of a protein product, such as a metabolic enzyme or transcription factor (for downstream operation of a transcription pathway). This application may have significant implications in synthetic biology (see above) or large-scale bioproduction (product production only under certain conditions).
[0199] Inducible, Administered, and Self-Inactivating Systems In one embodiment, a fusion complex comprising an RNA targeting effector protein of the invention and an effector component is designed to be inducible, e.g., light- or chemically inducible, allowing the effector component to be activated at a desired time.
[0200] Light-inducibility can be achieved, for example, by designing fusion complexes that utilize CRY2PHR / CIBN pairing for fusion, a system particularly useful for light-induction of protein interactions in living cells (Konermann S, et al. Nature. 2013;500:472-476).
[0201] Chemical inducibility can be provided, for example, by designing fusion complexes that use FKBP / FRB (FK506 binding protein / FKBP rapamycin binding) pairing for fusion. Using this system, rapamycin is required for protein binding (Zetsche et al. Nat Biotechnol. 2015;33(2):139-42 describes the use of this system with Cas9).
[0202] Furthermore, when introduced into cells as DNA, the RNA targeting effector proteins of the invention can be regulated by inducible promoters, such as tetracycline- or doxycycline-controlled transcriptional activation (Tet-On and Tet-Off expression systems), hormone-inducible gene expression systems, such as ecdysone-inducible gene expression systems, and arabinose-inducible gene expression systems. When delivered as RNA, expression of the RNA targeting effector proteins can be regulated by riboswitches, which can sense small molecules such as tetracycline (as described in Goldfless et al. Nucleic Acids Res. 2012;40(9):e64).
[0203] In one embodiment, delivery of the RNA targeting effector proteins of the present invention can be modulated to alter the amount of protein or crRNA in the cell, thereby altering the magnitude of the desired effect or any undesired off-target effects.
[0204] In one embodiment, the RNA targeting effector proteins described herein can be designed to self-inactivate, which when delivered to a cell as RNA, either mRNA or a replication RNA therapeutic (Wrobleska et al Nat Biotechnol. 2015 Aug;33(8):839-841), can self-inactivate expression and subsequent effects by destroying self-RNA, thereby reducing resident and potentially undesirable effects.
[0205] For further in vivo applications of RNA targeting effector proteins as described herein, see Mackay JP et al (Nat Struct Mol Biol. 2011 Mar;18(3):256-61), Nelles et al (Bioessays. 2015 Jul;37(7):732-9), and Abil Z and Zhao H (Mol Biosyst. 2015 Oct;11(10):2658-65), which are incorporated herein by reference. In particular, in certain embodiments of the invention, preferably by using catalytically inactive C2c2, the following applications are envisaged: translation enhancement (e.g., C2c2-translation enhancing factor fusions (e.g., eIF4 fusions)); translation repression (e.g., gRNAs targeting ribosome binding sites); exon skipping (e.g., gRNAs targeting splice donor and / or acceptor sites); exon inclusion (e.g., gRNAs or spliceosome components (e.g., U1) targeting specific exon splice donor and / or acceptor sites for inclusion). snRNA); access to RNA localization (e.g., C2c2-marker fusions (e.g., EGFP fusions)); alteration of RNA localization (e.g., C2c2-localization signal fusions (e.g., NLS or NES fusions)); RNA degradation (in this case, catalytically inactive C2c2 would not be used if relying on C2c2 activity, or, alternatively, for increased specificity, split C2c2 may be used); inhibition of non-coding RNA function (e.g., miRNA), such as by degradation of gRNA or binding to functional sites (potentially titrated out at specific sites by relocalization via C2c2-signal sequence fusions).
[0206] As described hereinabove and demonstrated in the Examples, C2c2 function is robust to 5' or 3' extension of the crRNA and extension of the crRNA loop. Therefore, it is envisioned that MS2 loops and other recruitment domains can be added to the crRNA without affecting complex formation and binding to target transcripts. Such modifications to the crRNA to recruit various effector domains are applicable in the use of the RNA-targeting effector proteins described above.
[0207] As demonstrated in the Examples, C2c2, specifically LshC2c2, has the ability to mediate RNA phage resistance. It is therefore envisioned that C2c2 can be used to immunize, for example, animals, humans, and plants against RNA-only pathogens, including but not limited to retroviruses (e.g., lentiviruses such as HIV), HCV, Ebola virus, and Zika virus.
[0208] We have shown that C2c2 can process (cleave) its own array. This applies to both wild-type C2c2 protein and mutant C2c2 proteins containing one or more mutant amino acid residues R597, H602, R1278, and H1283, such as one or more of the modifications selected from R597A, H602A, R1278A, and H1283A. It is therefore envisioned that multiple crRNAs designed for different target transcripts and / or applications can be delivered as a single pre-crRNA or as a single transcript driven by one promoter. Such a delivery method has the advantages of being substantially more compact, easier to synthesize, and easier to deliver in viral systems. Preferably, the amino acid numbering as described herein refers to the Lsh C2c2 protein. It will be understood that the exact amino acid positions for orthologs of Lsh C2c2 may vary and can be appropriately determined by protein alignment, as known in the art and described elsewhere herein.
[0209] Aspects of the present invention also include methods and uses of the compositions and systems described herein for, for example, altering or manipulating (protein) expression of one or more genes or one or more gene products in genome or transcriptome engineering in vitro, in vivo or ex vivo in prokaryotic or eukaryotic cells.
[0210] In one aspect, the present invention provides methods and compositions for modulating, e.g., reducing, target RNA (protein) expression in a cell, wherein the C2c2 system of the present invention is provided to interfere with RNA transcription, stability, and / or translation.
[0211] In certain embodiments, an effective amount of the C2c2 system is used to cleave RNA or otherwise inhibit RNA expression. In this regard, this system has similar uses to siRNA and shRNA, and therefore can also replace such methods. This method includes, but is not limited to, the use of the C2c2 system instead of an interfering ribonucleic acid (such as siRNA or shRNA) or its transcription template, such as a DNA encoding shRNA. The C2c2 system is introduced into target cells, for example, by administration to a mammal containing the target cells.
[0212] Advantageously, the C2c2 systems of the present invention are specific: for example, whereas interfering ribonucleic acid (such as siRNA or shRNA) polynucleotide systems suffer from design and stability issues as well as off-target binding, the C2c2 systems of the present invention can be designed with high specificity.
[0213] Destabilized C2c2 In certain embodiments, an effector protein of the present invention as described herein (CRISPR enzyme; C2c2) is associated with or fused to a destabilization domain (DD). In some embodiments, the DD is ER50. The corresponding stabilizing ligand of this DD is, in some embodiments, 4HT. Thus, in some embodiments, one of the at least one DD is ER50, and thus the stabilizing ligand is 4HT or CMP8. In some embodiments, the DD is DHFR50. The corresponding stabilizing ligand of this DD is, in some embodiments, TMP. Thus, in some embodiments, one of the at least one DD is DHFR50, and thus the stabilizing ligand is TMP. In some embodiments, the DD is ER50. The corresponding stabilizing ligand of this DD is, in some embodiments, CMP8. Thus, CMP8 can be an alternative stabilizing ligand to 4HT in the ER50 system. While it is possible that CMP8 and 4HT can / should be used competitively, some cell types may be more susceptible to one of these two ligands, and given this disclosure and knowledge in the art, one skilled in the art can use CMP8 and / or 4HT.
[0214] In some embodiments, one or two DDs may be fused to the N-terminal end of the CRISPR enzyme, and one or two DDs may be fused to the C-terminal end of the CRISPR enzyme. In some embodiments, at least two DDs are associated with the CRISPR enzyme, and the DDs are the same DD, i.e., the DDs are homologous. Thus, both (or two or more) of the DDs may be ER50 DDs. This is preferred in some embodiments. Alternatively, both (or two or more) of the DDs may be DHFR50 DDs. This is also preferred in some embodiments. In some embodiments, at least two DDs are associated with the CRISPR enzyme, and the DDs are different DDs, i.e., the DDs are heterologous. Thus, one of the DDs may be ER50, while one or more of the DDs or any other DD may be DHFR50. Having two or more heterologous DDs may be advantageous, as it may result in a higher level of degradation control. Tandem fusion of two or more DDs at the N- or C-terminus can enhance degradation; and such tandem fusions can be, for example, ER50-ER50-C2c2 or DHFR-DHFR-C2c2. It is envisioned that high levels of degradation occur in the absence of either stabilizing ligand, intermediate levels of degradation can occur in the absence of one stabilizing ligand and the presence of the other (or another) stabilizing ligand, while low levels of degradation can occur in the presence of both (or more) stabilizing ligands. Control can also be provided by having an N-terminal ER50 DD and a C-terminal DHFR50 DD.
[0215] In some embodiments, the fusion of the CRISPR enzyme and the DD includes a linker between the DD and the CRISPR enzyme. In some embodiments, the linker is a GlySer linker. In some embodiments, the DD-CRISPR enzyme further includes at least one nuclear export signal (NES). In some embodiments, the DD-CRISPR enzyme includes two or more NESs. In some embodiments, the DD-CRISPR enzyme includes at least one nuclear localization signal (NLS), which may be included in addition to an NES. In some embodiments, the CRISPR enzyme includes, consists essentially of, or consists of a localization (nuclear import or nuclear export) signal as, or as part of, the linker between the CRISPR enzyme and the DD. HA or Flag tags are also within the scope of the present invention as linkers. Applicants use NLSs and / or NESs as linkers, as well as glycine-serine linkers as short as GS up to (GGGGS)3.
[0216] Destabilizing domains have general utility in conferring instability to a wide range of proteins; see, e.g., Miyazaki, J Am Chem Soc. Mar 7, 2012;134(9):3942-3945 (incorporated herein by reference). CMP8 or 4-hydroxytamoxifen can be destabilizing domains. More generally, a temperature-sensitive mutant of mammalian DHFR (DHFRts), which contains a destabilizing residue due to the N-end rule, was found to be stable at permissive temperatures but unstable at 37°C. Addition of methotrexate, a high-affinity ligand for mammalian DHFR, to cells expressing DHFRts partially inhibited protein degradation. This was an important demonstration that small molecule ligands can stabilize proteins that are normally targeted for degradation in cells. A rapamycin derivative stabilized a destabilizing mutant of the FRB domain of mTOR (FRB*) and restored function of the fused kinase GSK-3β. 6,7 This system demonstrated that ligand-dependent stability represents an attractive strategy for modulating the function of specific proteins in complex biological environments. Regulation of protein activity may involve a DD that becomes functional upon ubiquitin complementation via rapamycin-induced dimerization of FK506-binding protein with FKBP12. Mutants of human FKBP12 or ecDHFR proteins can be engineered to be metabolically unstable in the absence of their high-affinity ligands, Shield-1 or trimethoprim (TMP), respectively. These mutants are some of the potential destabilization domains (DDs) useful in the practice of the present invention, and the instability of the DD as a fusion with a CRISPR enzyme leads to CRISPR proteolysis of the entire fusion protein by the proteasome. Shield-1 and TMP bind to and stabilize the DD in a dose-dependent manner. The estrogen receptor ligand-binding domain (ERLBD, residues 305–549 of ERS1) can also be engineered as a destabilization domain. Because the estrogen receptor signaling pathway is involved in various diseases, including breast cancer, this pathway has been extensively studied and numerous estrogen receptor agonists and antagonists have been developed.Thus, compatible pairs of ERLBD and drugs are known. There are ligands that bind to mutant forms of ERLBD but not to wild-type forms of ERLBD. By using one of these mutant domains, encoding three mutations (L384M, M421G, G521R),12 it is possible to modulate the stability of the ERLBD-derived DD with a ligand that does not disrupt the endogenous estrogen-sensitive network. An additional mutation (Y537S) can be introduced to further destabilize the ERLBD, making it a potential DD candidate. This quadruple mutant is an advantageous DD development. This mutant ERLBD can be fused to a CRISPR enzyme, and its stability can be modulated or disrupted using a ligand, resulting in the CRISPR enzyme having a DD. Another DD could be a 12 kDa (107 amino acids) tag based on a mutant FKBP protein that is stabilized by the Shield1 ligand; see, for example, Nature Methods 5, (2008).For example, the DD can be a modified FK506-binding protein 12 (FKBP12) that binds to and is reversibly stabilized by a synthetic, biologically inert small molecule, Shield-1; see, e.g., Banaszynski LA, Chen LC, Maynard-Smith LA, Ooi AG, Wandless TJ. "A rapid, reversible, and tunable method to regulate protein function in living cells using synthetic small molecules." Cell. 2006;126:995-1004; Banaszynski LA, Sellmyer MA, Contag CH, Wandless TJ, Thorne SH. "Chemical control of protein stability and function in living mice." Nat Med. 2008;14:1123-1127; Maynard-Smith LA, Chen See LC, Banaszynski LA, Ooi AG, Wandless TJ. "A directed approach for engineering conditional protein stability using biologically silent small molecules." The Journal of biological chemistry. 2007;282:24866-24872; and Rodriguez, Chem Biol. Mar 23, 2012;19(3):391-398, all of which are incorporated herein by reference, and may be used in the practice of the invention in selecting a DD to associate with a CRISPR enzyme.As can be appreciated, the art includes numerous DDs, which can be associated with, e.g., fused to, a CRISPR enzyme, preferably with a linker, such that the DD can be stabilized in the presence of a ligand and destabilized in its absence, thereby destabilizing the CRISPR enzyme as a whole, or the DD can be stabilized in the absence of a ligand and destabilized in the presence of a ligand; the DD allows for the CRISPR enzyme and, in turn, the CRISPR-Cas complex or system to be regulated or controlled—in effect, switched on or off, so to speak—providing a means of regulating or controlling the system, e.g., in vivo or in vitro. For example, expressing a protein of interest as a fusion with a DD tag destabilizes it in cells and causes rapid degradation, e.g., by the proteasome. Thus, the absence of a stabilizing ligand leads to degradation of the D-associated Cas. Fusing a novel DD to a protein of interest confers instability to the protein of interest, resulting in rapid degradation of the entire fusion protein. Peak Cas activity is sometimes beneficial for reducing off-target effects. Therefore, a short burst of high activity is desirable. The present invention can provide such a peak. In one sense, the system is inducible. In another sense, the system is repressed in the absence of a stabilizing ligand and derepressed in the presence of a stabilizing ligand.
[0217] Application of RNA-targeting CRISPR systems to plants and yeast Definition: In general, the term "plant" refers to any of a variety of photosynthetic, eukaryotic, unicellular or multicellular organisms of the kingdom Plantae that grow characteristically by cell division, contain chloroplasts, and have cell walls composed of cellulose. The term plant includes monocotyledonous and dicotyledonous plants. Specific plants include, without limitation, acacia, alfalfa, amaranth, apple, apricot, artichoke, ash tree, asparagus, avocado, banana, barley, beans, sugar beet, birch, beech, blackberry, blueberry, broccoli, Brussels sprouts, cabbage, canola, cantaloupe, carrot, cassava, cauliflower, cedar, grains, celery, chestnut, cherry, Chinese cabbage, citrus fruits, clementine, clover, coffee, corn, cotton, cowpea, cucumber, cypress, eggplant, elm, endive, eucalyptus, fennel, fig, fir, geranium, grape, grapefruit, peanut, ground cherry, gum hemlock, and the like. hemlock), hickory, kale, kiwi fruit, kohlrabi, larch, lettuce, chives, lemon, lime, black locust, pine, maidenhair, corn, mango, maple, melon, millet, mushrooms, mustard, nuts, oak, oats, oil palm, okra, onion, orange, ornamental plants or decorative flowers or trees, papaya, palm, parsley, parsnip, pea, peach, peanut, pear, peat, pepper, persimmon, pigeon pea, pine, pineapple, plantain, It is intended to include angiosperms and gymnosperms such as plum, pomegranate, potato, pumpkin, radish, rapeseed, raspberry, rice, rye, sorghum, safflower, wild willow, soybean, spinach, spruce, pumpkin, strawberry, sugar beet, sugarcane, sunflower, sweet potato, sweet corn, tangerine, tea, tobacco, tomato, trees, triticale, turfgrass, turnip, vines, walnut, watercress, watermelon, wheat, yam, yew, and zucchini. The term plant also encompasses algae, which are mostly photoautotrophs, united primarily by their lack of roots, leaves, and other organs that characterize higher plants.
[0218] The method of regulating gene expression using the RNA targeting system as described herein can be used to confer desired traits to essentially any plant. A wide variety of plants and plant cell lines can be engineered for the desired physiological and agronomic characteristics described herein using the nucleic acid constructs of the present disclosure and the various transformation methods described above. In preferred embodiments, plants and plant cells targeted for engineering include monocotyledonous and dicotyledonous plants, such as crops including, but not limited to, cereal crops (e.g., wheat, corn, rice, millet, barley), fruit crops (e.g., tomato, apple, pear, strawberry, orange), forage crops (e.g., alfalfa), root crops (e.g., carrot, potato, sugar beet, yam), leafy vegetable crops (e.g., lettuce, spinach); flowering plants (e.g., petunia, rose, chrysanthemum), coniferous and pine trees (e.g., fir, spruce); plants used in phytoremediation (e.g., heavy metal accumulating plants); oil crops (e.g., sunflower, rapeseed) and plants used for experimental purposes (e.g., Arabidopsis). Thus, the present methods and CRISPR-Cas systems can be used across a wide range of plants, including, for example, Magniolales, Illiciales, Laurales, Piperales, Aristochiales, Nymphaeales, Ranunculales, Papeverales, Sarraceniaceae, Trochodendrales, Mammales, and the like. Hamamelidales, Eucomiales, Leitneriales, Myricales, Fagales, Casuarinales, Caryophyllales, Batales, Polygonales, Plumbaginales, Dilleniales, Theales, Malvales,Urticales (Urticales), Lecythidales (Lecythidales), Violales (Violales), Salicales (Salicales), Capparales (Capparales), Ericales (Ericales), Diapensales (Diapensales), Ebenales (Ebenales), Primulaceae (Primulales), Rosales (Rosales), Fabaceae (Fabales), Podostemales (Podostemales), Haloragales (Haloragales), Myrtales (Myrtales), Cornales (Cornales), Proteales (Proteales), Sandaleales (San tales, Rafflesiales, Celastraceae, Euphorbiales, Rhamnales, Sapindales, Juglandales, Geraniales, Polygalales, Umbellales, Gentianales, Polemoniales, Lamiales, Plantaginales, Scrophulariales, Campanulales, Rubiales, Dipsacales, and Asterales ), and dicotyledonous plants belonging to the following orders: Alismatales, Hydrocharitales, Najadales, Triuridales, Commelinales, Eriocaulales, Restionales, Poales, Juncales, Cyperales, Typhales, Bromeliales, Zingiberales, Arecales, Cyclanthales, Pandanales, Arales,Monocotyledons such as those belonging to the orders Liliales and Orchids, or Gymnospermae, such as those belonging to the orders Pinales, Ginkgoales, Cycadales, Araucariales, Cupressales, and Gnetales, can be used.
[0219] The RNA-targeting CRISPR systems and methods of use described herein can be used in a wide range of plant species within the following non-limiting list of dicotyledonous, monocotyledonous, or gymnosperm genera: Atropa, Alseodaphne, Anacardium, Arachis, Beilschmiedia, Brassica, Carthamus, and Mentha. Cocculus, Croton, Cucumis, Citrus, Citrullus, Capsicum, Catharanthus, Cocos, Coffea, Cucurbita, Daucus, Duguetia, Eschscholzia, Ficus, Fragaria, Horn poppy Glaucium, Glycine, Gossypium, Helianthus, Hevea, Hyoscyamus, Lactuca, Landolphia, Linum, Litsea, Lycopersicon, Lupin, Manihot, Majorana, Malus, Medicago , Nicotiana, Olea, Parthenium, Papaver, Persea, Phaseolus, Pistacia, Pisum, Pyrus, Prunus, Raphanus, Ricinus, Senecio, Sinomenium, Stephania, Sinapis,Solanum, Theobroma, Trifolium, Trigonella, Vicia, Vinca, Vilis, and Vigna; and Allium, Andropogon, Aragrostis, Asparagus, Avena, Cynodon, Elaeis, Festuca, Festulolium, Daylily, and Daylily. Heterocallis, barley (Hordeum), duckweed (Lemna), Lolium, Musa, rice (Oryza), millet (Panicum), pennesetum, timothy grass (Phleum), Poa, rye (Secale), sorghum (Sorghum), wheat (Triticum), corn (Zea), fir (Abies), Chinese fir (Cunninghamia), ephedra, spruce (Picea), pine (Pinus), and Pseudotsuga.
[0220] The RNA targeting CRISPR systems and methods of use can also be used with a wide range of "algae" or "algal cells," including, for example, algea selected from several eukaryotic phyla, including Rhodophyta (red algae), Chlorophyta (green algae), Phaeophyta (brown algae), Bacillariophyta (diatoms), Eustigmatophyta, and Dinoflagellates, and the prokaryotic phylum Cyanobacteria (blue-green algae).The term "algae" includes, for example, species of the genera Amphora, Anabaena, Anikstrodesmis, Botryococcus, Chaetoceros, Chlamydomonas, Chlorella, Chlorococcum, Cyclotella, Cylindrotheca, Dronata, and the like. Dunaliella, Emiliana, Euglena, Hematococcus, Isochrysis, Monochrysis, Monoraphidium, Nannochloris, Nannochloropsis, Navicula, Nep hrochloris, Nephroselmis, Nitzschia, Nodularia, Nostoc, Oochromonas, Oocystis, Oscillartoria, Pavlova, Phaeodactylum, Platymonas, Pleurococcus hrysis, Porhyra, Pseudoanabaena, Pyramimonas, Stichococcus, Synechococcus, Synechocystis, Tetraselmis, Thalassiosira, and Trichodesmium.
[0221] Plant parts, i.e., "plant tissue," can be treated according to the methods of the present invention to produce improved plants. Plant tissue also includes plant cells. The term "plant cell," as used herein, refers to an individual unit of a living plant, either an intact whole plant, or in isolated form grown in in vitro tissue culture on media or agar, in suspension in a growth medium or buffer, or as part of a more highly organized unit, such as a plant tissue, plant organ, or whole plant.
[0222] "Protoplast" refers to a plant cell whose protective cell wall has been completely or partially removed, e.g., using mechanical or enzymatic means, resulting in an intact, biochemically competent unit of a living plant that can reform its cell wall, grow, and regenerate to develop into a whole plant under appropriate growth conditions.
[0223] The term "transformation" broadly refers to a method in which a plant host is genetically modified by the introduction of DNA using Agrobacteria or one of a variety of chemical or physical methods. As used herein, the term "plant host" refers to a plant, including any cell, tissue, organ, or progeny of the plant. Many suitable plant tissues or plant cells can be transformed, including, but not limited to, protoplasts, somatic embryos, pollen, leaves, seedlings, stems, callus, stolon, microtubules, and shoots. Plant tissue also refers to any clones of such plants, seeds, progeny, propagules, and progeny of any of these, such as cuttings or seeds, whether produced sexually or asexually.
[0224] The term "transformed," as used herein, refers to a cell, tissue, organ, or organism into which an exogenous DNA molecule, such as a construct, has been introduced. The introduced DNA molecule may be integrated into the genomic DNA of the recipient cell, tissue, organ, or organism such that the introduced DNA molecule is passed on to subsequent progeny. In these embodiments, a "transformed" or "transgenic" cell or plant can also include the progeny of that cell or plant, and progeny produced from breeding programs using such transformed plants as parents in crosses and exhibiting phenotypic changes resulting from the presence of the introduced DNA molecule. Preferably, the transgenic plant is fertile, capable of passing on the introduced DNA to progeny through sexual reproduction.
[0225] The term "progeny," such as the offspring of a transgenic plant, is one that is born from, produced from, or derived from a plant or transgenic plant. The introduced DNA molecule may also be transiently introduced into a recipient cell, such that the introduced DNA molecule is not inherited by subsequent progeny and is therefore not considered "transgenic." Thus, as used herein, a "non-transgenic" plant or plant cell is one that does not contain foreign DNA stably integrated into its genome.
[0226] The term "plant promoter," as used herein, is a promoter capable of initiating transcription in plant cells, regardless of whether its origin is a plant cell. Exemplary suitable plant promoters include, but are not limited to, those obtained from plants, plant viruses, and bacteria such as Agrobacterium or Rhizobium that contain genes that are expressed in plant cells.
[0227] As used herein, "fungal cell" refers to any type of eukaryotic cell within the kingdom Fungi. Phylums within the kingdom Fungi include Ascomycota, Basidiomycota, Blastocladiomycota, Chytridiomycota, Glomeromycota, Microsporidia, and Neocallimastigomycota. Fungal cells can include yeast, mold, and filamentous fungi. In some embodiments, the fungal cell is a yeast cell.
[0228] As used herein, the term "yeast cell" refers to any fungal cell within the phyla Ascomycota and Basidiomycota. Yeast cells can include budding yeast cells, fission yeast cells, and mold cells. Many types of yeast used in laboratory and industrial settings are part of the phylum Ascomycota, although not limited to these organisms. In some embodiments, the yeast cell is a S. cerevisiae, Kluyveromyces marxianus, or Issatchenkia orientalis cell. Other yeast cells include, without limitation, Candida spp. (e.g., Candida albicans), Yarrowia spp. (e.g., Yarrowia lipolytica), Pichia spp. (e.g., Pichia pastoris), Kluyveromyces spp. (e.g., Kluyveromyces lactis and Kluyveromyces marxianus), Neurospora spp. (e.g., Neurospora crassa), Fusarium spp. spp. (e.g., Fusarium oxysporum), and Issatchenkia spp. (e.g., Issatchenkia orientalis, also known as Pichia kudriavzevii and Candida acidothermophilum). In some embodiments, the fungal cell is a filamentous fungal cell. As used herein, the term "filamentous fungal cell" refers to any type of fungal cell that grows in a filamentous manner, i.e., as a hypha or mycelium.Examples of filamentous fungal cells include, without limitation, Aspergillus spp. (e.g., Aspergillus niger), Trichoderma spp. (e.g., Trichoderma reesei), Rhizopus spp. (e.g., Rhizopus oryzae), and Mortierella spp. (e.g., Mortierella isabellina).
[0229] In some embodiments, the fungal cell is an industrial strain. As used herein, "industrial strain" refers to any strain of fungal cell used in or isolated from an industrial process, e.g., the production of a product on a commercial or industrial scale. An industrial strain may refer to a fungal species typically used in industrial processes, or it may refer to an isolate of a fungal species that may also be used for non-industrial purposes (e.g., laboratory research). Examples of industrial processes can include fermentation (e.g., in the production of food or beverage products), distillation, biofuel production, chemical compound production, and polypeptide production. Examples of industrial strains can include, without limitation, JAY270 and ATCC4124.
[0230] In some embodiments, the fungal cell is a polyploid cell. As used herein, a "polyploid" cell can refer to any cell in which the genome exists in two or more copies. A polyploid cell can refer to a type of cell naturally found in a polyploid state, or it can refer to a cell that has been induced to exist in a polyploid state (e.g., by specific regulation, alteration, inactivation, activation, or modification of meiosis, cytokinesis, or DNA replication). A polyploid cell can refer to a cell whose entire genome is polyploid, or it can refer to a cell that is polyploid at a particular genomic locus of interest. Without wishing to be bound by theory, it is believed that the methods using the C2c2 CRISPRS system described herein can take advantage of the use of certain fungal cell types, as guide RNA abundance can often be a rate-limiting component in genome engineering of polyploid cells compared to haploid cells.
[0231] In some embodiments, the fungal cell is a diploid cell. As used herein, a "diploid" cell can refer to any cell in which the genome exists in two copies. A diploid cell can refer to a type of cell naturally found in a diploid state, or it can refer to a cell that has been induced to exist in a diploid state (e.g., by specific regulation, alteration, inactivation, activation, or modification of meiosis, cytokinesis, or DNA replication). For example, S. cerevisiae strain S228C can be maintained in a haploid or diploid state. A diploid cell can refer to a cell in which the entire genome is diploid, or it can refer to a cell that is diploid at a specific genomic locus of interest. In some embodiments, the fungal cell is a haploid cell. As used herein, a "haploid" cell can refer to any cell in which the genome exists in one copy. A haploid cell may refer to a type of cell naturally found in a haploid state, or it may refer to a cell that has been induced to exist in a haploid state (e.g., by specific regulation, alteration, inactivation, activation, or modification of meiosis, cytokinesis, or DNA replication). For example, S. cerevisiae strain S228C can be maintained in a haploid or diploid state. A haploid cell may refer to a cell whose entire genome is haploid, or it may refer to a cell that is haploid at a particular genomic locus of interest.
[0232] As used herein, "yeast expression vector" refers to a nucleic acid containing one or more sequences encoding an RNA and / or a polypeptide, and may further contain any desired elements that control expression of the nucleic acid(s), as well as any elements that allow for replication and maintenance of the expression vector within a yeast cell. Many suitable yeast expression vectors and their characteristics are known in the art; for example, various vectors and techniques are exemplified in "Yeast Protocols, 2nd edition," Xiao, W., ed. (Humana Press, New York, 2007) and Buckholz, RG and Gleeson, MA (1991) Biotechnology (NY) 9(11):1067-72. Yeast vectors may contain, without limitation, a centromere (CEN) sequence, an autonomously replicating sequence (ARS), a promoter such as an RNA polymerase III promoter operably linked to a sequence or gene of interest, a terminator such as an RNA polymerase III terminator, an origin of replication, and a marker gene (e.g., an auxotroph, antibiotic, or other selectable marker). Examples of expression vectors used in yeast include plasmids, yeast artificial chromosomes, 2μ plasmids, yeast integrating plasmids, yeast replicating plasmids, shuttle vectors, and episomal plasmids.
[0233] Stable integration of RNA-targeting CRISP system components in the genomes of plants and plant cells In particular embodiments, it is envisioned that polynucleotides encoding components of an RNA-targeting CRISPR system are introduced for stable integration into the genome of a plant cell. In these embodiments, the design of the transformation vector or expression system can be tailored depending on when, where, and under what conditions the guide RNA and / or one or more RNA-targeting genes are to be expressed.
[0234] In particular embodiments, it is contemplated that components of an RNA-targeting CRISPR system are stably introduced into the genomic DNA of a plant cell. Additionally or alternatively, it is contemplated that components of an RNA-targeting CRISPR system are introduced for stable integration into the DNA of a plant organelle, such as, but not limited to, a plastid, mitochondria, or chloroplast.
[0235] Expression systems for stable integration into the genome of plant cells may contain one or more of the following elements: a promoter element that can be used to express the guide RNA and / or RNA targeting enzyme in plant cells; a 5' untranslated region that enhances expression; an intron element that further enhances expression in certain cells, such as monocotyledonous plant cells; a multiple cloning site that provides convenient restriction sites for inserting one or more guide RNA and / or RNA targeting gene sequences and other desired elements; and a 3' untranslated region that provides efficient termination of the expressed transcript.
[0236] The elements of the expression system may be on one or more expression constructs that are either circular, such as a plasmid or transformation vector, or non-circular, such as linear double-stranded DNA. In particular embodiments, the RNA targeting CRISPR expression system comprises at least: (a) a nucleotide sequence encoding a guide RNA (gRNA) that hybridizes with a target sequence in a plant, the guide RNA comprising a guide sequence and a direct repeat sequence; and (b) a nucleotide sequence encoding an RNA targeting protein; Including, wherein components (a) or (b) are located on the same or different constructs, and whereby the different nucleotide sequences can be under the control of the same or different regulatory elements operable in plant cells.
[0237] One or more DNA constructs containing the components of the RNA-targeting CRISPR system, and, if applicable, the template sequence, can be introduced into the genome of a plant, plant part, or plant cell by a variety of conventional techniques. This process generally includes selecting a suitable host cell or tissue, introducing the one or more constructs into the host cell or tissue, and regenerating a plant cell or plant therefrom. In particular embodiments, DNA constructs may be introduced into plant cells using techniques such as, but not limited to, electroporation, microinjection, aerosol beam injection of plant cell protoplasts, or DNA constructs may be introduced directly into plant tissue using biolistic methods such as DNA particle bombardment (see also Fu et al., Transgenic Res. 2000 Feb;9(1):11-9). The basis of particle bombardment is to accelerate particles coated with one or more genes of interest toward the cell, allowing the particles to penetrate the cytoplasm and typically result in stable integration into the genome (see, e.g., Klein et al., Nature (1987); Klein et al., Bio / Technology (1992); Casas et al., Proc. Natl. Acad. Sci. USA (1993)).
[0238] In a detailed embodiment, a DNA construct containing components of an RNA-targeting CRISPR system can be introduced into a plant via Agrobacterium-mediated transformation. The DNA construct can be combined with suitable T-DNA flanking regions and introduced into a conventional Agrobacterium tumefaciens host vector. Foreign DNA can be incorporated into the plant genome by infecting the plant or by incubating plant protoplasts with Agrobacterium bacteria containing one or more Ti (tumor-inducing) plasmids (see, e.g., Fraley et al., (1985), Rogers et al., (1987), and U.S. Patent No. 5,563,055).
[0239] Plant promoters To ensure proper expression in plant cells, the components of the C2c2 CRISPR system described herein are typically placed under the control of a plant promoter, i.e., a promoter that is operable in plant cells. The use of various types of promoters is contemplated.
[0240] A constitutive plant promoter is a promoter that is capable of causing the open reading frame (ORF) it controls to be expressed in all or nearly all plant tissues at all or nearly all developmental stages of the plant (referred to as "constitutive expression"). One non-limiting example of a constitutive promoter is the cauliflower mosaic virus 35S promoter. The present invention also contemplates methods for modifying RNA sequences and thereby regulating the expression of plant biomolecules. Thus, in specific embodiments of the present invention, it is advantageous to place one or more elements of an RNA targeting CRISPR system under the control of a promoter that may be regulated. A "regulated promoter" refers to a promoter that directs gene expression in a non-constitutive but temporally and / or spatially regulated manner, including tissue-specific, tissue-preferred, and inducible promoters. Different promoters may direct the expression of genes in different tissues or cell types, at different developmental stages, or in response to different environmental conditions. In particular embodiments, one or more of the RNA targeting CRISPR components are expressed under the control of a constitutive promoter, such as the Cauliflower Mosaic Virus 35S promoter, and tissue-preferred promoters can be utilized to target enhanced expression to certain cell types within specific plant tissues, such as vascular cells in leaves or roots, or specific cells in seeds. For detailed examples of promoters used in RNA targeting CRISPR systems, see Kawamata et al., (1997) Plant Cell Physiol 38:792-803; Yamamoto et al., (1997) Plant J 12:255-65; Hire et al., (1992) Plant Mol Biol 20:207-18; Kuster et al., (1995) Plant Mol Biol 29:759-72, and Capana et al., (1994) Plant Mol Biol 25:681-91. Examples of promoters that are inducible and allow spatiotemporal control of gene editing or gene expression can use some form of energy.The form of energy can include, but is not limited to, acoustic energy, electromagnetic radiation, chemical energy, and / or thermal energy. Examples of inducible systems include tetracycline-inducible promoters (Tet-On or Tet-Off), small molecule two-hybrid transcription activation systems (FKBP, ABA, etc.), or light-inducible systems (phytochrome, LOV domain, or cryptochrome), such as light-inducible transcription effectors (LITEs) that induce sequence-specific changes in transcriptional activity. Components of light-inducible systems can include RNA-targeting CRISPR enzymes, light-responsive cytochrome heterodimers (e.g., from Arabidopsis thaliana), and transcription activation / repression domains. Further examples of inducible DNA-binding proteins and methods of use thereof are provided in U.S. Provisional Patent Application Nos. 61 / 736465 and 61 / 721,283 (hereby incorporated by reference in their entirety).
[0241] In particular embodiments, transient or inducible expression can be achieved using, for example, chemically regulated promoters, whereby gene expression is induced upon addition of an exogenous chemical. Regulation of gene expression can also be achieved through chemically repressible promoters, whereby gene expression is repressed upon addition of a chemical. Chemically inducible promoters include, but are not limited to, the maize ln2-2 promoter, which is activated by benzenesulfonamide herbicide antidotes (De Veylder et al., (1997) Plant Cell Physiol 38:568-77), the maize GST promoter (GST-II-27, WO 93 / 01294), which is activated by hydrophobic electrophilic compounds used as pre-emergence herbicides, and the tobacco PR-1a promoter, which is activated by salicylic acid (Ono et al., (2004) Biosci Biotechnol Biochem 68:803-7). Antibiotic-regulated promoters, such as tetracycline-inducible and tetracycline-repressible promoters (Gatz et al., (1991) Mol Gen Genet 227:229-37; U.S. Pat. Nos. 5,814,618 and 5,789,156), can also be used herein.
[0242] Transport to and / or expression in specific plant organelles The expression system may contain elements for translocation to and / or expression in specific plant organelles.
[0243] Chloroplast targeting In particular embodiments, it is envisioned that RNA targeting CRISPR system is used to specifically modify the expression and / or translation of chloroplast gene, or ensure the expression in chloroplast.For this purpose, chloroplast transformation method or RNA targeting CRISPR components are used to compartmentalize into chloroplast.For example, introducing genetic modification into plastid genome can reduce biosafety issues such as gene flow through pollen.
[0244] Chloroplast transformation methods are known in the art and include particle bombardment, PEG treatment, and microinjection. In addition, methods involving the transfer of transformation cassettes from the nuclear genome to plastids can be used, as described in WO2010061186.
[0245] Alternatively, it is envisioned that one or more of the RNA-targeting CRISPR components may be targeted to plant chloroplasts. This is achieved by incorporating into the expression construct a sequence encoding a chloroplast transit peptide (CTP) or plastid transit peptide operably linked to the 5' region of the sequence encoding the RNA-targeting protein. The CTP is removed in a processing step during translocation to the chloroplasts. Chloroplast targeting of expressed proteins is well known to those skilled in the art (see, for example, Protein Transport into Chloroplasts, 2010, Annual Review of Plant Biology, Vol. 61:157-180). In such embodiments, it may also be desirable to target one or more guide RNAs to plant chloroplasts. Methods and constructs that can be used to translocate guide RNAs to chloroplasts using chloroplast-localizing sequences are described, for example, in U.S. Patent Application Publication No. 20040142476 (incorporated herein by reference). By incorporating such various constructs into the expression system of the present invention, one or more RNA-targeting guide RNAs can be efficiently transferred.
[0246] Introduction of polynucleotides encoding the CRISPR-RNA targeting system into algal cells Transgenic algae (or other plants, such as rapeseed) may be particularly useful for the production of vegetable oils or biofuels, such as alcohols (especially methanol and ethanol), or other products. They can be engineered to express or overexpress high levels of oils or alcohols for use in the oil or biofuel industry.
[0247] U.S. Patent No. 8,945,839 describes a method for engineering microalgae (Chlamydomonas reinhardtii cells) using Cas9. Using similar tools, the RNA-targeting CRISPR-based method described herein can be applied to Chlamydomonas species and other algae. In a detailed embodiment, an RNA-targeting protein and one or more guide RNAs are introduced and expressed in algae using a vector expressing the RNA-targeting protein under the control of a constitutive promoter, such as Hsp70A-Rbc S2 or β2-tubulin. The guide RNA is optionally delivered using a vector containing a T7 promoter. Alternatively, RNA-targeting mRNA and in vitro transcribed guide RNA can be delivered to algae cells. Electroporation protocols are available to those skilled in the art, such as the standard recommended protocol from the GeneArt Chlamydomonas Engineering Kit.
[0248] Introduction of polynucleotides encoding RNA targeting components into yeast cells In a specific embodiment, the present invention relates to the use of RNA targeting CRISPR system for RNA editing in yeast cells. The transformation method of yeast cells that can be used to introduce polynucleotides encoding RNA targeting CRISPR system components is well known to those skilled in the art and is reviewed by Kawai et al., 2010, Bioeng Bugs. 2010 Nov-Dec; 1 (6): 395-403). Non-limiting examples include transformation of yeast cells by lithium acetate treatment (which may further include carrier DNA and PEG treatment), bombardment or electroporation.
[0249] Transient expression of RNA-targeting CRISP system components in plants and plant cells. In particular embodiments, it is envisioned that the guide RNA and / or RNA targeting gene is transiently expressed in plant cells. In these embodiments, the RNA targeting CRISPR system will only modify the RNA target molecule when both the guide RNA and the RNA targeting protein are present in the cell, thus further controlling gene expression. Because the expression of the RNA targeting enzyme is transient, plants regenerated from such plant cells typically do not contain foreign DNA. In particular embodiments, the RNA targeting enzyme is stably expressed by the plant cell, and the guide sequence is transiently expressed.
[0250] In a particularly preferred embodiment, RNA-targeting CRISPR system components can be introduced into plant cells using plant viral vectors (Scholthof et al. 1996, Annu Rev Phytopathol. 1996;34:299-323). In further detailed embodiments, the viral vector is a vector derived from a DNA virus. For example, a geminivirus (e.g., cabbage leaf curl virus, bean yellows virus, wheat dwarf virus, tomato leaf curl virus, corn streak virus, tobacco leaf curl virus, or tomato golden mosaic virus) or a nanovirus (e.g., broad bean spotted wilt virus). In other detailed embodiments, the viral vector is a vector derived from an RNA virus. For example, a tobravirus (e.g., tobacco rattle virus, tobacco mosaic virus), a potexvirus (e.g., potato virus X), or a hordeivirus (e.g., wheat stripe mosaic virus). Replicating genomes of plant viruses are non-integrating vectors, which is advantageous in terms of avoiding the creation of GMO plants.
[0251] In a detailed embodiment, the vector used for transient expression of RNA targeting CRISPR constructs is, for example, a pEAQ vector, which has been adapted for Agrobacterium-mediated transient expressi...
Claims
1. 1. A non-naturally occurring or engineered composition for modifying a target RNA sequence in a mammalian cell, the composition comprising: (a) a catalytically inactive Cas13a that exhibits reduced collateral RNase activity in a mammalian cell; (b) a guide molecule that forms a complex with the Cas13a and is capable of directing sequence-specific binding of the complex to a target sequence of a target RNA sequence; and (c) one or more heterologous functional domains associated with the Cas13a, wherein the one or more functional domains convert adenosine to inosine or cytosine to uracil, and the Cas13a is codon-optimized for expression in a mammalian cell.
2. 2. The composition of claim 1, wherein the Cas13a is Leptotrichia wadei F0279 (Lw2).
3. The Cas13a is effective against Blautia sp. Marseille-P2398, Chloroflexus aggregans, Demequina aurantiaca, Thalassospira sp. TSL5-1, Pseudobutyrivibrio sp. OR37, Butyrivibrio sp. YAB3001, Leptotrichia sp. Marseille-P3007, Bacteroides ihuae, Porphyromonadaceae bacterium KH3CP3RA, Listeria riparia, and Insolitispirillum peregrinum; 2. The composition of claim 1, wherein the Cas13a comprises a mutation in the HEPN domain corresponding to R597A, H602A, R1278A and / or H1283A of Leptotrichia shahii, or the corresponding amino acids in a Cas13a ortholog.
4. 2. The composition of claim 1, wherein the Cas13a comprises at least two higher eukaryotic and prokaryotic nucleotide-binding (HEPN) domains.
5. The composition of claim 1, wherein the Cas13a comprises one or more localization signals, and at least one localization signal is a nuclear localization signal (NLS) or a nuclear export signal (NES).
6. 2. The composition of claim 1, wherein the one or more heterologous functional domains are an ADAR enzyme, an APOBEC enzyme, or an AID enzyme.
7. 10. The composition of claim 1, wherein the guide molecule comprises a dual direct repeat sequence.
8. 10. The composition of Claim 1, further comprising one or more polynucleotide molecules encoding Cas13a and / or a guide molecule, wherein said one or more polynucleotide molecules comprise one or more regulatory elements operably configured to express a polypeptide and / or nucleic acid component.
9. The composition of claim 8 , wherein the one or more regulatory elements comprise one or more promoters or inducible promoters.
10. The composition of claim 8 , wherein the one or more polynucleotide molecules are contained within one or more vectors.
11. The composition of claim 10 , wherein the one or more vectors are viral vectors.
12. 12. The composition of claim 11, wherein the one or more viral vectors comprise one or more retroviral, lentiviral, adenoviral, adeno-associated, or herpes simplex viral vectors.
13. The composition of claim 1 , wherein the composition is contained in a delivery system.
14. 10. The composition of claim 1, wherein the non-naturally occurring or engineered composition is delivered by a delivery vehicle comprising one or more liposomes, one or more particles, one or more exosomes, one or more microvesicles, a gene gun, or one or more viral vectors.
15. 15. A eukaryotic cell engineered to contain or express the composition of any one of claims 1 to 14.
16. The cell of claim 15 , wherein the cell comprises a mammalian cell.
17. A multicellular organism comprising one or more cells according to claim 15.
18. 16. A plant or animal model comprising one or more cells according to claim 15, wherein said one or more cells express said composition or a component thereof.
19. 15. A method of modifying a target sequence comprising delivering a composition of any one of claims 1 to 14 to a locus of interest containing the mammalian target sequence.
20. 20. The method of claim 19, wherein modifying the target sequence comprises converting adenosine to inosine or cytosine to uracil without strand cleavage.
21. 20. The method of claim 19, wherein the Cas13a CRISPR-Cas effector protein is associated with one or more functional domains.
22. 22. The method of claim 21, wherein the one or more functional domains alter transcription or translation of the target locus.
23. 20. The method of claim 19, wherein the Cas13a comprises one or more localization signals, and at least one of the one or more localization signals is a nuclear localization signal (NLS) or a nuclear export signal (NES).
24. 20. The method of Claim 19, further comprising one or more polynucleotide molecules encoding Cas13a and / or a guide molecule, wherein said one or more polynucleotide molecules comprise one or more regulatory elements operably configured to express the polypeptide and / or nucleic acid components.
25. 25. The method of claim 24, wherein the one or more regulatory elements comprise one or more promoters or inducible promoters.
26. 15. A method of treatment comprising administering to a non-human mammalian subject in need thereof a composition according to any one of claims 1 to 14.