Crystal Structure of CRISPR-CPF1
The Cpf1 effector protein complex with engineered nucleic acid components addresses the need for scalable and precise genome editing by inducing targeted modifications in eukaryotic genomes, enhancing gene expression control and reducing off-target effects.
Patent Information
- Application Number
- JP2023079002
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-03-31
- Filing Date
- 2023-05-12
- Publication Date
- 2025-07-16
- Estimated Expiration
- 2037-01-23
Smart Images

Figure 0007709487000422 
Figure 0007709487000423 
Figure 0007709487000424
Abstract
Description
Technical Field
[0001] Related Applications and Incorporation by Reference This application claims the benefit and priority of U.S. Provisional Patent Application No. 62 / 281,947, filed on January 22, 2016, and U.S. Provisional Patent Application No. 62 / 316,240, filed on March 31, 2016.
[0002] All documents cited therein or cited during its examination procedure ("application cited documents") and all documents cited or referenced in the documents cited in this specification are hereby incorporated by reference into this specification and may be used in the practice of the invention, together with any manufacturer's instructions, specifications, product specifications, and product sheets for any product related to any of the documents mentioned in this specification or incorporated herein by reference. More specifically, all documents referenced are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference.
[0003] Description of Federally Sponsored Research This invention was made with government support under Grant Nos. MH100706, MH110049, and DK097768 awarded by the National Institutes of Health. The government has certain rights in this invention.
[0004] This invention was made with the support of PRESTO (Precursory Research for Embryonic Science and Technology: Young Individual Research Promotion Program) 15H01463, granted by JST (Japan Science and Technology Agency). JST has certain rights in this invention. This research was supported by JSPS KAKENHI Grant Number 26291010.
[0005] The present invention generally relates to systems, methods, and compositions for use in the control of gene expression, including sequence targeting that can utilize vector systems related to clustered regularly interspaced short palindromic repeats (CRISPR) and its components, such as perturbation of gene transcripts or nucleic acid editing.
Background Art
[0006] With recent advances in genome sequencing techniques and analysis methods, the ability to catalog and map genetic factors associated with various biological functions and diseases has been rapidly improving. Accurate genome targeting technologies are needed to enable systematic reverse engineering of causative genetic mutations by selectively perturbing individual genetic elements and to advance synthetic biology, biotechnology, and medical applications. Although genome editing techniques such as designer zinc fingers, transcription activator-like effectors (TALEs), or homing meganucleases are available to cause targeted genome perturbation, there is still a need for new genome engineering technologies that utilize novel strategies and molecular mechanisms, are inexpensive, easy to set up, scalable, and suitable for targeting multiple positions within eukaryotic genomes. It will be a major resource for new applications in genome engineering and biotechnology.
[0007] The CRISPR-Cas systems of bacterial and archaeal adaptive immunity exhibit extremely high diversity in terms of protein composition and genomic locus organization. Since the CRISPR-Cas locus has more than 50 gene families and no strictly universal genes, rapid evolution and extremely high diversity of locus organization are suggested. So far, by taking a multi-directional approach, approximately 395 profiles of cas genes for 93 Cas proteins have been comprehensively identified. Included in the classification are signature gene profiles + signatures of locus organization. A new classification of CRISPR-Cas systems has been proposed, where these systems are roughly divided into two classes: class 1, which has a multi-subunit effector complex, and class 2, which has a single-subunit effector module exemplified by the Cas9 protein. Novel effector proteins associated with class 2 CRISPR-Cas systems can be developed as powerful genome engineering tools, and the prediction of putative novel effector proteins as well as their engineering and optimization are important.
[0008] The citation or identification of any document in this application does not admit that such document is available as prior art for the present invention.
Summary of the Invention
Problems to be Solved by the Invention
[0009] Alternative and robust systems and techniques for targeting nucleic acids or polynucleotides (e.g., DNA or RNA or any hybrids or derivatives thereof) with broad applicability are urgently needed. The present invention addresses this need and provides related advantages. The novel DNA or RNA targeting systems of the present application, when added to the repertoire of genomic and epigenomic targeting technologies, can change the study and perturbation or editing of specific target sites through direct detection, analysis, and manipulation. To effectively utilize the DNA or RNA targeting systems of the present application for genomic or epigenomic targeting without harmful effects, it is critically important to understand the aspects of engineering and optimization of these DNA or RNA targeting tools.
Means for Solving the Problems
[0010] The present invention provides a method for modifying a sequence associated with or at a target locus of interest, the method comprising delivering to the locus a non-naturally occurring or engineered composition comprising a Cpf1 effector protein and one or more nucleic acid components, wherein the effector protein forms a complex with the one or more nucleic acid components, and when the complex binds to the target locus of interest, the effector protein induces a modification of the sequence associated with or at the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break.
[0011] The terms Cas enzyme, CRISPR enzyme, CRISPR protein, Cas protein, and CRISPR Cas are generally used synonymously and will be understood to refer to the novel CRISPR effector proteins further described in the present application by analogy when referred to herein, unless specifically indicated otherwise, such as by specific reference to Cas9. The CRISPR effector proteins described herein are preferably Cpf1 effector proteins.
[0012] The present invention provides a method of modifying a sequence associated with or at a target locus of interest, the method comprising delivering a non-naturally occurring or engineered composition comprising a Cpf1 locus effector protein and one or more nucleic acid components to the sequence associated with or at the locus, wherein the Cpf1 effector protein forms a complex with the one or more nucleic acid components, and when the complex binds to the target locus of interest, the effector protein induces modification of a sequence associated with or at the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break. In a preferred embodiment, the Cpf1 effector protein forms a complex with one nucleic acid component; advantageously an engineered or non-naturally occurring nucleic acid component. Induction of modification of a sequence associated with or at a target locus of interest can be under a Cpf1 effector protein-nucleic acid guide. In a preferred embodiment, the one nucleic acid component is a CRISPR RNA (crRNA). In a preferred embodiment, the one nucleic acid component is a mature crRNA or guide RNA, wherein the mature crRNA or guide RNA comprises a spacer sequence (or guide sequence) and a direct repeat sequence or derivatives thereof. In a preferred embodiment, the spacer sequence or a derivative thereof comprises a seed sequence, wherein the seed sequence is critically important for recognition and / or hybridization with a sequence at the target locus. In a preferred embodiment, the seed sequence of the Cpf1 guide RNA is within the first approximately 5 nt on the 5' end of the spacer sequence (or guide sequence). In a preferred embodiment, the strand break is a sticky end type break with a 5' overhang. In a preferred embodiment, the sequence associated with or at the target locus of interest comprises linear DNA or supercoiled DNA.
[0013] Aspects of the invention relate to non-naturally occurring or engineered compositions comprising a Cpf1 locus effector protein and one or more nucleic acid components, where the Cpf1 effector protein is capable of forming a complex with one or more nucleic acid components, preferably an engineered or non-naturally occurring nucleic acid component. In a preferred embodiment, one nucleic acid component is a mature crRNA or guide RNA, where the mature crRNA or guide RNA comprises a spacer sequence (or guide sequence) and a direct repeat sequence or derivatives thereof. In a preferred embodiment, the spacer sequence or a derivative thereof comprises a seed sequence, where the seed sequence is capable of hybridizing to a sequence within a target DNA. In a particular embodiment, the DNA molecule is a DNA molecule encoding a gene product in a cell. Hybridization of the guide RNA to the target sequence targets the complex to the target DNA, ensuring modification of the target sequence.
[0014] In a preferred embodiment, the modification is the introduction of a strand break. In a preferred embodiment, the Cpf1 effector protein forms a complex with one nucleic acid component;
[0015] Induction of modification of a sequence associated with or at a target locus of interest can be under a Cpf1 effector protein-nucleic acid guide. In a preferred embodiment, said one nucleic acid component is a CRISPR RNA (crRNA). Aspects of the invention relate to a Cpf1 effector protein complex having one or more non-naturally occurring or engineered or modified or optimized nucleic acid components. In a preferred embodiment, the nucleic acid component of this complex can comprise a guide sequence linked to a direct repeat sequence, where the direct repeat sequence comprises one or more stem-loops or an optimized secondary structure. In a preferred embodiment, the direct repeat has a minimum length of 16 nt and a single stem-loop. In a further embodiment, the direct repeat is longer than 16 nt, preferably longer than 17 nt, and has two or more stem-loops or an optimized secondary structure. In a preferred embodiment, the direct repeat may be modified to include one or more protein-binding RNA aptamers. In a preferred embodiment, one or more aptamers may be included as part of an optimized secondary structure. Such aptamers may have the ability to bind to a bacteriophage coat protein. The bacteriophage coat protein may be selected from the group consisting of Qβ, F2, GA, fr, JP501, MS2, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s and PRR1. In a preferred embodiment, the bacteriophage coat protein is MS2. The invention also provides nucleic acid components of the complex that are 30 or more, 40 or more or 50 or more nucleotides in length.
[0016] The present invention provides a genome editing method, which method includes two or more rounds of Cpf1 effector protein targeting and cleavage. In certain embodiments, the first round includes cleaving a sequence associated with a target locus that is distal from the seed sequence by a Cpf1 effector protein, and the second round includes cleaving a sequence at the target locus by a Cpf1 effector protein. In a preferred embodiment of the present invention, the first targeting round by the Cpf1 effector protein results in indels, and the second targeting round by the Cpf1 effector protein can be repaired by homology-directed repair (HDR). In the most preferred embodiment of the present invention, one or more targeting rounds by the Cpf1 effector protein result in sticky-end cleavage that can be repaired by the insertion of a repair template.
[0017] The present invention provides a method for genome editing or modification of a sequence associated with or at a target locus of interest, which method includes introducing a Cpf1 effector protein complex into any desired cell type, prokaryotic or eukaryotic cell, whereby the Cpf1 effector protein complex functions effectively to incorporate a DNA insert into the genome of the eukaryotic or prokaryotic cell. In a preferred embodiment, the cell is a eukaryotic cell and the genome is a mammalian genome. In a preferred embodiment, the incorporation of the DNA insert is facilitated by a non-homologous end joining (NHEJ)-based gene insertion mechanism. In a preferred embodiment, the DNA insert is an exogenously introduced DNA template or repair template. In a preferred embodiment, the exogenously introduced DNA template or repair template is delivered together with a polynucleotide vector for the expression of the Cpf1 effector protein complex or one of the components or components of the complex. In a more preferred embodiment, the eukaryotic cell is a non-dividing cell (e.g., a non-dividing cell in which genome editing by HDR is particularly problematic). In a preferred genome editing method in human cells, the Cpf1 effector protein can include, but is not limited to, FnCpf1, AsCpf1, and LbCpf1 effector proteins.
[0018] The present invention also provides a method for modifying a target locus of interest, the method comprising delivering to the locus a non-naturally occurring or engineered composition comprising a Cpf1 effector protein and one or more nucleic acid components, wherein the Cpf1 effector protein forms a complex with the one or more nucleic acid components, and when the complex binds to the target locus of interest, the effector protein induces modification of the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break.
[0019] In such a method, the target locus of interest can be contained in a DNA molecule in vitro. In a preferred embodiment, the DNA molecule is a plasmid. In such a method, the target locus of interest can be contained in a DNA molecule within a cell. The cell can be a prokaryotic cell or a eukaryotic cell. The cell can be a mammalian cell.
[0020] The mammalian cells may be non-human mammalian cells, such as primate, bovine, ovine, porcine, canine, rodent, Leporidae (e.g., monkey, female bovine, ovine, porcine, canine, rabbit, rat or mouse cells). The cells may also be non-mammalian eukaryotic cells, such as avian birds (e.g., chicken), vertebrate fish (e.g., salmon) or crustacean (e.g., oyster, clam, lobster, shrimp) cells. The cells may also be plant cells. The plant cells may be monocotyledonous or dicotyledonous or of crops or cereal plants, such as cassava, corn, sorghum, soybean, wheat, oat or rice. The plant cells may also be of algae, trees or production plants, fruits or vegetables (e.g., citrus trees, such as orange, grapefruit or lemon trees; peach or nectarine trees; apple or pear trees; nut trees, such as almond, walnut or pistachio trees; solanaceous plants; Brassica plants; Lactuca plants; Spinacia plants; Capsicum plants; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.).
[0021] The modifications introduced into the cells according to the present invention can be such as to change the cells and the progeny of the cells for the improvement of the production of biological products, such as antibodies, starches, alcohols or other desired cell products. The modifications introduced into the cells according to the present invention can be such that the changes that alter the biological products produced are included in the cells and the progeny of the cells.
[0022] In any of the methods described, the target locus of interest may be a genomic or epigenomic locus of interest. In any of the methods described, the complex may be delivered together with a plurality of guides for multiplexed use. In any of the methods described, two or more proteins may be used.
[0023] In a preferred embodiment of the present invention, biochemical or in vitro or in vivo cleavage of a sequence associated with or at a target locus of interest, for example cleavage by an AsCpf1 effector protein, occurs without a putative trans-activating crRNA (tracrRNA) sequence. In other embodiments of the present invention, cleavage, for example cleavage by other CRISPR family effector proteins, may occur with a putative trans-activating crRNA (tracrRNA) sequence. However, it has been found that tracrRNA is not required for target DNA cleavage by the Cpf1 effector protein complex, and more specifically, that a Cpf1 effector protein complex containing only the Cpf1 effector protein and crRNA (guide RNA containing direct repeat sequences and a guide sequence) was sufficient for cleavage of target DNA (Zetsche et al, 2015, Cell 163, 759-771).
[0024] In any of the methods described, the effector protein (e.g., Cpf1) and nucleic acid component(s) may be provided by that protein and / or one or more polynucleotide molecules encoding one or more nucleic acid components, and where the one or more polynucleotide molecules are operably configured to express the protein and / or one or more nucleic acid components. The one or more polynucleotide molecules may include one or more regulatory elements operably configured to express the protein and / or one or more nucleic acid components. The one or more polynucleotide molecules may be contained within one or more vectors. The present invention encompasses such one or more polynucleotide molecules, such as polynucleotide molecules operably configured to express the protein and / or one or more nucleic acid components, and such one or more vectors.
[0025] In any of the methods described, the strand cleavage may be a single-strand cleavage or a double-strand cleavage.
[0026] The regulatory element may comprise an inducible promoter. The polynucleotide and / or vector system may comprise an inducible system.
[0027] In any of the methods described, one or more polynucleotide molecules may be included in the delivery system, or one or more vectors may be included in the delivery system.
[0028] In any of the methods described, a non-naturally occurring or engineered composition may be delivered by liposomes, particles (e.g., nanoparticles), exosomes, microvesicles, gene guns, or one or more vectors, such as nucleic acid molecules or viral vectors.
[0029] The present invention also provides a non-naturally occurring or engineered composition that has the features as discussed herein or is a composition defined by any of the methods described herein.
[0030] The present invention also provides a vector system comprising one or more vectors, wherein the one or more vectors comprise one or more polynucleotide molecules encoding a component of a non-naturally occurring or engineered composition that has the features as discussed herein or is a composition defined by any of the methods described herein.
[0031] The present invention also provides a delivery system comprising one or more vectors or one or more polynucleotide molecules, wherein the one or more vectors or polynucleotide molecules comprise one or more polynucleotide molecules encoding a component of a non-naturally occurring or engineered composition that has the features as discussed herein or is a composition defined by any of the methods described herein.
[0032] The present invention also provides a composition that does not occur naturally or has been engineered, or one or more polynucleotides encoding a component of said composition, or a vector or delivery system comprising one or more polynucleotides encoding a component of said composition, for use in a therapeutic treatment method. The therapeutic treatment method may include gene or genome editing, or gene therapy.
[0033] The present invention also provides methods and compositions in which one or more amino acid residues of an effector protein may be modified, e.g., an engineered or non-naturally occurring effector protein or Cpf1. In certain embodiments, the modification may include a mutation of one or more amino acid residues of the effector protein. The one or more mutations may be in one or more catalytically active domains of the effector protein. The effector protein may have reduced or ablated nuclease activity as compared to an effector protein lacking the one or more mutations. The effector protein may not induce cleavage of either DNA or RNA strand at a target locus of interest. The effector protein may not induce cleavage of any DNA or RNA strand at a target locus of interest. In preferred embodiments, the one or more mutations may include two mutations. In preferred embodiments, one or more amino acid residues are modified in a Cpf1 effector protein, e.g., an engineered or non-naturally occurring effector protein or Cpf1. In preferred embodiments, the Cpf1 effector protein is an AsCpf1 effector protein. In preferred embodiments, the one or more modified or mutated amino acid residues are D908, E993, D1263 based on the amino acid numbering of the AsCpf1 effector protein. In a more preferred embodiment, the one or more mutated amino acid residues are D908A, E993A, D1263A based on the amino acid positions in AsCpf1.
[0034] In a preferred embodiment, the one or more modified or mutated amino acid residues are selected from any one of the amino acids at D861, R862, R863, W382, E993, D1263, D908, W958, K968, R951, R1226, S1228, D1235, K548, M604, K607, T167, N631, N630, K547, K163, Q571, K1017, R955, K1009, R909, R912, R1072, E372, K15, K810, H755, K557, E857, K943, K1022, K1029, K942, K949, R84, K87, K200, H206, R210, R301, R699, K705, K887, R891, K1086, K1089, R1094, R1127, R1220, Q1224, N178, N197, N204, N259, N278, N282, N519, N747, N759, N878, N889, and / or in the regions of 1189 - 1197, 1200 - 1208, 398 - 400, 380 - 383, 362 - 420, 1163 - 1173, 1230 - 1233, 1152 - 1148, 1076 - 1249, based on the amino acid numbering of Acidaminococcus sp. BV3L6. In a preferred embodiment, the one or more modified or mutated amino acid residues are selected from the list consisting of R862A, E993A, D1263A, D908A, W958A, R951A, R1226A, S1228A, D1235A, K548A, M604A, K607A, K607R, T167S, N631K, N613R, N630K, N630R, K547R, K163R, Q571K, Q571R, K1009A, R909A, R1072A, E327A, K15A, K810A, H755A, K557A, E857A, K943A, K1022A, K1029A, K942A, K949A, R84A, K87A, K200A, H206A, R210A, R301A, R699A, K705A, K887A, R891A, K1086A, K1089A, R1094A, R1127A, R1220A and Q1224A.In a preferred embodiment, the one or more modified or mutated amino acid residues are selected from the list consisting of R862A, E993A, D1263A, D908A, W958A, R951A, K548A, M604A, K607A, K607R, N631K, N613R, N630K, N630R, K547R, K163R, Q571K, Q571R, K1009A, R909A, R1072A, E327A, K15A, K810A, H755A, K557A, E857A, K943A, K1022A, K1029A, K942A, K949A, R84A, K87A, K200A, H206A, R210A, R301A, R699A, K705A, K887A, R891A, K1086A, K1089A, R1094A, R1127A, R1220A and Q1224A; In a preferred embodiment, the one or more modified or mutated amino acid residues are selected from N178, N197, N204, N259, N278, N282, N519, N747, N759, N878, N889. In a preferred embodiment, the one or more modified or mutated amino acid residues are selected from the list consisting of R862A, W958A, R951A, R1226A, S1228A, D1235A, K548A, M604A, K607A, K607R, T167S, N631K, N613R, N630K, N630R, K547R, K163R, Q571K, Q571R, K1009A, R909A, R1072A, E327A, K15A, K810A, H755A, K557A, E857A, K943A, K1022A, K1029A, K942A, K949A, R84A, K87A, K200A, H206A, R210A, R301A, R699A, K705A, K887A, R891A, K1086A, K1089A, R1094A, R1127A, R1220A and Q1224A.In a preferred embodiment, the one or more modified or mutated amino acid residues are selected from D861, W958, S1228, D1235, T167, N631, N630, K547, K163, Q571, R1226, E372, K15, K810, H755, K557, E857, K943, K1022, K1029, K942, K949, R84, K87, K200, H206, R210, R301, R699, K705, K887, R891, K1086, K1089, R1094, R1127, R1220, Q1224, N178, N197, N204, N259, N278, N282, N519, N747, D749, N759, H761, H872, N878, N889, and / or any one amino acid in the regions of 1189 - 1197, 1200 - 1208, 398 - 400, 380 - 383, 362 - 420, 1163 - 1173, 1230 - 1233, 1152 - 1148, 1076 - 1249. In a particular embodiment, the mutation is R862A and the Cpf1 enzyme no longer binds to RNA. In a particular embodiment, the one or more mutations are selected from K15A, D749A, H761A, H872A, K810A, H755A, K557A, E857A, R862A, K943A, K1022A, and K1029A, where the Cpf1 enzyme is no longer capable of RNA binding and / or processing. In a particular embodiment, the one or more mutations are selected from K547A, K607A, M604A, and T176S, where the TTT specificity is reduced or removed. In a particular embodiment, the one or more mutations are selected from N631K, N613R, N630K, N630R, K547R, K163R, Q571K, Q571R, and K607R, where the non-specific DNA interaction of the Cpf1 enzyme is increased. In a particular embodiment, the one or more mutations are selected from R84A, K87A, K200A, H206A, R210A, R301A, R699A, K705A, K887A, R891A, K1086A, K1089A, R1094A, R1127A, R1220A, and Q1224A, whereby the specificity of the enzyme is increased or reduced.In certain embodiments, one or more of D861, R862, R863, and W382 are mutated and the RNA binding of said Cpf1 is disrupted. In certain embodiments, one or more of amino acids W958, K968, R951, R1226, D1253, and T167 and the stability of Cpf1 are affected. In certain embodiments, one or more of K968 and R951 are mutated and the DNA binding of said Cpf1 is disrupted. In certain embodiments, one or more of N631 and N630 are mutated and the interaction with phosphate in the DNA backbone is increased. In certain embodiments, one or more of the following amino acids: based on the amino acid numbering of AsCpf1 (Acidaminococcus sp. BV3L6), L117, T118, D119, T150, T151, T152, R341, N342, E343, T398, G399, K400, D451, Q452, P453, L454, P455, T456, T457, L458, K459, V486, D487, E488, S489, N490, E491, V492, D493, P494, E506, M507, E508, Q571, K572, G573, R574, Y575, T621, E649, K650, E651, D665, T737, D749, F750, K815, N848, V1108, K1109, T1110, G1111, S1124, A1195, A1196, A1197, N1198, L1244, N1245, and / or G1246 are mutated, whereby the stability and / or activity of the Cpf1 enzyme is not substantially affected.
[0035] The present invention also provides one or more mutations or two or more mutations that should be in the catalytically active domain of an effector protein containing an RuvC domain. In some embodiments of the present invention, the RuvC domain may include a catalytically active domain that is homologous to the RuvCI, RuvCII, or RuvCIII domain, or any relevant domain homologous to the RuvCI, RuvCII, or RuvCIII domain or as described in any of the methods described herein. The effector protein may include one or more heterologous functional domains. The one or more heterologous functional domains may include one or more nuclear localization signal (NLS) domains. The one or more heterologous functional domains may include at least two or more NLS domains. The one or more NLS domains may be located at or near or in proximity to the terminus of the effector protein (e.g., Cpf1), and in the case where there are two or more NLSs, each of the two may be located at or near or in proximity to the terminus of the effector protein (e.g., Cpf1). The one or more heterologous functional domains may include one or more transcriptional activation domains. In a preferred embodiment, the transcriptional activation domain may include VP64. The one or more heterologous functional domains may include one or more transcriptional repression domains. In a preferred embodiment, the transcriptional repression domain includes a KRAB domain or a SID domain (e.g., SID4X). The one or more heterologous functional domains may include one or more nuclease domains. In a preferred embodiment, the nuclease domain includes Fok1.
[0036] The present invention also provides one or more heterologous functional domains so as to have one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional derepression factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity. At least one or more heterologous functional domains may be at or near the amino terminus of the effector protein, and / or here at least one or more heterologous functional domains may be at or near the carboxy terminus of the effector protein. One or more heterologous functional domains may be fused to the effector protein. One or more heterologous functional domains may be tethered to the effector protein. One or more heterologous functional domains may be linked to the effector protein via a linker moiety.
[0037] The present invention also provides an effector protein (e.g., Cpf1) containing an effector protein (e.g., Cpf1) derived from an organism of a genus including Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus, Eubacterium, Corynebacter, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Clostridiaridium, Leptotrichia, Francisella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus, Letospira, Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium or Acidaminococcus.
[0038] The present invention also provides effector proteins (such as Cpf1) comprising effector proteins (such as Cpf1) derived from organisms of S. mutans, S. agalactiae, S. equisimilis, S. sanguinis, Streptococcus pneumoniae; C. jejuni, C. coli; N. salsuginis, N. tergarcus; S. auricularis, S. carnosus; N. meningitides, Neisseria gonorrhoeae; Listeria monocytogenes, L. ivanovii; Clostridium botulinum, C. difficile, Clostridium tetani, C. sordellii.
[0039] The effector protein can comprise a chimeric effector protein comprising a first fragment from a first effector protein (e.g., Cpf1) ortholog and a second fragment from a second effector (e.g., Cpf1) protein ortholog, wherein the first and second effector protein orthologs are different. At least one of the first and second effector protein (e.g., Cpf1) orthologs is from the genus Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus, Eubacterium, Corynebacter, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Clostridiaridium, Leptotrichia, Francisella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus, Letospira, Desulfovibrio, Desulfonatronum, Opitutaceae,It can contain effector proteins (such as Cpf1) derived from organisms including the genus Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium, or Acidaminococcus; for example, a chimeric effector protein containing a first fragment and a second fragment, where each of the first and second fragments is from the genus Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus, Eubacterium, Corynebacterium, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Clostridiaridium, Leptotrichia, Francisella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, the phylum Bacteroidetes, Helcococcus, Letospira, Desulfovibrio, Desulfonatronum, Opitutaceae,Selected from Cpf1 of an organism comprising a genus Tuberibacillus, a genus Bacillus, a genus Brevibacilus, a genus Methylobacterium or a genus Acidaminococcus, wherein the first and second fragments are not from the same bacterium; for example, a chimeric effector protein comprising a first fragment and a second fragment, wherein each of the first and second fragments is from S. mutans, S. agalactiae, S. equisimilis, S. sanguinis, Streptococcus pneumoniae; C. jejuni, C. coli; N. salsuginis, N. tergarcus; S. auricularis, S. carnosus; N. meningitides, Neisseria gonorrhoeae; Listeria monocytogenes, L. ivanovii; Clostridium botulinum, C. difficile, Clostridium tetani, C. sordellii; Francisella tularensis 1, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020,Selected from Cpf1 of Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, and Porphyromonas macacae, wherein the first and second fragments are not from the same bacterium.,
[0040] In a preferred embodiment of the present invention, the effector protein is derived from the Cpf1 locus (hereinafter, such effector protein is also referred to as "Cpf1p") and is, for example, a Cpf1 protein (and such effector protein or Cpf1 protein or protein derived from the Cpf1 locus is also referred to as "CRISPR enzyme"). The Cpf1 locus includes, but is not limited to, the Cpf1 loci of the bacterial species listed in FIG. 64. In a more preferred embodiment, Cpf1p is derived from a bacterial species selected from Francisella tularensis 1, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, and Porphyromonas macacae.In certain embodiments, Cpf1p is derived from a bacterial species selected from the genus Acidaminococcus sp. BV3L6.
[0041] In a further embodiment of the invention, a protospacer adjacent motif (PAM) or PAM-like motif directs the binding of the effector protein complex to the target locus of interest. In a preferred embodiment of the invention, the PAM is 5’ NTTT [wherein N is A / C or G], and the effector protein is AsCpf1p. In another preferred embodiment of the invention, the PAM is 5’ TTTV [wherein V is A / C or G], and the effector protein is PaCpf1p. In certain embodiments, the PAM is 5’ TTN [wherein N is A / C / G or T], the effector protein is FnCpf1p, and the PAM is located upstream of the 5’ end of the protospacer. In a particular embodiment of the invention, the PAM is 5’ CTA, where the effector protein is FnCpf1p, and the PAM is located upstream of the 5’ end of the protospacer or target locus. In a preferred embodiment, the invention provides an expansion of the targeting range of an RNA-guided genome editing nuclease that enables targeting and editing of AT-rich genomes by T-rich PAMs of the Cpf1 family. In certain embodiments, the CRISPR enzyme can be engineered to contain one or more mutations that reduce or abolish nuclease activity.
[0042] Amino acid positions in the AsCpf1p RuvC domain include, but are not limited to, 908, 993, and 1263. In a preferred embodiment, the mutations in the AsCpf1p RuvC domain are D908A, E993A, and D1263A, where the D908A, E993A, and D1263A mutations completely inactivate the DNA cleavage activity of the AsCpf1 effector protein.
[0043] Mutations can also be created in adjacent residues, such as amino acids in the vicinity of those indicated above that are involved in nuclease activity. In some embodiments, only the RuvC domain is inactivated, and in other embodiments, another putative nuclease domain is inactivated, where the effector protein complex functions as a nickase to cleave only one DNA strand. In a preferred embodiment, the other putative nuclease domain is a HincII-like endonuclease domain. In some embodiments, two AsCpf1 mutants (each a different nickase) are used to increase specificity and two nickase mutants are used to cleave the DNA at the target (where both nickases cleave the DNA strand while minimizing or eliminating off-target modifications, where only one DNA strand is cleaved and subsequently repaired). In a preferred embodiment, the Cpf1 effector protein cleaves a sequence associated with or present at the target locus of interest as a homodimer comprising two Cpf1 effector protein molecules. In a preferred embodiment, the homodimer can comprise two Cpf1 effector protein molecules each containing a different mutation in their respective RuvC domain.
[0044] In certain embodiments, the CRISPR enzyme can include one or more mutations that are engineered to modify its activity, specificity, and / or stability. Amino acid positions in the AsCpf1p enzyme include, but are not limited to: based on the amino acid numbering of AsCpf1 (Acidaminococcus sp. BV3L6), D861, R862, R863, W382, E993, D1263, D908, W958, K968, R951, R1226, S1228, D1235, K548, M604, K607, T167, N631, N630, K547, K163, Q571, K1017, R955, K1009, R909, R912, R1072, E372, K15, K810, H755, K557, E857, K943, K1022, K1029, K942, K949, R84, K87, K200, H206, R210, R301, R699, K705, K887, R891, K1086, K1089, R1094, R1127, R1220, Q1224, N178, N197, N204, N259, N278, N282, N519, N747, N759, N878, N889, and / or any one amino acid in the regions of 1189 - 1197, 1200 - 1208, 398 - 400, 380 - 383, 362 - 420, 1163 - 1173, 1230 - 1233, 1152 - 1148, 1076 - 1249. In preferred embodiments, these one or more mutations are selected from, but are not limited to: R862A, E993A, D1263A, D908A, W958A, R951A, R1226A, S1228A, D1235A, K548A, M604A, K607A, K607R, T167S, N631K, N613R, N630K, N630R, K547R, K163R, Q571K, Q571R, K1009A, R909A, R1072A, E327A, K15A, K810A, H755A, K557A, E857A, K943A, K1022A, K1029A, K942A, K949A, R84A, K87A, K200A, H206A, R210A, R301A, R699A, K705A, K887A, R891A, K1086A, K1089A, R1094A, R1127A, R1220A, and Q1224A.
[0045] In other preferred embodiments, the one or more mutations are selected from R862A, E993A, D1263A, D908A, W958A, R951A, K548A, M604A, K607A, K607R, N631K, N613R, N630K, N630R, K547R, K163R, Q571K, Q571R, K1009A, R909A, R1072A, E327A, K15A, K810A, H755A, K557A, E857A, K943A, K1022A, K1029A, K942A, K949A, R84A, K87A, K200A, H206A, R210A, R301A, R699A, K705A, K887A, R891A, K1086A, K1089A, R1094A, R1127A, R1220A, and Q1224A.
[0046] In certain embodiments, the one or more Cpf1 mutations result in nickase activity. In certain embodiments, the mutation is located at the position of the second nuclease domain, and more specifically, the mutation corresponds to R1226 of AsCpf1. In certain embodiments, the one or more mutations result in cleavage of only the non-target strand and non-cleavage of the target strand. In certain embodiments, the mutation is R1226A.
[0047] The present invention contemplates methods of using two or more nickases, specifically the dual or double nickase approach. In some aspects and embodiments, a single type of AsCpf1 nickase, such as a modified AsCpf1 or modified AsCpf1 nickase as described herein, may be delivered. This will result in two AsCpf1 nickases binding to the target DNA. Additionally, it is also envisioned that different orthologs, such as an AsCpf1 nickase for one strand of DNA (e.g., the coding strand) and an ortholog for the non-coding strand or the opposite DNA strand, may be used. It may be advantageous to use two different orthologs that require different PAMs and also have different guide requirements, thus allowing for greater user control. In certain embodiments, DNA cleavage will involve at least four types of nickases, each type being guided to a different sequence of the target DNA, where each pair introduces a first nick in one DNA strand and the second pair introduces a nick in the second DNA strand. In such a method, at least two single-strand cleavage pairs are introduced into the target DNA, and when the first and second single-strand cleavage pairs are introduced, the target sequence between the first and second single-strand cleavage pairs is excised. In certain embodiments, one or both of the orthologs are controllable, i.e., inducible.
[0048] In a detailed embodiment, the present invention provides a method of modifying an organism or non-human organism by minimizing off-target modification by manipulating first and second target sequences on the reverse strand of a DNA double strand at a target genomic locus of a cell, comprising delivering a non-naturally occurring or engineered composition, the composition comprising: - A first guide sequence linked to a direct repeat sequence, the polynucleotide sequence encoding a first type V CRISPR-Cas polynucleotide sequence comprising a guide RNA comprising a guide sequence capable of hybridizing to the first target sequence; - A second guide RNA comprising a guide sequence linked to a direct repeat array and comprising a guide sequence capable of hybridizing with the second target sequence, a polynucleotide sequence encoding a second type V CRISPR-Cas polynucleotide sequence, and - A polynucleotide sequence encoding a Cpf1 effector protein comprising at least one or more nuclear localization sequences and comprising one or more mutations, Upon transcription, the first and second guide RNAs direct the sequence-specific binding of the first and second CRISPR complexes to the first and second target sequences, respectively. The first CRISPR complex comprises a Cpf1 enzyme complexed with a first guide RNA comprising a first guide sequence capable of hybridizing to the first target sequence, and the second CRISPR complex comprises a Cpf1 enzyme complexed with a second guide RNA comprising a guide sequence capable of hybridizing to the second target sequence. The polynucleotide sequence encoding the CRISPR enzyme is DNA or RNA, and the first guide sequence directs cleavage of one strand of the DNA duplex near the first target sequence, and the second guide sequence directs cleavage of the other strand near the second target sequence to induce a double-strand break, thereby modifying an organism or a non-human organism by minimizing off-target modifications. In a detailed embodiment, when the first guide sequence directs cleavage of one strand of the DNA duplex near the first target sequence and the second guide sequence directs cleavage of the other strand near the second target sequence, a 5' overhang is generated. In a detailed embodiment, the 5' overhang is at most 200 base pairs. In a detailed embodiment, the 5' overhang is at most 100 base pairs, or at most 50 base pairs. In a detailed embodiment, the 5' overhang is at least 26 or at least 30 base pairs. In a detailed embodiment, the 5' overhang is 1 to 100, 1 to 34 base pairs or 34 to 50 base pairs. In a detailed embodiment, the 5' overhang is at least 1, at least 10, or at least 15 base pairs. In a detailed embodiment, when the first guide sequence directs cleavage of one strand of the DNA duplex near the first target sequence and the second guide sequence directs cleavage of the other strand near the second target sequence, blunt-end cleavage occurs. In a detailed embodiment, the Cpf1 mutation is R1226A.In a detailed embodiment, the present invention provides a method for modifying an organism or a non-human organism by minimizing off-target modification by manipulating first and second target sequences located on the reverse strand of a DNA double strand at a genomic locus of interest in a cell, the method comprising delivering an unnatural or engineered composition comprising a vector system comprising one or more vectors, the vector comprising: I. a first regulatory element operably linked to a first guide RNA comprising a first guide sequence capable of hybridizing to a first target sequence; II. a second regulatory element operably linked to a second guide RNA comprising a second guide sequence capable of hybridizing to a second target sequence; and III. a third regulatory element operably linked to an enzyme coding sequence encoding a Cpf1 enzyme, wherein components I, II, and III are located on the same or different vectors of the system and, when transcribed, the first and second guide sequences direct sequence-specific binding of the first and second CRISPR complexes to the first and second target sequences, respectively, the first CRISPR complex comprising a Cpf1 enzyme complexed with a first guide RNA comprising a first guide sequence capable of hybridizing to the first target sequence, the second CRISPR complex comprising a Cpf1 enzyme complexed with a second guide RNA comprising a second guide sequence capable of hybridizing to the second target sequence, the polynucleotide sequence encoding the Cpf1 enzyme being DNA or RNA, and the first guide sequence directs cleavage of one strand of the DNA double strand in the vicinity of the first target sequence and the second guide sequence directs cleavage of the other strand in the vicinity of the second target sequence to induce a double-strand break, thereby modifying the organism or non-human organism by minimizing off-target modification.In a detailed embodiment, the present invention provides a method for modifying a target genomic locus by minimizing off-target modifications by introducing an engineered, non-naturally occurring CRISPR-Cas system that comprises a Cpf1 effector protein having one or more mutations into a cell that contains and expresses a double-stranded DNA molecule encoding a gene product, and two guide RNAs that target the first and second strands, respectively, of the DNA molecule, whereby the guide RNAs target the DNA molecule encoding the gene product and the Cpf1 effector protein makes nicks in each of the first and second strands of the DNA molecule encoding the gene product, thereby changing the expression of the gene product; and wherein the Cpf1 effector protein and the two guide RNAs do not naturally occur together.
[0049] The present invention further provides an engineered, non-naturally occurring CRISPR-Cpf1 system comprising a Cpf1 protein having one or more mutations, and two guide RNAs that target, respectively, the first and second strands of a double-stranded DNA molecule encoding a cellular gene product, whereby the guide RNAs target the DNA molecule encoding the gene product and the Cpf1 protein makes nicks in each of the first and second strands of the DNA molecule encoding the gene product, whereby the expression of the gene product is altered; and wherein the Cpf1 protein and the two guide RNAs do not naturally occur together. In a specific embodiment, the Cpf1 mutation is R1226A. The present invention further provides an engineered, non-naturally occurring vector system comprising one or more vectors comprising: a) a first regulatory element operably linked to each of two CRISPR-Cpf1 system guide RNAs that target, respectively, the first and second strands of a double-stranded DNA molecule encoding a gene product; b) a second regulatory element operably linked to the Cpf1 protein, wherein components (a) and (b) are located on the same or different vectors of the system, whereby the guide RNAs target the DNA molecule encoding the gene product and the Cpf1 protein makes nicks in each of the first and second strands of the DNA molecule encoding the gene product, whereby the expression of the gene product is altered; and wherein the Cpf1 protein and the two guide RNAs do not naturally occur together.
[0050] The present invention further provides a method for modifying an organism comprising promoting homologous recombination repair to include delivering a non-naturally occurring or engineered composition, the composition comprising first and second target sequences on the reverse strand of a DNA double strand at a target genomic locus of a cell, the composition comprising: I. a first CRISPR-Cpf1-based guide RNA polynucleotide sequence comprising a first guide sequence hybridizable to the first target sequence and a direct repeat sequence; II. a second CRISPR-Cpf1-based RNA polynucleotide sequence comprising a second guide sequence hybridizable to the second target sequence and a direct repeat sequence; III. a polynucleotide sequence encoding a Cpf1 enzyme comprising at least one or more nuclear localization sequences and comprising one or more mutations; and IV. a repair template comprising a synthetic or engineered single-stranded oligonucleotide, which when transcribed, the first and second Cpf1 guide RNAs direct sequence-specific binding of the first and second CRISPR complexes to the first and second target sequences, respectively, the first CRISPR complex comprising a Cpf1 enzyme complexed with a first Cpf1 guide RNA comprising a first guide sequence hybridizable to the first target sequence, the second CRISPR complex comprising a Cpf1 enzyme complexed with a second Cpf1 guide RNA comprising a second guide sequence hybridizable to the second target sequence, the polynucleotide sequence encoding the Cpf1 enzyme being DNA or RNA, the first guide sequence directs cleavage of one strand of the DNA double strand near the first target sequence, and the second guide sequence directs cleavage of the other strand near the second target sequence to induce a double-strand break, and the repair template is introduced into the DNA double strand by homologous recombination, thereby modifying the organism.
[0051] The present invention further provides a method for modifying an organism comprising promoting non-homologous end joining (NHEJ)-mediated ligation to include delivering a non-naturally occurring or engineered composition, the composition comprising first and second target sequences on the reverse strand of a DNA double strand at a target genomic locus of a cell, the composition comprising I. A first Cpf1 guide RNA polynucleotide sequence, a first polynucleotide sequence comprising a first guide sequence capable of hybridizing to a first target sequence and a direct repeat sequence; II. A second Cpf1 guide RNA polynucleotide sequence, a second polynucleotide sequence comprising a second guide sequence capable of hybridizing to a second target sequence and a direct repeat sequence; III. A polynucleotide sequence encoding a Cpf1 enzyme comprising at least one or more nuclear localization sequences and comprising one or more mutations; and IV. A repair template comprising a first overhang set, which, when transcribed, causes the first and second guide sequences to direct sequence-specific binding of the first and second CRISPR complexes to the first and second target sequences, respectively, the first CRISPR complex comprising a Cpf1 enzyme complexed with a first guide RNA comprising a first guide sequence capable of hybridizing to the first target sequence, the second CRISPR complex comprising a Cpf1 enzyme complexed with a second guide RNA comprising a second guide sequence capable of hybridizing to the second target sequence, the polynucleotide sequence encoding the Cpf1 enzyme being DNA or RNA, the first guide sequence causing cleavage of one strand of the DNA duplex in the vicinity of the first target sequence, and the second guide sequence causing cleavage of the other strand in the vicinity of the second target sequence to induce a double-strand break having a second overhang set, the first overhang set being compatible with and matching the second overhang set, and the repair template being introduced into the DNA duplex by ligation, thereby modifying the organism.
[0052] The present invention further provides I. A first polynucleotide comprising a first guide sequence capable of hybridizing to a first target sequence and a direct repeat sequence; II. A second polynucleotide comprising a second guide sequence capable of hybridizing to a second target sequence and a direct repeat sequence; and III. A third polynucleotide comprising a sequence encoding a Cpf1 enzyme and one or more nuclear localization sequences, wherein the first target sequence is on the first strand of a DNA duplex and the second target sequence is on the reverse strand of the DNA duplex, and when the first and second guide sequences hybridize to the target sequences of the duplex, the 5' ends of the first polynucleotide and the second polynucleotide are offset from each other by at least one base pair of the duplex, and optionally, each of I, II, and III is provided in the same or different vectors, A kit or composition is provided. The present invention further relates to the use of a kit as described herein in the methods described herein. The present invention further provides a composition as described herein for use as a medicament, more particularly for the treatment or prevention of a disease caused by a defect in a locus corresponding to a target sequence.
[0053] A Cpf1 enzyme as defined herein can utilize two or more RNA guides without loss of activity. This enables targeting of multiple DNA targets, genes, or loci with a single enzyme, system, or complex as defined herein using a Cpf1 enzyme, system, or complex as defined herein. Guide RNAs may be arranged tandemly and optionally separated by a nucleotide sequence, but preferably the guide RNAs are directly linked, i.e., two or more guide RNAs are directly linked to each other such that in each guide RNA the direct repeat is on the 5' side of the guide sequence, such that each guide sequence is adjacent to the direct repeat of the adjacent guide RNA. When the Cpf1 enzyme used is R1226A of AsCpf1, the non-target strand will be cleaved and there will be no cleavage of the target strand. This information is relevant for guide design. The positions of these different guide RNAs are tandem and do not affect activity. As a further guide, the following specific aspects and embodiments are provided.
[0054] In one aspect, the invention provides for the use of a Cpf1 enzyme, complex, or system as defined herein for targeting multiple loci. In one embodiment, this can be achieved by using multiple (tandem or multiplexed) guide RNA (gRNA) sequences. A Cpf1 enzyme, system, or complex as defined herein provides an effective means for modifying multiple target polynucleotides. A Cpf1 enzyme, system, or complex as defined herein has a wide range of utility, including modifying (e.g., deleting, inserting, translocating, inverting, activating) one or more target polynucleotides in a number of cell types. Thus, a Cpf1 enzyme, system, or complex as defined herein has a broad range of applicability, including targeting multiple loci within a single CRISPR system, for example, in gene therapy, drug screening, disease diagnosis, and prognosis determination.
[0055] The present invention encompasses a guide RNA comprising guide sequences arranged in tandem. The present invention further encompasses a Cpf1 protein coding sequence codon-optimized for expression in eukaryotic cells. In preferred embodiments, the eukaryotic cells are mammalian cells, plant cells or yeast cells, and in more preferred embodiments, the mammalian cells are human cells. Expression of the gene product may be reduced. The Cpf1 enzyme may form part of a CRISPR system or complex, and the CRISPR system or complex further comprises a guide RNA (gRNA) arranged in tandem, each containing a series of 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 25, 25, 30, or more than 30 guide sequences capable of specifically hybridizing to a target sequence at a target genomic locus of the cell. In some embodiments, the functional Cpf1 CRISPR system or complex binds to multiple target sequences. In some embodiments, the functional CRISPR system or complex can edit multiple target sequences, for example, the target sequences may include genomic loci, and in some embodiments, there may be a change in gene expression. In some embodiments, the functional CRISPR system or complex may include additional functional domains. In some embodiments, the present invention provides a method of altering or modifying the expression of multiple gene products. The method may include introducing into a cell containing the target nucleic acid, for example a DNA molecule, or expressing the target nucleic acid, for example a DNA molecule; for example, the target nucleic acid may encode a gene product or may result in the expression of a gene product (e.g., regulatory sequences).
[0056] In preferred embodiments, the CRISPR enzyme used for multiplex targeting is AsCpf1, or the CRISPR system or complex used for multiplex targeting comprises AsCpf1. In some embodiments, the CRISPR enzyme is LbCpf1, or the CRISPR system or complex comprises LbCpf1. In some embodiments, the Cpf1 enzyme used for multiplex targeting cleaves both strands of DNA to create a double-strand break (DSB). In some embodiments, the CRISPR enzyme used for multiplex targeting is a nickase. In some embodiments, the Cpf1 enzyme used for multiplex targeting is a dual nickase.
[0057] In certain embodiments of the invention, the guide RNA or mature crRNA comprises, consists essentially of, or consists of a direct repeat sequence and a guide sequence or spacer sequence. In certain embodiments, the guide RNA or mature crRNA comprises, consists essentially of, or consists of a direct repeat sequence linked to a guide sequence or spacer sequence. In certain embodiments, the guide RNA or mature crRNA comprises a 19 nt partial direct repeat and a subsequent guide sequence or spacer sequence of 20 - 30 nt, preferably about 20 nt, 23 - 25 nt, or 24 nt. In certain embodiments, the effector protein is an AsCpf1 effector protein, requires at least a 16 nt guide sequence to achieve detectable DNA cleavage, and requires at least a 17 nt guide sequence to achieve efficient DNA cleavage in vitro. In certain embodiments, the direct repeat sequence is located upstream (i.e., 5' side) of the guide sequence or spacer sequence. In preferred embodiments, the seed sequence of the AsCpf1 guide RNA (i.e., the important sequence essential for recognition and / or hybridization with the sequence of the target locus) is within the first approximately 5 nt at the 5' end of the guide sequence or spacer sequence.
[0058] In a preferred embodiment of the present invention, the mature crRNA comprises a stem-loop or an optimized stem-loop structure or an optimized secondary structure. In a preferred embodiment, the mature crRNA comprises a stem-loop or an optimized stem-loop structure in the direct repeat sequence, where the stem-loop or the optimized stem-loop structure is important for cleavage activity. In certain embodiments, the mature crRNA preferably comprises a single stem-loop. In certain embodiments, the direct repeat sequence preferably comprises a single stem-loop. In certain embodiments, the cleavage activity of the effector protein complex is modified by introducing mutations that affect the stem-loop RNA duplex structure. In a preferred embodiment, mutations that maintain the RNA duplex of the stem-loop may be introduced, thereby maintaining the cleavage activity of the effector protein complex. In other preferred embodiments, mutations that disrupt the RNA duplex structure of the stem-loop may be introduced, thereby completely abolishing the cleavage activity of the effector protein complex.
[0059] The present invention also provides a nucleotide sequence encoding an effector protein codon-optimized for expression in a eukaryote or eukaryotic cell in any of the methods or compositions described herein. In certain embodiments of the invention, the codon-optimized effector protein is AsCpf1p and is codon-optimized for operability in a eukaryotic cell or organism, such as a cell or organism as listed in other parts of this specification, such as, without limitation, a yeast cell, or a mammalian cell or organism, such as a mouse cell, a rat cell, and a human cell or a non-human eukaryote, such as a plant.
[0060] In certain embodiments of the invention, at least one nuclear localization signal (NLS) is added to the nucleic acid sequence encoding the Cpf1 effector protein. In preferred embodiments, at least one or more C-terminal or N-terminal NLSs are added (thus one or more nucleic acid molecules encoding the Cpf1 effector protein can include the coding of one or more NLSs, so that one or more NLSs are added or connected to the expressed product). In preferred embodiments, a C-terminal NLS is added for optimal expression and nuclear targeting in eukaryotic cells, preferably human cells. In preferred embodiments, the codon-optimized effector protein is AsCpf1p and the spacer length of the guide RNA is 15 - 35 nt. In certain embodiments, the spacer length of the guide RNA is at least 16 nucleotides, such as at least 17 nucleotides. In certain embodiments, the spacer length is 15 - 17 nt, 17 - 20 nt, 20 - 24 nt, such as 20, 21, 22, 23, or 24 nt, 23 - 25 nt, such as 23, 24, or 25 nt, 24 - 27 nt, 27 - 30 nt, 30 - 35 nt, or 35 nt or more. In certain embodiments of the invention, the codon-optimized effector protein is AsCpf1p and the length of the direct repeat of the guide RNA is at least 16 nucleotides. In certain embodiments, the codon-optimized effector protein is AsCpf1p and the length of the direct repeat of the guide RNA is 16 - 20 nt, such as 16, 17, 18, 19, or 20 nucleotides. In certain preferred embodiments, the length of the direct repeat of the guide RNA is 19 nucleotides.
[0061] The present invention also encompasses a method of delivering a plurality of nucleic acid components, wherein each nucleic acid component is specific for a different target locus of interest, thereby modifying a plurality of target loci of interest. The nucleic acid components of the complex can include one or more protein-binding RNA aptamers. The one or more aptamers can have the ability to bind to a bacteriophage coat protein. The bacteriophage coat protein can be selected from the group consisting of Qβ, F2, GA, fr, JP501, MS2, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s, and PRR1. In a preferred embodiment, the bacteriophage coat protein is MS2. The present invention also provides nucleic acid components of the complex that are 30 or more, 40 or more, or 50 or more nucleotides in length.
[0062] The present invention also encompasses the cells, components, and / or systems of the present invention in which trace amounts of cations are present. Advantageously, the cation is magnesium, such as Mg 2+ . The cation can be present in trace amounts. A preferred range can be from about 1 mM to about 15 mM of the cation, which is advantageously Mg 2+ . Preferred concentrations can be about 1 mM for human-based cells, components, and / or systems, and about 10 mM to about 15 mM for bacteria-based cells, components, and / or systems. See, for example, Gasiunas et al., PNAS, published online September 4, 2012, www.pnas.org / cgi / doi / 10.1073 / pnas.1208507109.
[0063] Accordingly, it is an object of the present invention not to encompass within the scope of the invention any previously known product, process for making the product, or method of using the product which the applicants reserve the right to disclaim and which is hereby disclosed by this specification as a disclaimer of any previously known product, process, or method. Further, it is noted that the present invention is intended not to encompass within the scope of the invention any product, process for making the product, or method of using the product which does not meet the written description and enablement requirements of the specification of the United States Patent and Trademark Office (USPTO) (35 U.S.C. § 112, first paragraph) or the European Patent Office (EPO) (Article 83 EPC), which the applicants reserve the right to disclaim and which is hereby disclosed by this specification as a disclaimer of any previously described product, process, or method. In practicing the present invention, it may be advantageous to comply with Article 53(c) EPC and Rules 28(b) and (c) EPC. Nothing in this specification should be construed as a promise.
[0064] In this disclosure, and particularly in the claims and / or paragraphs thereof, the terms "comprises," "comprised of," "comprising," etc. may have the meaning ascribed to them in the United States Patent Laws; for example, these may mean "includes," "included," "including," etc.; and it is noted that the terms "consisting essentially of" and "consists essentially of" may have the meaning ascribed to them in the United States Patent Laws.
[0065] The above and other embodiments are disclosed or are apparent from and included in the following detailed description.
[0066] The novel features of the present invention are set forth in detail in the appended claims. A further understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description which illustrates exemplary embodiments in which the principles of the invention are utilized, and to the accompanying drawings.
Brief Description of the Drawings
[0067]
FIG. 1A-C
FIG. 2A
FIG. 2B
FIG. 3
FIG. 4A
FIG. 4B
FIG. 5
FIG. 6A
FIG. 6B
FIG. 7
FIG. 8
FIG. 9A
FIG. 9B
FIG. 10
FIG. 11
FIG. 12A
FIG. 12B
FIG. 13
FIG. 14
FIG. 15A-D
FIG. 16A-I
FIG. 17
FIG. 18A-E
FIG. 19A-F
FIG. 20A-F
FIG. 21A-F
FIG. 22
FIG. 23A-B
FIG. 24A-C
FIG. 25A-B
FIG. 26A-C
FIG. 27
FIG. 28A-B
FIG. 29A-B
Best Mode for Carrying Out the Invention
[0068] The figures in this specification are for illustrative purposes only and are not necessarily drawn to scale.
[0069] This application describes the crystal structure of a Cpf1 effector protein. The Cpf1 effector protein is functionally different from the previously described CRISPR-Cas9 system, and thus the terminology for elements related to these novel endonucleases is accordingly modified herein. The Cpf1-related CRISPR arrays described herein are processed into mature crRNAs without the need for an additional tracrRNA. The crRNAs described herein contain a spacer sequence (or guide sequence) and a direct repeat sequence, and only the Cpf1p-crRNA complex is sufficient for efficient cleavage of target DNA. The seed sequence described herein, for example, the seed sequence of the AsCpf1 guide RNA, is within the first approximately 5 nt at the 5' end of the spacer sequence (or guide sequence), and mutations within the seed sequence adversely affect the cleavage activity of the Cpf1 effector protein complex.
[0070] Generally, the CRISPR system is characterized by elements that promote the formation of a CRISPR complex at the site of a target sequence (also referred to as a protospacer in the context of an endogenous CRISPR system). In the context of CRISPR complex formation, a "target sequence" refers to a sequence that the guide sequence targets, e.g., is designed to have complementarity therewith, where hybridization between the target sequence and the guide sequence promotes the formation of the CRISPR complex. The section of the guide sequence where complementarity with the target sequence is important for cleavage activity is referred to herein as the seed sequence. The target sequence can comprise any polynucleotide, such as a DNA or RNA polynucleotide, and is contained within the target locus of interest. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell.
[0071] The term "nucleic acid targeting system", wherein the nucleic acid is DNA or RNA and in some embodiments can also refer to a DNA-RNA hybrid or derivative thereof, collectively refers to transcripts and other elements that are involved in the expression of DNA or RNA targeting CRISPR-associated ("Cas") genes or induce their activity, where the gene includes a sequence encoding a DNA or RNA targeting Cas protein and a CRISPR RNA (crRNA) sequence and (although not in all systems, in the CRISPR-Cas9 system) a trans-activating CRISPR-Cas system RNA (tracrRNA) sequence, a DNA or RNA targeting guide RNA, or other sequences and transcripts from a DNA or RNA targeting CRISPR locus. In the Cpf1 DNA targeting RNA guide endonuclease system described herein, the tracrRNA sequence is not required. Generally, an RNA targeting system is characterized by elements that promote the formation of an RNA targeting complex at the site of a target RNA sequence. In the context of the formation of a DNA or RNA targeting complex, a "target sequence" refers to a DNA or RNA sequence that is designed such that a DNA or RNA targeting guide RNA has complementarity thereto, where hybridization between the target sequence and the RNA targeting guide RNA promotes the formation of the RNA targeting complex. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell. In some embodiments, the target sequence can be within an organelle of a eukaryotic cell, such as within a mitochondrion or chloroplast. A sequence or template that can be used for recombination into a target locus containing the target sequence is referred to as an "editing template" or "editing RNA" or "editing sequence". In aspects of the invention, an exogenous template RNA can be referred to as an editing template. In certain aspects of the invention, the recombination is homologous recombination.
[0072] The nucleic acid targeting systems, vector systems, vectors, and compositions described herein can be used in various nucleic acid targeting applications, alteration or modification of the synthesis of gene products such as proteins, nucleic acid cleavage, nucleic acid editing, nucleic acid splicing; transport of target nucleic acids, tracking of target nucleic acids, isolation of target nucleic acids, visualization of target nucleic acids, and the like.
[0073] As used herein, a Cas protein or CRISPR enzyme refers to any of the proteins presented in the new classification of the CRISPR-Cas system. In advantageous embodiments, the invention encompasses effector proteins identified in type V CRISPR-Cas loci, such as the Cpf1-encoding locus, also referred to as subtype V-A. Currently, the subtype V-A locus contains distinct genes and a CRISPR array designated as cas1, cas2, cpf1. Cpf1 (CRISPR-associated protein Cpf1, subtype PREFRAN) is a large protein (about 1300 amino acids) that includes a RuvC-like nuclease domain homologous to the corresponding domain of Cas9 along with what corresponds to the characteristic arginine-rich cluster of Cas9. However, Cpf1 lacks the HNH nuclease domain present in all Cas9 proteins, and in contrast to Cas9 that contains a long insert including the HNH domain, the RuvC-like domain is continuous in the Cpf1 sequence. Thus, in a detailed embodiment, the CRISPR-Cas enzyme contains only the RuvC-like nuclease domain.
[0074] The Cpf1 gene is found in several diverse bacterial genomes and is typically at the same locus as the cas1, cas2, and cas4 genes and the CRISPR cassette (e.g., FNFX1_1431 - FNFX1_1428 of Francisella cf. novicida Fx1). Thus, the layout of this putative novel CRISPR-Cas system seems to be similar to type II-B. Furthermore, like Cas9, the Cpf1 protein contains an easily identifiable C-terminal region homologous to transposon ORF-B and includes an active RuvC-like nuclease, an arginine-rich region, and a Zn finger (absent in Cas9). However, unlike Cas9, Cpf1 is also present in several genomes without a CRISPR-Cas context, and its relatively high similarity to ORF-B suggests that it may be a transposon component. If this were a bona fide CRISPR-Cas system, it was suggested that Cpf1 would be a functional analog of Cas9 and a novel CRISPR-Cas type, namely type V (see Annotation and Classification of CRISPR-Cas Systems. Makarova KS, Koonin EV. Methods Mol Biol. 2015;1311:47 - 75).
[0075] Aspects of the invention also include methods and uses of the compositions and systems described herein for, e.g., altering or manipulating the expression of one or more genes or one or more gene products in in vitro, in vivo, or ex vivo genome engineering in prokaryotic or eukaryotic cells.
[0076] In embodiments of the present invention, the terms mature crRNA, guide RNA, and single guide RNA are used interchangeably as described in the foregoing cited references such as WO 2014 / 093622 (PCT / US2013 / 074667). Generally, a guide sequence is any polynucleotide sequence having complementarity with a target polynucleotide sequence sufficient to hybridize with the target sequence and direct sequence-specific binding of the CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between the guide sequence and its corresponding target sequence is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or higher when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any algorithm suitable for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transform (e.g., the Burrows-Wheeler Aligner), ClustalW, ClustalX, BLAT, Novoalign (available at www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequence is about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 nucleotides in length or longer. In some embodiments, the guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12 nucleotides in length or shorter. Preferably, the guide sequence is 10-30 nucleotides in length. The ability of the guide sequence to direct sequence-specific binding of the CRISPR complex to the target sequence can be evaluated by any suitable assay.For example, the components of the CRISPR system sufficient to form a CRISPR complex may be provided to a host cell having a corresponding target sequence, such as by transfection of a vector encoding the components of the CRISPR array, including the guide sequence to be tested, and subsequently, the preferential cleavage within the target sequence may be evaluated by, for example, the Surveyor assay as described herein. Similarly, cleavage of the target polynucleotide sequence can be determined in vitro by providing the target sequence, the components of the CRISPR complex including the guide sequence to be tested, and a control guide sequence different from the test guide sequence, and comparing the binding or cleavage rates in the target sequence between the reactions of the test guide sequence and the control guide sequence. Other assays are possible and will be apparent to those skilled in the art. The guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the genome of the cell. Exemplary target sequences include those that are unique in the target genome.
[0077] Generally, and throughout this specification, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is ligated. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules without free ends (e.g., circular) that contain one or more free ends; nucleic acid molecules containing DNA, RNA, or both; and other types of polynucleotides known in the art. Certain vectors are "plasmids", which refer to circular double-stranded DNA loops into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, where the vector contains viral-derived DNA or RNA sequences for packaging into a virus (e.g., retrovirus, replication-defective retrovirus, adenovirus, replication-defective adenovirus, and adeno-associated virus). Viral vectors also include polynucleotides carried by a virus for transfection of a host cell. Certain vectors have the ability to self-replicate in the host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of the host cell upon introduction into the host cell and are thereby replicated with the host genome. Further, certain vectors are capable of directing the expression of a gene to which they are operably linked. Such vectors are referred to herein as "expression vectors". Vectors for expression in and that effect expression in eukaryotic cells can be referred to herein as "eukaryotic cell expression vectors". Common expression vectors useful in recombinant DNA techniques are often in the form of plasmids.
[0078] The recombinant expression vector can contain the nucleic acid of the present invention in a form suitable for the expression of the nucleic acid in a host cell, which means that the recombinant expression vector contains one or more regulatory elements operably linked to the nucleic acid sequence to be expressed (which may be selected based on the host cell used for expression). Within the scope of the recombinant expression vector, "operably linked" is intended to mean that the nucleotide sequence of interest is linked to one or more regulatory elements in such a way that expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell) is possible.
[0079] The term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRESs), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and polyU sequences). Such regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in a particular host cell (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in a desired target tissue, such as muscle, neuron, bone, skin, blood, a particular organ (e.g., liver, pancreas), or a particular cell type (e.g., lymphocyte). Regulatory elements can also direct expression in a time-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, and this expression can also be tissue- or cell-type specific or not. In some embodiments, the vector includes one or more pol III promoters (e.g., 1, 2, 3, 4, 5, or more pol III promoters), one or more pol II promoters (e.g., 1, 2, 3, 4, 5, or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5, or more pol I promoters), or combinations thereof. Examples of pol III promoters include, but are not limited to, the U6 and H1 promoters.Examples of pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer) [see, e.g., Boshart et al, Cell, 41:521-530 (1985)], the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. Also included within the term "regulatory element" are enhancer elements such as WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), p. 466-472, 1988); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), p. 1527-31, 1981). Those skilled in the art will understand that the design of the expression vector can depend on factors such as the choice of host cell to be transformed and the desired level of expression. The vector can be introduced into the host cell, whereby transcripts, proteins, or peptides encoded by the nucleic acids as described herein can be produced, including fusion proteins or peptides (e.g., clustered regularly interspaced short palindromic repeats (CRISPR) transcripts, proteins, enzymes, their mutants, their fusion proteins, etc.).
[0080] Advantageous vectors include lentiviruses and adeno-associated viruses, and the type of such vectors can also be selected to target specific types of cells.
[0081] As used herein, the term "crRNA" or "guide RNA" or "single guide RNA" or "sgRNA" or "one or more nucleic acid components" of a type V or type VI CRISPR-Cas locus effector protein includes any polynucleotide sequence having complementarity with a target nucleic acid sequence sufficient to hybridize to the target nucleic acid sequence and direct sequence-specific binding of a nucleic acid targeting complex to the target nucleic acid sequence. In some embodiments, the degree of complementarity is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or higher when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any algorithm suitable for alignment of sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transform (e.g., the Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (available at www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). The ability of a guide sequence (within a nucleic acid targeting guide RNA) to direct sequence-specific binding of a nucleic acid targeting complex to a target nucleic acid sequence can be evaluated by any suitable assay. For example, the components of a nucleic acid targeting CRISPR system sufficient to form a nucleic acid targeting complex may be provided, including the guide sequence to be tested, by transfection of a vector encoding the components of the nucleic acid targeting complex into a host cell having the corresponding target nucleic acid sequence, and subsequent evaluation of preferential targeting (e.g., cleavage) within the target nucleic acid sequence by, for example, the Surveyor assay as described herein.Similarly, cleavage of a target nucleic acid sequence can be determined in vitro by providing the target nucleic acid sequence, the components of the nucleic acid targeting complex including the guide sequence to be tested, and a control guide sequence different from the test guide sequence, and comparing the binding or cleavage rates at the target sequence between the reaction of the test guide sequence and the control guide sequence. Other assays are possible and will be apparent to those skilled in the art. The guide sequence, and thus the nucleic acid targeting guide RNA, can be selected to target any target nucleic acid sequence. The target sequence can be DNA. The target sequence can be any RNA sequence. In some embodiments, the target sequence can be a sequence within an RNA molecule selected from the group consisting of messenger RNA (mRNA), pre-mRNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), double-stranded RNA (dsRNA), non-coding RNA (ncRNA), long non-coding RNA (lncRNA), and small cytoplasmic RNA (scRNA). In some preferred embodiments, the target sequence can be a sequence within an RNA molecule selected from the group consisting of mRNA, pre-mRNA, and rRNA. In some preferred embodiments, the target sequence can be a sequence within an RNA molecule selected from the group consisting of ncRNA and lncRNA. In some more preferred embodiments, the target sequence can be a sequence within an mRNA molecule or a pre-mRNA molecule.
[0082] In some embodiments, the nucleic acid targeting guide RNA is selected such that the degree of secondary structure within the RNA targeting guide RNA is reduced. In some embodiments, when optimally folded, about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1% or less of the nucleotides of the nucleic acid targeting guide RNA are involved in self-complementary base pairing, or less. Optimal folding can be determined by any suitable polynucleotide folding algorithm. Some programs are based on the calculation of minimum Gibbs free energy. An example of such an algorithm is mFold as described by Zuker and Stiegler (Nucleic Acids Res. 9(1981), 133-148). Another exemplary folding algorithm is the online web server RNAfold using the centroid structure prediction algorithm developed at the Institute for Theoretical Chemistry, University of Vienna (see, for example, A.R. Gruber et al., 2008, Cell 106(1):23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12):1151-62).
[0083] The "tracrRNA" sequence or similar terms include any polynucleotide sequence having sufficient complementarity to hybridize to the crRNA sequence. As pointed out above herein, in embodiments of the invention, tracrRNA is not required for the cleavage activity of the Cpf1 effector protein complex.
[0084] To minimize toxicity and off-target effects, it can be important to control the concentration of the nucleic acid targeting guide RNA delivered. The optimal concentration of the nucleic acid targeting guide RNA can be determined by testing various concentrations in cell models or non-human eukaryotic animal models and analyzing the degree of modification at potential off-target genomic loci using deep sequencing. The concentration that results in the highest level of on-target modification while minimizing the off-target modification level should be selected for in vivo delivery. The nucleic acid targeting system preferably derives from type V / VI CRISPR systems. In some embodiments, one or more elements of the nucleic acid targeting system are derived from a particular organism that includes an endogenous RNA targeting system. In a preferred embodiment of the invention, the RNA targeting system is a type V / VI CRISPR system. Homologs and orthologs can be identified by homology modeling (e.g., Greer, Science vol. 228 (1985) 1055, and Blundell et al. Eur J Biochem vol 172 (1988), 513) or "structural BLAST" (Dey F, Cliff Zhang Q, Petrey D, Honig B. "Toward a "structural BLAST": using structural relationships to infer function". Protein Sci. 2013 Apr;22(4):359-66. doi:10.1002 / pro.2225. See also Shmakov et al. (2015) with respect to applications in the field of CRISPR-Cas loci. However, homologous proteins may not be structurally related or may only be partially structurally related. In a detailed embodiment, a homolog or ortholog of Cpf1 as referred to herein has sequence homology or identity with Cpf1 of at least 80%, more preferably at least 85%, even more preferably at least 90%, such as at least 95%, etc.In further embodiments, homologs or orthologs of Cpf1 as referred to herein have sequence identity of at least 80%, more preferably at least 85%, even more preferably at least 90%, such as at least 95% etc. with wild-type Cpf1. When Cpf1 has one or more mutations (mutant type), the homologs or orthologs of said Cpf1 as referred to herein have sequence identity of at least 80%, more preferably at least 85%, even more preferably at least 90%, such as at least 95% etc. with mutant Cpf1.
[0085] In a detailed embodiment, homologs or orthologs of type V / VI proteins such as Cpf1 as referred to herein have sequence homology or identity of at least 80%, more preferably at least 85%, even more preferably at least 90%, such as at least 95% etc. with AsCpf1. In further embodiments, homologs or orthologs of type V / VI proteins such as AsCpf1 as referred to herein have sequence identity of at least 80%, more preferably at least 85%, even more preferably at least 90%, such as at least 95% etc. with AsCpf1.
[0086] In certain embodiments, the type V / VI RNA-targeting Cas protein may be a Cpf1 ortholog of an organism of a genus including, but not limited to, Corynebacter, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus, Lactobacillus, Mycoplasma, Bacteroides, Flaviivola, Flavobacterium, Sphaerochaeta, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus, Nitratifractor, Mycoplasma, and Campylobacter. The species of organisms of such genera may be as otherwise contemplated herein.
[0087] It will be understood that any of the functions described herein can be engineered into CRISPR enzymes from other orthologs, including chimeric enzymes that include fragments from multiple orthologs. Examples of such orthologs are described in other parts of this specification. Thus, chimeric enzymes can include, but are not limited to, fragments of CRISPR enzyme orthologs from organisms of genera including Corynebacter, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus, Lactobacillus, Mycoplasma, Bacteroides, Flaviivola, Flavobacterium, Sphaerochaeta, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus, Nitratifractor, Mycoplasma, and Campylobacter. The chimeric enzyme can include a first fragment and a second fragment, and these fragments can be those of CRISPR enzyme orthologs from organisms of the genera listed herein or species listed herein; advantageously the fragments are from CRISPR enzyme orthologs of different species.
[0088] In embodiments, the Cpf1 proteins as referred to herein also include functional variants of AsCpf1 or its homologs or orthologs. A "functional variant" of a protein, as used herein, refers to a variant of such protein that retains at least partially the activity of the said protein. Functional variants can include mutants (which can be insertion, deletion, or substitution mutants), including polymorphisms and the like. Also included within the scope of functional variants are fusion products of such proteins with another, usually unrelated nucleic acid, protein, polypeptide, or peptide. Functional variants may be naturally occurring or artificial. Advantageous embodiments may include engineered or non-naturally occurring AsCpf1 or its orthologs or homologs.
[0089] In certain embodiments, one or more nucleic acid molecules encoding AsCpf1 or its orthologs or homologs can be codon-optimized for expression in eukaryotic cells. The eukaryotes may be those as contemplated herein. The one or more nucleic acid molecules may be engineered or non-naturally occurring.
[0090] In certain embodiments, AsCpf1 or its orthologs or homologs can contain one or more mutations (and thus one or more nucleic acid molecules encoding it can have one or more mutations). The mutations may be artificially introduced mutations and can include, but are not limited to, one or more mutations in the catalytic domain. Examples of catalytic domains associated with the Cas9 enzyme can include, but are not limited to, the RuvC I, RuvC II, RuvC III, and HNH domains.
[0091] In certain embodiments, Cpf1 or an ortholog or homolog thereof can be used as a general nucleic acid binding protein fused to or operably linked to a functional domain. Exemplary functional domains include, but are not limited to, translation initiation factors, translation activation factors, translation repressors, nucleases, particularly ribonucleases, spliceosomes, beads, light-inducible / controllable domains or chemical-inducible / controllable domains.
[0092] In some embodiments, the non-modified nucleic acid targeting effector protein may have cleavage activity. In some embodiments, the RNA targeting effector protein may direct cleavage of one or both nucleic acid (DNA or RNA) strands at or near the position of the target sequence, such as within the target sequence and / or within the complement of the target sequence or in a sequence associated with the target sequence. In some embodiments, the nucleic acid targeting effector protein may direct cleavage of one or both DNA or RNA strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the cleavage is of the blunt-end type, i.e., it can result in blunt ends. In some embodiments, the cleavage is a blunt-end type cleavage with a 5' overhang. In some embodiments, the cleavage is a blunt-end type cleavage with a 5' overhang of 1 to 5 nucleotides, preferably 4 or 5 nucleotides. In some embodiments, the cleavage site is distant from the PAM. For example, cleavage occurs after the 18th nucleotide on the non-target strand and after the 23rd nucleotide on the target strand (Zetsche et al., 2015). In some embodiments, the cleavage site occurs after the 18th nucleotide (counting from the PAM) on the non-target strand and after the 23rd nucleotide (counting from the PAM) on the target strand. In some embodiments, the vector may encode a nucleic acid targeting effector protein that is mutated compared to the corresponding wild-type enzyme, such that the mutant nucleic acid targeting effector protein lacks the ability to cleave one or both DNA or RNA strands of the target polynucleotide containing the target sequence. As a further example, mutant Cas proteins lacking substantially all DNA cleavage activity may be generated by mutating two or more catalytic domains of the Cas protein (e.g., RuvC and, optionally, a second nuclease domain identified herein).As described herein, the corresponding catalytic domain of the Cpf1 effector protein can also be mutated to generate a mutant Cpf1 effector protein that lacks all DNA cleavage activity or has substantially reduced DNA cleavage activity. In some embodiments, a nucleic acid targeting effector protein is considered to lack substantially all RNA cleavage activity when the RNA cleavage activity of the mutant enzyme is about 25%, 10%, 5%, 1%, 0.1%, 0.01% or less, or less than that of the non-mutant enzyme; one example can be when the mutant nucleic acid cleavage activity is zero or negligible compared to the non-mutant form. Effector proteins can be identified by reference to a common enzyme class that shares homology with the largest nucleases having multiple nuclease domains from type V / VI CRISPR systems. Most preferably, the effector protein is a type V / VI protein such as Cpf1. In further embodiments, the effector protein is a type V protein. By "derived from," Applicants generally mean based on a wild-type enzyme in the sense that the derived enzyme has a high degree of sequence homology with the wild-type enzyme, but that it has been mutated (modified) in some way as known in the art or as described herein.
[0093] Also in this case, the terms Cas and CRISPR enzymes and CRISPR proteins and Cas proteins are generally used synonymously and, unless otherwise specified, such as by specifically referring to Cas9, will be understood to refer to the novel CRISPR effector proteins further described herein by analogy whenever mentioned in this specification. As described above, many of the residue numberings used herein refer to effector proteins from type V / VI CRISPR loci. However, it will be understood that the present invention includes many more effector proteins from other microbial species. In certain embodiments, the effector protein may be constitutively present, or inducibly present, or conditionally present, or administered, or delivered. Effector protein optimization may be used to enhance function or develop new functions, and chimeric effector proteins can be created. And as described herein, the effector protein may be modified to be used as a general nucleic acid binding protein.
[0094] Typically, in the context of a nucleic acid targeting system, formation of a nucleic acid targeting complex (comprising a guide RNA that hybridizes to a target sequence and forms a complex with one or more nucleic acid targeting effector proteins) results in cleavage of one or both DNA strands or RNA strands within or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50 base pairs, or more therefrom). As used herein, the term "one or more sequences associated with a target locus of interest" refers to sequences in the vicinity of the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50 base pairs, or more from the target sequence, where the target sequence is contained within the target locus of interest).
[0095] Examples of codon-optimized sequences are, in this case, sequences optimized for expression in eukaryotes, such as humans (i.e., optimized for expression in humans), or other eukaryotes, animals or mammals as discussed elsewhere herein; for example, see the SaCas9 human codon-optimized sequence of WO 2014 / 093622 (PCT / US2013 / 074667) as an example of a codon-optimized sequence (from the knowledge in the art and this disclosure, codon optimization of one or more coding nucleic acid molecules, particularly with respect to effector proteins such as Cpf1, is within the scope of those skilled in the art). This is preferred, but it is understood that other examples are possible, and codon optimization for host species other than humans, or for specific organs, is known. In some embodiments, the enzyme coding sequences encoding the DNA / RNA targeting Cas proteins are codon-optimized for expression in specific cells, such as eukaryotic cells. Eukaryotic cells can be from a specific organism, such as a plant or a mammal including but not limited to humans, or a non-human eukaryote, animal or mammal as discussed herein, such as a mouse, rat, rabbit, dog, livestock, or a non-human mammal or primate, or can be derived therefrom. In some embodiments, methods of modifying the germline gene identity of humans and / or methods of modifying the gene identity of animals that cause pain to them without any substantial medical benefit to humans or animals can be excluded, as well as animals obtained from such methods. Generally, codon optimization refers to a method of modifying a nucleic acid sequence to enhance expression in a target host cell by replacing at least one codon of a native sequence (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) with a codon that is used more frequently or most frequently in the genes of the host cell while maintaining the native amino acid sequence. Different species exhibit specific biases for certain codons of a particular amino acid.Codon bias (the differences in codon usage between organisms) often correlates with the translation efficiency of messenger RNA (mRNA), which in turn is thought to depend on, among other things, the properties of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The prevalence of a selected tRNA in a cell generally reflects the codons most frequently used in peptide synthesis. Thus, genes can be adjusted for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the "Codon Usage Database" at www.kazusa.orjp / codon / , and these tables can be adapted in several ways. See Nakamura, Y., et al. "Codon usage tabulated from the international DNA sequence databases: status for the year 2000" Nucl. Acids Res. 28:292 (2000). Computer algorithms are also available for codon-optimizing a particular sequence for expression in a particular host cell, such as Gene Forge (Aptagen; Jacobus, PA). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more, or all codons) in the sequence encoding the DNA / RNA targeting Cas protein correspond to the codons most frequently used for a particular amino acid. For codon usage in yeast, see the online yeast genome database available at http: / / www.yeastgenome.org / community / codon_usage.shtml, or "Codon selection in yeast", Bennetzen and Hall, J Biol Chem. 1982 Mar 25;257(6):3026-31.Regarding codon usage in plants, including algae, reference is made to "Codon usage in higher plants, green algae, and cyanobacteria", Campbell and Gowri, Plant Physiol. 1990 Jan; 92(1): 1-11; as well as "Codon usage in plant genes", Murray et al, Nucleic Acids Res. 1989 Jan 25; 17(2): 477-98; or "Selection on the codon bias of chloroplast and cyanelle genes in different plant and algal lineages", Morton BR, J Mol Evol. 1998 Apr; 46(4): 449-59.
[0096] In some embodiments, the vector encodes a nucleic acid targeting effector protein, such as AsCpf1 or an ortholog or homolog thereof, that includes one or more nuclear localization sequences (NLSs), such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. In some embodiments, the RNA targeting effector protein includes about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs near the amino terminus or at the carboxy terminus or near the carboxy terminus, or a combination thereof (e.g., zero or at least 1 or more NLSs at the amino terminus and zero or 1 or more NLSs at the carboxy terminus). When two or more NLSs are present, each may be selected independently of the others, and thus a single NLS may be present in two or more copies, and / or in combination with one or more other NLSs present in one or more copies. In some embodiments, an NLS is considered to be near the N terminus or C terminus when the closest amino acid of the NLS is within the range of about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more amino acids along the polypeptide chain from the N terminus or C terminus.Non-limiting examples of NLSs include the NLS of SV40 large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO: 2); the NLS of nucleoplasmin (e.g., the nucleoplasmin bipartite NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 3)); the c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 4) or RQRRNELKRSP (SEQ ID NO: 5); the hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 6); the sequence of the IBB domain of importin-α, RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 7); the sequences of the muscle tumor T protein, VSRKRPRP (SEQ ID NO: 8) and PPKKARED (SEQ ID NO: 9); the sequence of human p53, PQPKKKPL (SEQ ID NO: 10); the sequence of mouse c-abl IV, SALIKKKKKMAP (SEQ ID NO: 11); the sequences of influenza virus NS1, DRLRR (SEQ ID NO: 12) and PKQKKRK (SEQ ID NO: 13); the sequence of hepatitis delta antigen, RKLKKKIKKL (SEQ ID NO: 14); the sequence of mouse Mx1 protein, REKKKFLKRR (SEQ ID NO: 15); the sequence of human poly(ADP-ribose) polymerase, KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 16); and NLS sequences derived from the sequence of the steroid hormone receptor (human) glucocorticoid, RKCLQAGMNLEARKTKK (SEQ ID NO: 17). Generally, one or more NLSs are of sufficient strength to drive the accumulation of a detectable amount of DNA / RNA-targeting Cas protein in the nucleus of eukaryotic cells. Generally, the strength of nuclear localization activity can be derived from the number of NLSs within the nucleic acid-targeting effector protein, the specific NLSs used, or a combination of these factors. Detection of accumulation in the nucleus can be carried out by any suitable technique. For example, a detectable marker may be fused to the nucleic acid-targeting protein, and thereby combined with means for detecting the location of the nucleus (e.g., staining specific for the nucleus such as DAPI) to visualize the location within the cell.The cell nucleus may also be isolated from the cell and then its contents may be analyzed by any suitable protein detection method, such as immunohistochemistry, Western blot, or enzyme activity assay. The accumulation in the nucleus may also be determined indirectly by assays related to the effect of nucleic acid targeting complex formation (e.g., assays related to DNA or RNA cleavage or mutation at the target sequence, or assays related to gene expression activity that has changed in response to nucleic acid targeting complex formation and / or DNA or RNA targeting Cas protein activity), compared to a control not exposed to the nucleic acid targeting Cas protein or nucleic acid targeting complex, or a control exposed to a nucleic acid targeting Cas protein lacking one or more NLSs. In the preferred embodiments of the Cpf1 effector protein complex and system described herein, the codon-optimized Cpf1 effector protein comprises an NLS added to the C-terminus of the protein.
[0097] In some embodiments, one or more vectors driving the expression of one or more elements of a nucleic acid targeting system are introduced into a host cell, and when the elements of the nucleic acid targeting system are expressed, the formation of a nucleic acid targeting complex is induced at one or more target sites. For example, a nucleic acid targeting effector enzyme and a nucleic acid targeting guide RNA may each be operably linked to a separate regulatory element on a separate vector. One or more RNAs of the nucleic acid targeting system can be delivered, for example, by pre-administering to a transgenic nucleic acid targeting effector protein animal or mammal, such as an animal or mammal that constitutively or inducibly or conditionally expresses a nucleic acid targeting effector protein; or an animal or mammal that naturally expresses or has cells containing a nucleic acid targeting effector protein, one or more vectors that encode and express the nucleic acid targeting effector protein in vivo. Alternatively, two or more of the elements expressed from the same or different regulatory elements may be combined in a single vector, and one or more additional vectors provide any components of the nucleic acid targeting system not included in the first vector. The nucleic acid targeting system elements combined in a single vector may be arranged in any suitable orientation, such that one element is located 5' (its "upstream") or 3' (its "downstream") relative to a second element. The coding sequence of one element may be located on the same or opposite strand as the coding sequence of a second element and may be oriented in the same or opposite direction. In some embodiments, a single promoter drives the expression of transcripts encoding a nucleic acid targeting effector protein and a nucleic acid targeting guide RNA incorporated within one or more intron sequences (e.g., each in a different intron, two or more in at least one intron, or all in a single intron). In some embodiments, the nucleic acid targeting effector protein and the nucleic acid targeting guide RNA are operably linked to the same promoter and may be expressed therefrom.For delivery media, vectors, particles, nanoparticles, formulations and their components for expressing one or more elements of a nucleic acid targeting system, they are as used in the aforementioned documents, such as WO 2014 / 093622 pamphlet (Specification of PCT / US2013 / 074667). In some embodiments, the vector includes one or more insertion sites (also referred to as "cloning sites"), such as restriction endonuclease recognition sequences. In some embodiments, one or more insertion sites (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more, or more insertion sites) are located upstream and / or downstream of one or more sequence elements of one or more vectors. Using multiple different guide sequences, nucleic acid targeting activity can be targeted to multiple different corresponding target sequences in cells using a single expression construct. For example, a single vector can contain about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more, or more guide sequences. In some embodiments, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more, or more such guide sequence-containing vectors are provided and optionally delivered to cells. Multiple sgRNAs can also be expressed in an array format using an RNA polymerase III type promoter (e.g., U6 or H1 RNA). The non-coding RNA CRISPR-Cas9 components described above can be cloned into an AAV shuttle vector or there is sufficient space left to include other elements cloned into an AAV shuttle plasmid using standard methods, such as a reporter gene, an antibiotic resistance gene or other sequences. In certain embodiments, the guide RNA is provided as an array comprising guide RNAs that can be processed by an endogenous mechanism (e.g., cleaved or separated from the array). For example, Port et al. (http: / / dx.doi.org / 10.1101 / 046417) describe a system for expressing multiple guide RNAs utilizing cellular tRNA processing.More specifically, in certain embodiments, an array of guide RNA sequences each separated from its neighbor by a nucleotide sequence that can be processed (cleaved) either by a tRNA sequence or by the cell's endogenous tRNA processing system may be provided. When transcribed, this array is processed to release multiple guide RNAs, which can be used, for example, to introduce multiple changes to one or more target sequences. The guide RNAs expressed from the array may be provided in any desired combination. For example, there may be multiple copies of the same gRNA, multiple gRNAs exclusive of each other, or a combination of both. These guides can be used to direct the expression of an active Cpf1 enzyme that cleaves DNA, a modified Cpf1 enzyme such as a nickase, or other mutant Cpf1 enzymes or proteins. In certain embodiments, multiple guide RNAs are used to introduce multiple mutations into the same gene or other target DNA. In another embodiment, multiple guide RNAs are used to introduce changes into two or more genes or target DNA.
[0098] In some embodiments, the vector comprises a regulatory element operably linked to an enzyme coding sequence encoding a nucleic acid targeting effector protein. The nucleic acid targeting effector protein or one or more nucleic acid targeting guide RNAs may be delivered separately; and advantageously, at least one of these is delivered by a particle complex. To allow time for expression of the nucleic acid targeting effector protein, the nucleic acid targeting effector protein mRNA may be delivered prior to the nucleic acid targeting guide RNA. The nucleic acid targeting effector protein mRNA may be administered 1 to 12 hours (preferably about 2 to 6 hours) before administration of the nucleic acid targeting guide RNA. Alternatively, the nucleic acid targeting effector protein mRNA and the nucleic acid targeting guide RNA may be administered together. Advantageously, a second booster dose of the guide RNA may be administered 1 to 12 hours (preferably about 2 to 6 hours) after the first administration of the nucleic acid targeting effector protein mRNA + guide RNA. Additional administrations of the nucleic acid targeting effector protein mRNA and / or the guide RNA may be useful to achieve the most efficient genomic modification levels.
[0099] In one aspect, the present invention provides a method of using one or more elements of a nucleic acid targeting system. The nucleic acid targeting complex of the present invention provides an effective means of modifying a target DNA or RNA (single-stranded or double-stranded, linear or supercoiled). The nucleic acid targeting complex of the present invention has a wide range of utilities, including modification (e.g., deletion, insertion, translocation, inactivation, activation) of target DNA or RNA in a very large number of cell types. Thus, the nucleic acid targeting complex of the present invention has broad applicability, for example, in gene therapy, drug screening, disease diagnosis, and prognosis determination. Exemplary nucleic acid targeting complexes include a DNA or RNA targeting effector protein complexed with a guide RNA that hybridizes to a target sequence within a target locus of interest.
[0100] In one embodiment, the present invention provides a method for cleaving a target RNA. The method may include modifying the target RNA using a nucleic acid targeting complex that binds to the target RNA and causes cleavage of the target DNA. In certain embodiments, the nucleic acid targeting complex of the present invention can create a cleavage (e.g., single-stranded or double-stranded cleavage) in the RNA sequence when introduced into a cell. For example, the method can be used to cleave disease RNAs in a cell. For example, an exogenous RNA template in which an upstream sequence and a downstream sequence are adjacent to the sequence to be integrated may be introduced into the cell. These upstream and downstream sequences share sequence similarity with both sides of the integration site in the RNA. Optionally, the donor RNA may be mRNA. The exogenous RNA template includes the sequence to be integrated (e.g., mutant RNA). The integration sequence may be an endogenous or exogenous sequence to the cell. Examples of sequences to be integrated include RNA encoding a protein or non-coding RNA (e.g., microRNA). Thus, the integration sequence can be operably linked to one or more appropriate regulatory sequences. Alternatively, the integration sequence may provide a regulatory function. The upstream and downstream sequences in the exogenous RNA template are selected to promote recombination between the target RNA sequence and the donor RNA. The upstream sequence is an RNA sequence that shares sequence similarity with the RNA sequence upstream of the target integration site. Similarly, the downstream sequence is an RNA sequence that shares sequence similarity with the RNA sequence downstream of the target integration site. The upstream and downstream sequences in the exogenous RNA template can have 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with the target RNA sequence. Preferably, the upstream and downstream sequences in the exogenous RNA template have about 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the target RNA sequence. In some methods, the upstream and downstream sequences in the exogenous RNA template have about 99% or 100% sequence identity with the target RNA sequence.The upstream or downstream sequence can include from about 20 bp to about 2500 bp, for example, about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, or 2500 bp. In some methods, exemplary upstream or downstream sequences have from about 200 bp to about 2000 bp, from about 600 bp to about 1000 bp, or more particularly from about 700 bp to about 1000 bp. In some methods, the exogenous RNA template can further include a marker. Such markers can facilitate screening for target integration. Examples of suitable markers include restriction sites, fluorescent proteins, or selectable markers. The exogenous RNA templates of the invention can be constructed using recombinant techniques (see, e.g., Sambrook et al., 2001 and Ausubel et al., 1996). In methods of modifying target RNA by incorporating an exogenous RNA template, a cleavage (e.g., a double-stranded or single-stranded cleavage in double-stranded or single-stranded DNA or RNA) is introduced into the DNA or RNA sequence by a nucleic acid targeting complex, and when this cleavage is repaired by homologous recombination with the exogenous RNA template, the template is incorporated into the RNA target. The presence of a double-stranded cleavage promotes template integration. In other embodiments, the invention provides a method of modifying the expression of RNA in a eukaryotic cell. The method includes increasing or decreasing the expression of a target polynucleotide using a nucleic acid targeting complex that binds to DNA or RNA (e.g., mRNA or pre-mRNA). In some methods, the target RNA can be inactivated to effect a modification of cell expression. For example, when an RNA targeting complex binds to a target sequence within a cell, the target RNA is inactivated such that its sequence is not translated, the encoded protein is not produced, or the sequence does not function like the wild-type sequence. For example, a sequence encoding a protein or microRNA can be inactivated such that no protein or microRNA or pre-microRNA transcript is produced. The target RNA of the RNA targeting complex can be any RNA that is endogenous or exogenous to the eukaryotic cell.For example, the target RNA may be an RNA present in the nucleus of a eukaryotic cell. The target RNA may be a sequence encoding a gene product (e.g., a protein) (e.g., mRNA or pre-mRNA) or a non-coding sequence (e.g., ncRNA, lncRNA, tRNA, or rRNA). Examples of target RNAs include sequences related to biochemical signal transduction pathways, such as RNAs related to biochemical signal transduction pathways. Examples of target RNAs include disease-related RNAs. "Disease-related" RNA refers to any RNA that produces a translation product at an abnormal level or in an abnormal form in cells derived from diseased tissue compared to non-diseased control tissue or cells. It may be an RNA transcribed from a gene that comes to be expressed at an abnormally high level; it may be an RNA transcribed from a gene that comes to be expressed at an abnormally low level, where the change in expression correlates with the incidence and / or progression of the disease. Disease-related RNA also refers to an RNA transcribed from a gene that has one or more mutations or genetic variations that are directly involved in or in linkage disequilibrium with one or more genes involved in the etiology of the disease. The translation product may be known or unknown, and may be at a normal level or an abnormal level. The target RNA of the RNA targeting complex may be any RNA that is endogenous or exogenous to the eukaryotic cell. For example, the target RNA may be an RNA present in the nucleus of a eukaryotic cell. The target RNA may be a sequence encoding a gene product (e.g., a protein) (e.g., mRNA or pre-mRNA) or a non-coding sequence (e.g., ncRNA, lncRNA, tRNA, or rRNA).
[0101] In some embodiments, the method can include the step of binding a nucleic acid targeting complex to a target DNA or RNA to cause cleavage of the target DNA or RNA, thereby modifying the target DNA or RNA, wherein the nucleic acid targeting complex comprises a nucleic acid targeting effector protein complexed with a guide RNA that hybridizes to a target sequence within the target DNA or RNA. In one aspect, the present invention provides a method for modifying the expression of DNA or RNA in a eukaryotic cell. In some embodiments, the method includes the step of binding a nucleic acid targeting complex to a DNA or RNA, whereby the binding results in an increase or decrease in the expression of the DNA or RNA; wherein the nucleic acid targeting complex comprises a nucleic acid targeting effector protein complexed with a guide RNA. Similar considerations and conditions apply to the method of modifying a target DNA or RNA as described above. In fact, these options for sample collection, culture, and reintroduction apply to all aspects of the present invention. In one aspect, the present invention provides a method for modifying a target DNA or RNA of a eukaryotic cell, which may be in vivo, ex vivo, or in vitro. In some embodiments, the method includes the step of sampling a cell or cell population from a human or non-human animal, and the step of modifying one or more cells. The culture can be performed ex vivo at any stage. The one or more cells may further be reintroduced into a non-human animal or plant. For the cells to be reintroduced, it is particularly preferred that the cells are stem cells.
[0102] In fact, in any aspect of the present invention, the nucleic acid targeting complex can comprise a nucleic acid targeting effector protein complexed with a guide RNA that hybridizes to a target sequence.
[0103] The present invention relates to the engineering and optimization of systems, methods, and compositions for the control of gene expression involving DNA or RNA sequence targeting, related to nucleic acid targeting systems and their components. In advantageous embodiments, the effector enzyme is Cpf1, more specifically AsCpf1. An advantage of the method is that the CRISPR system minimizes or avoids off-target binding and the resulting side effects. This is achieved using a system configured to have a high degree of sequence specificity for the target DNA or RNA.
[0104] With respect to a nucleic acid targeting complex or system, preferably, the crRNA sequence has one or more stem-loops or hairpins and is 30 nucleotides or longer, 40 nucleotides or longer, or 50 nucleotides or longer; the crRNA sequence is 10-30 nucleotides long, and the nucleic acid targeting effector protein is a Cpf1 enzyme. In certain embodiments, the crRNA sequence is 42-44 nucleotides long, and the nucleic acid targeting Cas protein is Cpf1 of Francisella tularensis subsp. novocida U112. In certain embodiments, the crRNA comprises, consists essentially of, or consists of a 19-nucleotide direct repeat and a 23-25 nucleotide spacer sequence, and the nucleic acid targeting Cas protein is Cpf1 of Francisella tularensis subsp. novocida U112.
[0105] Crystallization and Structure of CRISPR-Cpf1 Crystallization and Characterization of the Crystal Structure of CRISPR-Cpf1: The crystals of the present invention can be obtained by protein crystallography techniques including batch method, liquid bridging method, dialysis method, vapor diffusion method, and hanging drop method. Generally, the crystals of the present invention are grown by dissolving substantially pure CRISPR Cpf1 and the nucleic acid molecule to which it binds in an aqueous buffer containing a precipitant at a concentration slightly lower than the concentration required for precipitation. Water is removed by evaporation controlled such that precipitation conditions occur, and this condition is maintained until crystal growth stops.
[0106] Use of Crystals, Crystal Structures, and Atomic Structure Coordinates: The crystals of the present invention, particularly the atomic structure coordinates obtained therefrom, have a wide range of uses. The crystals and structure coordinates are particularly useful for identifying compounds (nucleic acid molecules) that bind to CRJSPR-Cpf1 and CRISPR-Cpf1 that can bind to specific compounds (nucleic acid molecules). Accordingly, the structure coordinates described herein can be used as a phase model in determining the crystal structures of additional synthetic or mutant CRISPR-Cpf1, Cpf1, nickase, and binding domains. The provision of the crystal structure of CRISPR-Cpf1 complexed with a nucleic acid molecule as shown in the crystal structure tables and / or figures herein provides those skilled in the art with detailed insights into the mechanism of action of CRISPR-Cpf1. From this insight, means for designing modified CRISPR-Cpf1, such as by adding functional groups such as repressors or activators, are obtained. Functional groups such as repressors or activators can be added to the N-terminus or C-terminus of CRISPR-Cpf1, but according to what the crystal structure shows, the N-terminus appears to be covered or hidden, while the C-terminus is more accessible for functional groups such as repressors or activators. Furthermore, according to what the crystal structure shows, there is a flexible loop suitable for adding functional groups such as activators or repressors between residues approximately 534 and 676 of CRISPR-Cpf1 (Streptococcus pyogenes). The addition can be via a linker, such as a flexible glycine-serine (GlyGlyGlySer) or (GGGS)3 or a rigid α-helical linker, such as (Ala(GluAlaAlaAlaLys)Ala). In addition to the flexible loop, there are also nuclease or H3 regions, H2 regions, and helical regions. "Helix" or "helical" means a helix known in the art, including but not limited to an α-helix. In addition, the terms helix or helical may also be used to refer to a c-terminal helical element having an N-terminal turn.
[0107] Providing the crystal structure of CRISPR-Cpf1 complexed with a nucleic acid molecule enables novel drug or compound discovery, identification, and design methods for compounds that can bind to CRISPR-Cpf1. Thus, the present invention provides tools useful for the diagnosis, treatment, or prevention of pathological conditions or diseases in multicellular organisms such as algae, plants, invertebrates, fish, amphibians, reptiles, birds, mammals; for example, cultivated plants, livestock (e.g., production animals such as pigs, cows, chickens; pet animals such as cats, dogs, rodents (rabbits, mice, hamsters); experimental animals such as mice, rats), and humans. Accordingly, provided herein is a computer-based method for rationally designing a CRISPR-Cpf1 complex. This rational design includes: providing the structure of the CRISPR-Cpf1 complex as defined by some or all of the coordinates in the crystal structure tables and / or figures herein (e.g., at least 2 or more, e.g., at least 5, preferably at least 10, more preferably at least 50, even more preferably at least 100 atoms of the structure); providing the structure of the desired nucleic acid molecule with respect to which the CRISPR-Cpf1 complex is desired; and fitting the structure of the CRISPR-Cpf1 complex as defined by some or all of the coordinates in the crystal structure tables and / or figures herein to the desired nucleic acid molecule (said fitting may include achieving one or more putative modifications of the CRISPR-Cpf1 complex as defined by some or all of the coordinates in the crystal structure tables and / or figures herein such that the desired nucleic acid molecule binds to one or more CRISPR-Cpf1 complexes involved with the desired nucleic acid molecule). This method or the fitting of this method may use the atomic coordinates of interest of the CRISPR-Cpf1 complex as defined by some or all of the coordinates in the crystal structure tables and / or figures herein (e.g., at least 2 or more, e.g., at least 5, preferably at least 10, more preferably at least 50, even more preferably at least 100 atoms of the structure) in the vicinity of the active site or binding region for modeling the vicinity of the active site or binding region.These coordinates can be used to define a space, which is then “in silico” screened against a desired or candidate nucleic acid molecule. Accordingly, the present invention provides a computer-based method for rationally designing a CRISPR-Cpf1 complex. The method may include: providing the coordinates of at least two atoms (the “selected coordinates”) from the crystal structure tables herein; providing the structure of a candidate or desired nucleic acid molecule; and fitting the structure of the candidate to the selected coordinates. In this way, one skilled in the art can also fit functional groups and candidate or desired nucleic acid molecules. For example, providing the structure of the CRISPR-Cpf1 complex as defined by some or all of the coordinates in the crystal structure tables and / or figures herein (e.g., at least two or more, e.g., at least five, advantageously at least ten, more advantageously at least fifty, even more advantageously at least one hundred atoms of the structure); providing the structure of a desired nucleic acid molecule for which the CRISPR-Cpf1 complex is desired; fitting the structure of the CRISPR-Cpf1 complex as defined by some or all of the coordinates in the crystal structure tables and / or figures herein to the desired nucleic acid molecule (said fitting includes achieving one or more putative modifications of the CRISPR-Cpf1 complex as defined by some or all of the coordinates in the crystal structure tables and / or figures herein such that the desired nucleic acid molecule binds to one or more CRISPR-Cpf1 complexes involved with the desired nucleic acid molecule); selecting one or more putative fit CRISPR-Cpf1-desired nucleic acid molecule complexes; fitting one or more such putative fit CRISPR-Cpf1-desired nucleic acid molecule complexes with respect to a functional group, e.g., fitting with respect to the location (e.g., a location within a mobile loop) where a functional group (e.g., an activator, a repressor) is to be positioned and / or with respect to one or more putative modifications of one or more putative fit CRISPR-Cpf1-desired nucleic acid molecule complexes to create a location for positioning the functional group.As suggested, the present invention can be practiced using the coordinates in the crystal structure tables and / or figures herein that are near the active site or binding region; thus, the methods of the present invention can utilize the desired subdomains of the CRISPR-Cpf1 complex. The methods disclosed herein can be practiced using the coordinates of domains or subdomains. The methods can optionally include synthesizing candidate or desired nucleic acid molecules and / or CRISPR-Cpf1 systems from "in silico" outputs, and testing the binding and / or activity of "wet" or actual functional groups linked to "wet" or actual CRISPR-Cpf1 systems bound to "wet" or actual candidate or desired nucleic acid molecules. The methods can include synthesizing a CRISPR-Cpf1 system (including a functional group) from an "in silico" output, and testing the binding and / or activity of "wet" or actual functional groups linked to "wet" or actual CRISPR-Cpf1 systems bound to "wet" or actual candidate or desired nucleic acid molecules in vivo, e.g., by contacting a "wet" or actual CRISPR-Cpf1 system including a functional group from an "in silico" output with a cell including a desired or candidate nucleic acid molecule. These methods can include observing a cell or an organism including the cell for a desired reaction, e.g., alleviation of a symptom or condition or disease. The step of providing the structure of a candidate nucleic acid molecule can include selecting a compound by computationally screening a nucleic acid molecule database including nucleic acid molecule data, e.g., such data related to a condition or disease. The 3D descriptors for binding of a candidate nucleic acid molecule can be derived from geometric and functional constraints derived from the structure and chemical properties of the CRISPR-Cpf1 complex or its domains or regions from the crystal structures herein. In fact, this descriptor can be a type of one or more virtual modifications of the CRISPR-Cpf1 complex crystal structure herein for binding CRISPR-Cpf1 to a candidate or desired nucleic acid molecule. This descriptor can then be used to query a nucleic acid molecule database to identify nucleic acid molecules in the database that are predicted to have good binding to the descriptor.Next, the "wet" steps of the present specification can be carried out using a descriptor and a nucleic acid molecule that is presumed to have good binding properties.
[0108] "Fitting" can mean determining, by automated or semi-automated means, the interaction between at least one atom of a candidate and at least one atom of the CRISPR-Cpf1 complex, and calculating how stable such an interaction is. The interaction can include attractive and repulsive forces caused by, for example, charge and steric factors. "Subdomain" can mean at least one, for example, 1, 2, 3, or 4 complete secondary structure elements. Particular regions or domains of CRISPR-Cpf1 include those identified in the crystal structure tables and figures of the present specification.
[0109] In any case, determination of the three-dimensional structure of the CRISPR-Cpf1 (AsCpf1) complex enables modification of the CRISPR-Cpf1 system to bind various nucleic acid molecules, induction systems that can interact with each other and can interact with CRISPR-Cpf1 (for example, an induction system that brings about functional self-activation and / or self-termination), modification of the CRISPR-Cpf1 system to be linked to any one or more of various functional groups that can interact with nucleic acid molecules (for example, the functional group may be a regulatory or functional domain selected from the group consisting of a transcriptional repressor, a transcriptional activator, a nuclease domain, a DNA methyltransferase, a protein transferase, a protein deacetylase, a protein methyltransferase, a protein deaminase, a protein kinase, and a protein phosphatase; and in some embodiments, the functional domain is an epigenetic regulator; for example, see Zhang et al., U.S. Patent No. 8,507,272, which is also hereby incorporated by reference in its entirety and all cited references and all application cited references herein), modification of Cpf1, design of novel and specific nucleic acid molecules that bind to CRISPR-Cpf1 (for example, AsCpf1) such as by novel nickases, and design of novel CRISPR-Cpf1 systems are provided. Indeed, the CRISPR-Cpf1 (AsCpf1) crystal structure according to the present specification has various uses. For example, from the information on the three-dimensional structure of the CRISPR-Cpf1 (AsCpf1) crystal structure, various molecules or other structural or functional features of the CRISPR-Cpf1 system (for example, AsCpf1) that are expected to interact with possible or confirmed sites such as binding sites can be designed or identified using a computer modeling program. Potentially binding compounds ("binders") can be investigated using computer modeling with a docking program.Docking programs are known; see, for example, GRAM, DOCK or AUTODOCK (see Walters et al. Drug Discovery Today, vol. 3, no. 4 (1998), 160-178, and Dunbrack et al. Folding and Design 2 (1997), 27-42). This procedure can include determining how well the shape and chemical structure of a potential conjugate binds to a CRISPR-Cpf1 system (e.g., AsCpf1) by computer fitting of the potential conjugate. Manual examination of the active or binding site of a CRISPR-Cpf1 system (e.g., AsCpf1) can be performed with computer assistance. Programs such as GRID (P. Goodford, J. Med. Chem, 1985, 28, 849-57) - a program that determines likely interaction sites between molecules with various functional groups - can also be used for analysis of the active or binding site to predict the partial structure of a binding compound. Using a computer program, the attractive, repulsive or steric hindrance between two binding partners, e.g., a CRISPR-Cpf1 system (e.g., AsCpf1) and a candidate nucleic acid molecule or a nucleic acid molecule and a candidate CRISPR-Cpf1 system (e.g., AsCpf1), can be estimated; and according to the CRISPR-Cpf1 crystal structure (AsCpf1) herein, such a method is possible. Generally, the closer the fit, the less steric hindrance and the greater the attractive force, resulting in a more potent potential conjugate, because these properties correlate with a tighter binding constant. Further, the higher the specificity in the design of a candidate CRISPR-Cpf1 system (e.g., AsCpf1), the lower the likelihood of also interacting with off-target molecules. Also, "wet" methods are made possible by the present application.For example, in one aspect, the present specification provides a method for determining the structure of a conjugate of a candidate CRISPR-Cpf1 system (e.g., AsCpf1) bound to a target nucleic acid molecule, the method comprising: (a) providing a first crystal of a candidate CRISPR-Cpf1 system (AsCpf1) as described herein or a second crystal of a candidate CRISPR-Cpf1 system (e.g., AsCpf1); (b) contacting the first crystal or the second crystal with the conjugate under conditions under which a complex can be formed; and (c) determining the structure of the candidate (e.g., the CRISPR-Cpf1 system (e.g., AsCpf1) or the CRISPR-Cpf1 system (AsCpf1) conjugate). The second crystal may have essentially the same coordinates as those discussed herein, but due to minor changes in the CRISPR-Cpf1 system, this crystal may be formed in a different space group.
[0110] Furthermore, the present specification provides other "wet" methods, including high-throughput screening of conjugates (e.g., target nucleic acid molecules) with candidate CRISPR-Cpf1 systems (e.g., AsCpf1), or candidate conjugates (e.g., target nucleic acid molecules) with CRISPR-Cpf1 systems (e.g., AsCpf1), or candidate conjugates (e.g., target nucleic acid molecules) with candidate CRISPR-Cpf1 systems (e.g., AsCpf1) (one or more of the above CRISPR-Cpf1 systems may or may not contain one or more functional groups), instead of or in addition to the "in silico" method. A pair of a conjugate showing binding activity and a CRISPR-Cpf1 system can be selected and further crystallized with a CRISPR-Cpf1 crystal having the structure of the present specification, for X-ray analysis, e.g., by co-crystallization or soaking. The obtained X-ray structure can be compared with those in the crystal structure table of the present specification and the information in the figures for various purposes, e.g., regarding the overlap range. When designing, identifying, or selecting a potential pair of a conjugate and a CRISPR-Cpf1 system by determining preferred fitting properties, e.g., those having strong predicted attraction, based on the pair of the conjugate and the CRISPR-Cpf1 crystal structure data of the present specification, these potential pairs can then be screened by "wet" methods for activity. As a result, in one aspect, the method can include: obtaining or synthesizing potential pairs; and contacting a conjugate (e.g., target nucleic acid molecule) with a candidate CRISPR-Cpf1 system (e.g., AsCpf1), or a candidate conjugate (e.g., target nucleic acid molecule) with a CRISPR-Cpf1 system (e.g., AsCpf1), or a candidate conjugate (e.g., target nucleic acid molecule) with a candidate CRISPR-Cpf1 system (e.g., AsCpf1) (one or more of the above CRISPR-Cpf1 systems may or may not contain one or more functional groups) to determine the binding ability. In the latter step, the contact is preferably under conditions for determining the function. Instead of or in addition to performing such an assay, the method can include: obtaining or synthesizing one or more complexes from the contact, and analyzing the one or more complexes by, e.g., X-ray diffraction or NMR or other means to determine the binding or interaction ability.Next, detailed structural information regarding the binding may be obtained, and based on that information, the structure or function of the candidate CRISPR-Cpf1 system or its components can be adjusted. These steps can be repeated and re-repeated as necessary. Alternatively or in addition, potential CRISPR-Cpf1 systems from or in the above-described methods can be used to confirm or demonstrate functionality, including, but not limited to, whether the desired result (e.g., symptom reduction, treatment) is obtained thereby, and can be combined with nucleic acid molecules in vivo, including by administration to an organism (including non-human animals and humans).
[0111] Furthermore, the present specification provides a method for determining the three-dimensional structure of one or more CRISPR-Cpf1 systems or complexes of unknown structure by using the structural coordinates in the crystal structure table of the present specification and the information in the figures. For example, when X-ray crystallographic data or NMR spectroscopic data is provided for a CRISPR system or complex of unknown crystal structure, by interpreting the data using the structure of the CRISPR-Cpf1 complex as defined in the crystal structure table and figures of the present specification, a putative structure of the unknown system or complex can be provided by techniques such as phase modeling in the case of X-ray crystallography. Thus, the method can include: aligning the representation of a CRISPR-cas system or complex having an unknown crystal structure so that homologous or similar regions (e.g., homologous or similar sequences) match the similar representations of the CRISPR-Cpf1 systems and complexes of the crystal structure of the present specification; modeling the structure of the aligned homologous or similar regions (e.g., sequences) of the CRISPR-cas system or complex of unknown crystal structure based on the structure as defined in the crystal structure table and / or figures of the present specification for the corresponding regions (e.g., sequences); and determining a conformation that substantially maintains the structure of the aligned homologous regions for the unknown crystal structure (e.g., taking into account that favorable interactions must be formed so that a low-energy conformation is formed). "Homologous regions" represent, for example with respect to amino acids, amino acid residues in two sequences having the same or similar, e.g., aliphatic, aromatic, polar, negatively charged, or positively charged, side-chain chemical groups. Homologous regions for nucleic acid molecules can include at least 85%, or 86%, or 87%, or 88%, or 89%, or 90%, or 91%, or 92%, or 93%, or 94%, or 95%), or 96%, or 97%, or 98%, or 99% homology or identity. Identical regions and similar regions may sometimes be described by those skilled in the art as "invariant" and "conserved", respectively. Advantageously, the first and third steps are carried out by computer modeling. Homology modeling is a technique well known to those skilled in the art (see, e.g., Greer, Science vol. 228 (1985) 1055, and Blundell et al. Eur J Biochem vol 172 (1988), 513).The computer representation of the conserved regions of the CRISPR-Cpf1 crystal structure and the computer representation of the CRISPR-Cas system of unknown crystal structure in this specification are useful for predicting and determining the crystal structure of the CRISPR-Cas system of unknown crystal structure. Furthermore, the embodiments described herein that utilize the CRISPR-Cpf1 crystal structure in silico can equally apply to novel CRISPR-Cas crystal structures inferred by using the CRISPR-Cpf1 crystal structure of this specification. Thus, a library of CRISPR-Cas crystal structures can be obtained. Accordingly, a rational CRISPR-Cas system design is provided herein. For example, when the conformation or crystal structure of a CRISPR-Cas system or complex is determined by the method described herein, such conformation can be used in the computer-based method herein to determine the conformation or crystal structure of other CRISPR-Cas systems or complexes whose crystal structures are still unknown. The data from all these crystal structures may be in a database, and the method herein can be made more robust by having a comparison herein involving the crystal structure of this specification or a part thereof compared to one or more crystal structures in the library. The present invention further includes a system such as a computer system intended for performing structure generation and / or rational design of a CRISPR-Cas system or complex. This system may include: atomic coordinate data related to or derived from, for example, by modeling, the crystal structure tables and figures herein, which define the three-dimensional structure of a CRISPR-Cas system or complex or at least one domain or subdomain thereof, or structure factor data that can be derived from the atomic coordinate data of the crystal structure tables and figures herein. This specification also provides a computer-readable medium including atomic coordinate data related to or derived from, for example, by homology modeling, the crystal structure tables and / or figures herein, which define the three-dimensional structure of a CRISPR-Cas system or complex or at least one domain or subdomain thereof, or structure factor data that can be derived from the atomic coordinate data of the crystal structure tables and / or figures herein."Computer-readable medium" refers to any medium that can be directly read and accessed by a computer, including but not limited to: magnetic storage media; optical storage media; electrical storage media; cloud storage and hybrids of these categories. By providing such a computer-readable medium, routine access to atomic coordinate data can be made available for modeling or other "in silico" methods. Further, this specification includes methods of conducting business by providing access to such a computer-readable medium, for example, via the Internet or a global communication / computer network in a membership registration manner; or the computer system may be made available to users in a membership registration manner. "Computer system" refers to the hardware means, software means, and data storage means used in the analysis of the atomic coordinate data of the present invention. The minimum hardware means of the computer-based system of the present invention may include a central processing unit (CPU), input means, output means, and data storage means. Desirably, a display or monitor is provided for visualizing the structural data. Further, this specification includes methods of transmitting any information obtained by any method or steps described herein or any information described herein, for example, by remote communication, telephone, mass communication, mass media, presentation, Internet, e-mail, etc. The crystal structures described herein can be analyzed to create one or more Fourier electron density maps of the CRISPR-cas system or complex; advantageously, the three-dimensional structure is as defined by the atomic coordinate data according to the crystal structure tables and / or figures herein. The Fourier electron density maps can be calculated based on the X-ray diffraction pattern. These density maps can then be used to determine various aspects of binding or other interactions. The electron density maps can be calculated using known programs such as those of the CCP4 computer package (Collaborative Computing Project, No. 4. The CCP4 Suite: Programs for Protein Crystallography, Acta Crystallographica, D50, 1994, 760-763).For the visualization of density maps and model construction, programs such as "QUANTA" (1994, San Diego, Calif.: Molecular Simulations, Jones et al., Acta Crystallography A47(1991), 110-119) can be used.
[0112] The crystal structure table in this specification provides the atomic coordinate data of CRISPR-Cpf1 (Acidaminococcus), and each atom is enumerated by a unique number; the chemical elements and their positions of each amino acid residue (as determined by electron density maps and antibody sequence comparison), the amino acid residues where the elements are located, the identification name of the chain, the number of residues, the atomic positions (in angstroms) of each atom, the coordinates (e.g., X, Y, Z) defined with respect to the crystal axes, the occupancy of the atoms at each position, "B", the isotropic displacement parameter (in angstroms) explaining the movement of the atoms around the atomic center, and the atomic number. See also the text and figures in this specification.
[0113] In a further aspect, the present invention provides a method, which may be computer-assisted, for identifying or designing potential compounds that fit within or bind to a CRISPR-Cpf1 system or a portion thereof, the method comprising: a) providing the coordinates of at least two atoms of the CRISPR-Cpf1 system in a crystal structure table; b) providing the structure of a candidate molecule for binding to or within the CRISPR-Cas9 system or for manipulating a portion of the CRISPR-Cas9 system; c) fitting the structure of the candidate molecule to at least two atoms of the CRISPR-Cas9 system, the fitting including determining the interactions between one or more atoms of the candidate molecule and the atoms of the CRISPR-SpCas9 system; and d) selecting the candidate molecule if it is predicted to bind to or within the CRISPR-Cas9 system. In certain embodiments of the method, Cpf1 in the crystal structure table further includes an amino acid substitution of aspartic acid at position 908. In certain embodiments, the candidate molecule includes atoms of the CRISPR-Cpf1 system in the crystal structure table. In one embodiment, the candidate molecule includes atoms of a crRNA:DNA heteroduplex, which includes comparing the atoms of the crRNA:DNA heteroduplex to the atoms of Cpf1. In one embodiment, the atoms of Cpf1 include atoms of the REC lobe and / or atoms of the NUC lobe. In one embodiment, the atoms of Cpf1 include atoms of the REC1 domain, atoms of the REC2 domain, and / or atoms of the RuvC domain. In one embodiment, the candidate molecule includes atoms of the PAM distal region of the crRNA:DNA heteroduplex, which includes comparing the atoms of the PAM distal region of the crRNA:DNA heteroduplex to the atoms of the REC1-REC2 domain. In one embodiment, the candidate molecule includes atoms of the PAM proximal region of the crRNA:DNA heteroduplex, which includes comparing the atoms of the PAM proximal region of the crRNA:DNA heteroduplex to the atoms of the WED-REC1-RuvC domain. In certain non-limiting embodiments, the atoms of Cpf1 include the atoms of R176, R192, G783, and / or R951.
[0114] In certain embodiments, the candidate molecule comprises atoms of the PAM duplex, which are compared to the atoms of the groove formed by the WED-REC and PI domains. In certain non-limiting embodiments, the candidate molecule comprises atoms of the PAM, which are compared to the atoms of Thr167, Lys607, Lys548, Pro599, and / or Met604 of Cpf1.
[0115] In certain embodiments, the candidate molecule comprises atoms of the target DNA strand and / or the non-target DNA strand, which comprises comparing the atoms of the target DNA strand and / or the non-target DNA strand to the atoms of Cpf1. In certain embodiments where the candidate molecule comprises atoms of the target DNA strand, the atoms of the target DNA strand are compared to the atoms of the Cpf1 Nuc domain. In certain embodiments where the candidate molecule comprises atoms of the target DNA strand, the atoms of the target DNA strand are compared to the atoms of Arg1226, Ser1228, and / or Asp1235 of Cpf1. In certain embodiments where the candidate molecule comprises atoms of the non-target DNA strand, the atoms of the non-target DNA strand are compared to the atoms of the Cpf1 RuvC domain. In certain embodiments where the candidate molecule comprises atoms of the non-target DNA strand, the atoms of the non-target DNA strand are compared to the atoms of Asp908, Trp958, Glu993, and / or Asp1263 of Cpf1. In certain such embodiments, the atoms of Leu467, Leu471, Tyr514, Arg518, Ala521 and / or Thr522 are also compared.
[0116] In certain embodiments, the candidate molecule comprises atoms of the protospacer adjacent motif (PAM), which are compared to the atoms of the PAM interaction (PI) domain of Cpf1.
[0117] In certain embodiments, the candidate molecule comprises atoms of the 5'-handle of the crRNA, which are compared to the atoms of the WED domain and / or the atoms of the RuvC domain.
[0118] In certain embodiments of the invention, the candidate molecule is synthesized and tested for binding or activity.
[0119] In certain embodiments, candidate molecules are tested in the CRISPR-Cpf1 system for changes in the expression of DNA molecules in cells.
[0120] In certain embodiments, the comparison or fitting of the structure of a candidate molecule involves atomic coordinates that include at least 2 atoms, or at least 5 atoms, or at least 10 atoms, or at least 50 atoms, or at least 100 atoms of the CRISPR-Cpf1 complex.
[0121] In certain embodiments of the invention, a candidate molecule comprises atoms of Cpf1 and a transcriptional repressor, transcriptional activator, nuclease domain, DNA methyltransferase, protein acetyltransferase, protein deacetylase, protein methyltransferase, protein deaminase, protein kinase, protein phosphatase, or epigenetic regulator.
[0122] In a further aspect, the invention includes a computer-aided method for identifying or designing a potential compound to fit or bind to a CRISPR-Cpf1 system or a functional portion thereof, or the reverse (a computer-aided method for identifying or designing a potential CRISPR-Cpf1 system or a functional portion thereof to bind to a desired compound), or for identifying or designing a potential CRISPR-Cpf1 system (e.g., in relation to the operable range prediction of the CRISPR-Cpf1 system - for example, based on crystal structure data or data of Cpf1 orthologs - or in relation to where a functional group such as an activator or repressor can be added to the CRISPR-Cpf1 system, or in relation to Cpf1 truncation or nickase design), the method comprising:
[0123] The following steps, using a computer system comprising a processor, a data storage system, an input device, and an output device, such as a programmed computer: (a) For example, based on differences between Cpf1 orthologs in or instead of or in addition to the CRISPR-Cpf1 binding domain, three-dimensional coordinates of a subset of atoms from or related to a CRISPR-Cpf1 crystal structure, such as the CRISPR-Cpf1 crystal structure of Example 3 ("Crystal Structure Table") in different domains, are input using the input device into a programmed computer, optionally together with structural information from one or more CRISPR-Cpf1 complexes, with respect to Cpf1 or with respect to nickase or with respect to a functional group, thereby creating a data set; (b) Using the processor, comparing the data set with a computer database of structures stored in the computer data storage system, such as the structure of a compound that binds to or is predicted to bind to or is desired to bind to a CRISPR-Cpf1 system, or with respect to Cpf1 orthologs (e.g., as Cpf1 or with respect to different domains or regions between Cpf1 orthologs) or with respect to a CRISPR-Cpf1 crystal structure such as the CRISPR-Cpf1 crystal structure of Example 3 ("Crystal Structure Table") or with respect to nickase or with respect to a functional group; (c) Using a computer method, selecting from the database one or more structures - for example, a CRISPR-Cpf1 structure that can bind to a desired structure, a desired structure that can bind to a specific CRISPR-Cpf1 structure, a portion of a CRISPR-Cpf1 system that can be manipulated based on data from other parts of the CRISPR-Cpf1 crystal structure and / or data from Cpf1 orthologs, a truncated Cpf1, a novel nickase or a specific functional group, or the position at which a functional group is added or a functional group-CRISPR-Cpf1 system; (d) Using a computer method, constructing a model of the selected one or more structures; and (e) Outputting the selected one or more structures to the output device; And optionally, synthesizing one or more of the selected structures; and further optionally including the step of testing one or more structures of said synthesized selection as or in a CRISPR-Cpf1 system; or the method provides at least two atoms of a CRISPR-Cpf1 crystal structure, such as the CRISPR-Cpf1 crystal structure of Example 3 (“Crystal Structure Table”), for example, the coordinates of at least two atoms of the crystal structure table of the CRISPR-Cpf1 crystal structure or the coordinates of at least a subdomain of the CRISPR-Cpf1 crystal structure (“coordinates of the selection”), the step of providing a structure of a candidate including a binding molecule or a structure of a portion of a CRISPR-Cpf1 system that can be engineered based on data from other parts of the CRISPR-Cpf1 crystal structure, such as, for example, and / or data from a Cpf1 ortholog, or the structure of a functional group, and fitting the candidate structure to the coordinates of the selection, thereby obtaining product data including a CRISPR-Cpf1 structure that can bind to a desired structure, a desired structure that can bind to a specific CRISPR-Cpf1 structure, a portion of a CRISPR-Cpf1 system that can be engineered, a truncated Cpf1, a novel nickase, or a specific functional group, or the position where a functional group is added or a functional group-CRISPR-Cpf1 system, together with its output; and optionally including the step of synthesizing one or more compounds from said product data, and further optionally including the step of testing said synthesized one or more compounds as or in a CRISPR-Cpf1 system.
[0124] The step of testing may include analyzing a CRISPR-Cpf1 system obtained from one or more structures of said synthesized selection, for example, with respect to binding or with respect to performing a desired function.
[0125] The output of the foregoing method may include the transmission of information by means of data transmission, such as remote communication, telephone, videoconference, mass communication, such as presentation by computer presentation (e.g., POWERPOINT), Internet, e-mail, document by computer program (e.g., WORD) document, etc. Accordingly, the present invention also relates to atomic coordinate data related to a crystal structure referred to herein, such as the CRISPR-Cpf1 crystal structure of Example 3 ("Crystal Structure Table"), which defines the three-dimensional structure of CRISPR-Cpf1 or at least one of its subdomains, or structure factor data of CRISPR-Cpf1, which can be derived from the atomic coordinate data of the crystal structure referred to herein, such as the CRISPR-Cpf1 crystal structure of Example 3 ("Crystal Structure Table"), and also includes a computer-readable medium. The computer-readable medium may also include any data of the foregoing method. The present invention further relates to a method for generating or implementing a rational design as in the foregoing method, including atomic coordinate data related to a crystal structure referred to herein, such as the CRISPR-Cpf1 crystal structure of Example 3 ("Crystal Structure Table"), which defines the three-dimensional structure of CRISPR-Cpf1 or at least one of its subdomains, or structure factor data of CRISPR-Cpf1, which can be derived from the atomic coordinate data of the crystal structure referred to herein, such as the CRISPR-Cpf1 crystal structure of Example 3 ("Crystal Structure Table"), and also includes a computer system. The present invention further relates to a method of conducting a business, including the step of providing a user with a computer system or medium or the three-dimensional structure of CRISPR-Cpf1 or at least one of its subdomains, or structure factor data of CRISPR-Cpf1, which defines the structure shown in the atomic coordinate data of the crystal structure referred to herein, such as the CRISPR-Cpf1 crystal structure of Example 3 ("Crystal Structure Table"), and structure factor data that can be derived therefrom, or the computer medium herein or the data transmission herein.A further aspect provides a CRISPR-Cpf1 system having the crystal structure of Example 3 (the “Crystal Structure Table”) and / or having an X-ray diffraction pattern corresponding to and / or obtained from some or all of the foregoing, and / or a crystal having a structure defined by at least 2, at least 50, at least 100 or all of the coordinates of the following Crystal Structure Table. JPEG0007709487000001.jpg214166JPEG0007709487000002.jpg214166JPEG0007709487000003.jpg214166JPEG0007709487000004.jpg214166JPEG0007709487000005.jpg214166JPEG0007709487000006.jpg214166JPEG0007709487000007.jpg214166JPEG0007709487000008.jpg214166JPEG0007709487000009.jpg214166JPEG0007709487000010.jpg214166JPEG0007709487000011.jpg214166JPEG0007709487000012.jpg214166JPEG0007709487000013.jpg214166JPEG0007709487000014.jpg214166JPEG0007709487000015.jpg214166JPEG0007709487000016.jpg214166JPEG0007709487000017.jpg214166JPEG0007709487000018.jpg214166JPEG0007709487000019.jpg214166JPEG0007709487000020.jpg214166JPEG0007709487000021.jpg214166JPEG0007709487000022.jpg214166JPEG0007709487000023.jpg214166JPEG0007709487000024.jpg214166JPEG0007709487000025.jpg214166JPEG0007709487000026.jpg214166JPEG0007709487000027.jpg214166JPEG0007709487000028.jpg214166JPEG0007709487000029.jpg214166JPEG0007709487000030.jpg214166JPEG0007709487000031.jpg214166JPEG0007709487000032.jpg214166JPEG0007709487000033.jpg214166JPEG0007709487000034.jpg214166JPEG0007709487000035.jpg214166JPEG0007709487000036.jpg214166JPEG0007709487000037.jpg214166JPEG0007709487000038.jpg214166JPEG0007709487000039.jpg214166JPEG0007709487000040.jpg214166JPEG0007709487000041.jpg214166JPEG0007709487000042.jpg214166JPEG0007709487000043.jpg214166JPEG0007709487000044.jpg214166JPEG0007709487000045.jpg214166JPEG0007709487000046.jpg214166JPEG0007709487000047.jpg214166JPEG0007709487000048.jpg214166JPEG0007709487000049.jpg214166JPEG0007709487000050.jpg214166JPEG0007709487000051.jpg214166JPEG0007709487000052.jpg214166JPEG0007709487000053.jpg214166JPEG0007709487000054.jpg214166JPEG0007709487000055.jpg214166JPEG0007709487000056.jpg214166JPEG0007709487000057.jpg214166JPEG0007709487000058.jpg214166JPEG0007709487000059.jpg214166JPEG0007709487000060.jpg214166JPEG0007709487000061.jpg214166JPEG0007709487000062.jpg214166JPEG0007709487000063.jpg214166JPEG0007709487000064.jpg214166JPEG0007709487000065.jpg214166JPEG0007709487000066.jpg214166JPEG0007709487000067.jpg214166JPEG0007709487000068.jpg214166JPEG0007709487000069.jpg214166JPEG0007709487000070.jpg214166JPEG0007709487000071.jpg214166JPEG0007709487000072.jpg214166JPEG0007709487000073.jpg214166JPEG0007709487000074.jpg214166JPEG0007709487000075.jpg214166JPEG0007709487000076.jpg214166JPEG0007709487000077.jpg214166JPEG0007709487000078.jpg214166JPEG0007709487000079.jpg214166JPEG0007709487000080.jpg214166JPEG0007709487000081.jpg214166JPEG0007709487000082.jpg214166JPEG0007709487000083.jpg214166JPEG0007709487000084.jpg214166JPEG0007709487000085.jpg214166JPEG0007709487000086.jpg214166JPEG0007709487000087.jpg214166JPEG0007709487000088.jpg214166JPEG0007709487000089.jpg214166JPEG0007709487000090.jpg214166JPEG0007709487000091.jpg214166JPEG0007709487000092.jpg214166JPEG0007709487000093.jpg214166JPEG0007709487000094.jpg214166JPEG0007709487000095.jpg214166JPEG0007709487000096.jpg214166JPEG0007709487000097.jpg214166JPEG0007709487000098.jpg214166JPEG0007709487000099.jpg214166JPEG0007709487000100.jpg214166JPEG0007709487000101.jpg214166JPEG0007709487000102.jpg214166JPEG0007709487000103.jpg214166JPEG0007709487000104.jpg214166JPEG0007709487000105.jpg214166JPEG0007709487000106.jpg214166JPEG0007709487000107.jpg214166JPEG0007709487000108.jpg214166JPEG0007709487000109.jpg214166JPEG0007709487000110.jpg214166JPEG0007709487000111.jpg214166JPEG0007709487000112.jpg214166JPEG0007709487000113.jpg214166JPEG0007709487000114.jpg214166JPEG0007709487000115.jpg214166JPEG0007709487000116.jpg214166JPEG0007709487000117.jpg214166JPEG0007709487000118.jpg214166JPEG0007709487000119.jpg214166JPEG0007709487000120.jpg214166JPEG0007709487000121.jpg214166JPEG0007709487000122.jpg214166JPEG0007709487000123.jpg214166JPEG0007709487000124.jpg214166JPEG0007709487000125.jpg214166JPEG0007709487000126.jpg214166JPEG0007709487000127.jpg214166JPEG0007709487000128.jpg214166JPEG0007709487000129.jpg214166JPEG0007709487000130.jpg214166JPEG0007709487000131.jpg214166JPEG0007709487000132.jpg214166JPEG0007709487000133.jpg214166JPEG0007709487000134.jpg214166JPEG0007709487000135.jpg214166JPEG0007709487000136.jpg214166JPEG0007709487000137.jpg214166JPEG0007709487000138.jpg214166JPEG0007709487000139.jpg214166JPEG0007709487000140.jpg214166JPEG0007709487000141.jpg214166JPEG0007709487000142.jpg214166JPEG0007709487000143.jpg214166JPEG0007709487000144.jpg214166JPEG0007709487000145.jpg214166JPEG0007709487000146.jpg214166JPEG0007709487000147.jpg214166JPEG0007709487000148.jpg214166JPEG0007709487000149.jpg214166JPEG0007709487000150.jpg214166JPEG0007709487000151.jpg214166JPEG0007709487000152.jpg214166JPEG0007709487000153.jpg214166JPEG0007709487000154.jpg214166JPEG0007709487000155.jpg214166JPEG0007709487000156.jpg214166JPEG0007709487000157.jpg214166JPEG0007709487000158.jpg214166JPEG0007709487000159.jpg214166JPEG0007709487000160.jpg214166JPEG0007709487000161.jpg214166JPEG0007709487000162.jpg214166JPEG0007709487000163.jpg214166JPEG0007709487000164.jpg214166JPEG0007709487000165.jpg214166JPEG0007709487000166.jpg214166JPEG0007709487000167.jpg214166JPEG0007709487000168.jpg214166JPEG0007709487000169.jpg214166JPEG0007709487000170.jpg214166JPEG0007709487000171.jpg214166JPEG0007709487000172.jpg214166JPEG0007709487000173.jpg214166JPEG0007709487000174.jpg214166JPEG0007709487000175.jpg214166JPEG0007709487000176.jpg214166JPEG0007709487000177.jpg214166JPEG0007709487000178.jpg214166JPEG0007709487000179.jpg214166JPEG0007709487000180.jpg214166JPEG0007709487000181.jpg214166JPEG0007709487000182.jpg214166JPEG0007709487000183.jpg214166JPEG0007709487000184.jpg214166JPEG0007709487000185.jpg214166JPEG0007709487000186.jpg214166JPEG0007709487000187.jpg214166JPEG0007709487000188.jpg214166JPEG0007709487000189.jpg214166JPEG0007709487000190.jpg214166JPEG0007709487000191.jpg214166JPEG0007709487000192.jpg214166JPEG0007709487000193.jpg214166JPEG0007709487000194.jpg214166JPEG0007709487000195.jpg214166JPEG0007709487000196.jpg214166JPEG0007709487000197.jpg214166JPEG0007709487000198.jpg214166JPEG0007709487000199.jpg214166JPEG0007709487000200.jpg214166JPEG0007709487000201.jpg214166JPEG0007709487000202.jpg214166JPEG0007709487000203.jpg214166JPEG0007709487000204.jpg214166JPEG0007709487000205.jpg214166JPEG0007709487000206.jpg214166JPEG0007709487000207.jpg214166JPEG0007709487000208.jpg214166JPEG0007709487000209.jpg214166JPEG0007709487000210.jpg214166JPEG0007709487000211.jpg214166JPEG0007709487000212.jpg214166JPEG0007709487000213.jpg214166JPEG0007709487000214.jpg214166JPEG0007709487000215.jpg214166JPEG0007709487000216.jpg214166JPEG0007709487000217.jpg214166JPEG0007709487000218.jpg214166JPEG0007709487000219.jpg214166JPEG0007709487000220.jpg214166JPEG0007709487000221.jpg214166JPEG0007709487000222.jpg214166JPEG0007709487000223.jpg214166JPEG0007709487000224.jpg214166JPEG0007709487000225.jpg214166JPEG0007709487000226.jpg214166JPEG0007709487000227.jpg214166JPEG0007709487000228.jpg214166JPEG0007709487000229.jpg214166JPEG0007709487000230.jpg214166JPEG0007709487000231.jpg214166JPEG0007709487000232.jpg214166JPEG0007709487000233.jpg214166JPEG0007709487000234.jpg214166JPEG0007709487000235.jpg214166JPEG0007709487000236.jpg214166JPEG0007709487000237.jpg214166JPEG0007709487000238.jpg214166JPEG0007709487000239.jpg214166JPEG0007709487000240.jpg214166JPEG0007709487000241.jpg214166JPEG0007709487000242.jpg214166JPEG0007709487000243.jpg214166JPEG0007709487000244.jpg214166JPEG0007709487000245.jpg214166JPEG0007709487000246.jpg214166JPEG0007709487000247.jpg214166JPEG0007709487000248.jpg214166JPEG0007709487000249.jpg214166JPEG0007709487000250.jpg214166JPEG0007709487000251.jpg214166JPEG0007709487000252.jpg214166JPEG0007709487000253.jpg214166JPEG0007709487000254.jpg214166JPEG0007709487000255.jpg214166JPEG0007709487000256.jpg214166JPEG0007709487000257.jpg214166JPEG0007709487000258.jpg214166JPEG0007709487000259.jpg214166JPEG0007709487000260.jpg214166JPEG0007709487000261.jpg214166JPEG0007709487000262.jpg214166JPEG0007709487000263.jpg214166JPEG0007709487000264.jpg214166JPEG0007709487000265.jpg214166JPEG0007709487000266.jpg214166JPEG0007709487000267.jpg214166JPEG0007709487000268.jpg214166JPEG0007709487000269.jpg214166JPEG0007709487000270.jpg214166JPEG0007709487000271.jpg214166JPEG0007709487000272.jpg214166JPEG0007709487000273.jpg214166JPEG0007709487000274.jpg214166JPEG0007709487000275.jpg214166JPEG0007709487000276.jpg214166JPEG0007709487000277.jpg214166JPEG0007709487000278.jpg214166JPEG0007709487000279.jpg214166JPEG0007709487000280.jpg214166JPEG0007709487000281.jpg214166JPEG0007709487000282.jpg214166JPEG0007709487000283.jpg214166JPEG0007709487000284.jpg214166JPEG0007709487000285.jpg214166JPEG0007709487000286.jpg214166JPEG0007709487000287.jpg214166JPEG0007709487000288.jpg214166JPEG0007709487000289.jpg214166JPEG0007709487000290.jpg214166JPEG0007709487000291.jpg214166JPEG0007709487000292.jpg214166JPEG0007709487000293.jpg214166JPEG0007709487000294.jpg214166JPEG0007709487000295.jpg214166JPEG0007709487000296.jpg214166JPEG0007709487000297.jpg214166JPEG0007709487000298.jpg214166JPEG0007709487000299.jpg214166JPEG0007709487000300.jpg214166JPEG0007709487000301.jpg214166JPEG0007709487000302.jpg214166JPEG0007709487000303.jpg214166JPEG0007709487000304.jpg214166JPEG0007709487000305.jpg214166JPEG0007709487000306.jpg214166JPEG0007709487000307.jpg214166JPEG0007709487000308.jpg214166JPEG0007709487000309.jpg214166JPEG0007709487000310.jpg214166JPEG0007709487000311.jpg214166JPEG0007709487000312.jpg214166JPEG0007709487000313.jpg214166JPEG0007709487000314.jpg214166JPEG0007709487000315.jpg214166JPEG0007709487000316.jpg214166JPEG0007709487000317.jpg214166JPEG0007709487000318.jpg214166JPEG0007709487000319.jpg214166JPEG0007709487000320.jpg214166JPEG0007709487000321.jpg214166JPEG0007709487000322.jpg214166JPEG0007709487000323.jpg214166JPEG0007709487000324.jpg214166JPEG0007709487000325.jpg214166JPEG0007709487000326.jpg214166JPEG0007709487000327.jpg214166JPEG0007709487000328.jpg214166JPEG0007709487000329.jpg214166JPEG0007709487000330.jpg214166JPEG0007709487000331.jpg214166JPEG0007709487000332.jpg214166JPEG0007709487000333.jpg214166JPEG0007709487000334.jpg214166JPEG0007709487000335.jpg214166JPEG0007709487000336.jpg214166JPEG0007709487000337.jpg214166JPEG0007709487000338.jpg214166JPEG0007709487000339.jpg214166JPEG0007709487000340.jpg214166JPEG0007709487000341.jpg214166JPEG0007709487000342.jpg214166JPEG0007709487000343.jpg214166JPEG0007709487000344.jpg214166JPEG0007709487000345.jpg214166JPEG0007709487000346.jpg214166JPEG0007709487000347.jpg214166JPEG0007709487000348.jpg214166JPEG0007709487000349.jpg214166JPEG0007709487000350.jpg214166JPEG0007709487000351.jpg214166JPEG0007709487000352.jpg214166JPEG0007709487000353.jpg214166JPEG0007709487000354.jpg214166JPEG0007709487000355.jpg214166JPEG0007709487000356.jpg214166JPEG0007709487000357.jpg214166JPEG0007709487000358.jpg214166JPEG0007709487000359.jpg214166JPEG0007709487000360.jpg214166JPEG0007709487000361.jpg214166JPEG0007709487000362.jpg214166JPEG0007709487000363.jpg214166JPEG0007709487000364.jpg214166JPEG0007709487000365.jpg214166JPEG0007709487000366.jpg214166JPEG0007709487000367.jpg214166JPEG0007709487000368.jpg214166JPEG0007709487000369.jpg214166JPEG0007709487000370.jpg214166JPEG0007709487000371.jpg214166JPEG0007709487000372.jpg214166JPEG0007709487000373.jpg214166JPEG0007709487000374.jpg214166JPEG0007709487000375.jpg214166JPEG0007709487000376.jpg214166JPEG0007709487000377.jpg214166JPEG0007709487000378.jpg214166JPEG0007709487000379.jpg214166JPEG0007709487000380.jpg214166JPEG0007709487000381.jpg214166JPEG0007709487000382.jpg214166JPEG0007709487000383.jpg214166JPEG0007709487000384.jpg214166JPEG0007709487000385.jpg214166JPEG0007709487000386.jpg214166JPEG0007709487000387.jpg214166JPEG0007709487000388.jpg214166JPEG0007709487000389.jpg214166JPEG0007709487000390.jpg214166JPEG0007709487000391.jpg214166JPEG0007709487000392.jpg214166JPEG0007709487000393.jpg214166JPEG0007709487000394.jpg214166JPEG0007709487000395.jpg214166JPEG0007709487000396.jpg214166JPEG0007709487000397.jpg214166JPEG0007709487000398.jpg214166JPEG0007709487000399.jpg214166JPEG0007709487000400.jpg214166JPEG0007709487000401.jpg214166JPEG0007709487000402.jpg214166.
[0126] Modified Cpf1 enzyme Zetsche et al. (2015) described individual regions of Cpf1. First, the C-terminal RuvC-like domain, which is the only functionally characterized domain. Second, the N-terminal α-helix region, and third, the mixed α and β region located between the RuvC-like domain and the α-helix region.
[0127] The crystal structure of Cpf1 provided herein provides further information regarding DNA-interacting amino acids (see Examples). Based on this information, mutants can be created that lead to enzyme inactivation or convert the double-strand nuclease to nickase activity. In alternative embodiments, this information is used to develop enzymes with reduced off-target effects (described elsewhere in this specification).
[0128] In certain embodiments of the Cpf1 enzyme described above, one or more modified or mutated amino acid residues are selected from any one amino acid in the regions of D861, R862, R863, W382, E993, D1263, D908, W958, K968, R951, R1226, S1228, D1235, K548, M604, K607, T167, N631, N630, K547, K163, Q571, K1017, R955, K1009, R909, R912, R1072, E372, K15, K810, H755, K557, E857, K943, K1022, K1029, K942, K949, R84, K87, K200, H206, R210, R301, R699, K705, K887, R891, K1086, K1089, R1094, R1127, R1220, Q1224, N178, N197, N204, N259, N278, N282, N519, N747, N759, N878, N889, and / or in the regions of 1189 - 1197, 1200 - 1208, 398 - 400, 380 - 383, 362 - 420, 1163 - 1173, 1230 - 1233, 1152 - 1148, 1076 - 1249, based on the amino acid numbering of Acidaminococcus sp. BV3L6. In preferred embodiments, the one or more modified or mutated amino acid residues are selected from the list consisting of R862A, E993A, D1263A, D908A, W958A, R951A, R1226A, S1228A, D1235A, K548A, M604A, K607A, K607R, T167S, N631K, N613R, N630K, N630R, K547R, K163R, Q571K, Q571R, K1009A, R909A, R1072A, E327A, K15A, K810A, H755A, K557A, E857A, K943A, K1022A, K1029A, K942A, K949A, R84A, K87A, K200A, H206A, R210A, R301A, R699A, K705A, K887A, R891A, K1086A, K1089A, R1094A, R1127A, R1220A and Q1224A.In a preferred embodiment, the one or more modified or mutated amino acid residues are selected from the list consisting of R862A, E993A, D1263A, D908A, W958A, R951A, K548A, M604A, K607A, K607R, N631K, N613R, N630K, N630R, K547R, K163R, Q571K, Q571R, K1009A, R909A, R1072A, E327A, K15A, K810A, H755A, K557A, E857A, K943A, K1022A, K1029A, K942A, K949A, R84A, K87A, K200A, H206A, R210A, R301A, R699A, K705A, K887A, R891A, K1086A, K1089A, R1094A, R1127A, R1220A and Q1224A; in a preferred embodiment, the one or more modified or mutated amino acid residues are selected from N178, N197, N204, N259, N278, N282, N519, N747, N759, N878, N889. In a preferred embodiment, the one or more modified or mutated amino acid residues are selected from the list consisting of R862A, W958A, R951A, R1226A, S1228A, D1235A, K548A, M604A, K607A, K607R, T167S, N631K, N613R, N630K, N630R, K547R, K163R, Q571K, Q571R, K1009A, R909A, R1072A, E327A, K15A, K810A, H755A, K557A, E857A, K943A, K1022A, K1029A, K942A, K949A, R84A, K87A, K200A, H206A, R210A, R301A, R699A, K705A, K887A, R891A, K1086A, K1089A, R1094A, R1127A, R1220A and Q1224A.In a preferred embodiment, the one or more modified or mutated amino acid residues are selected from D861, W958, S1228, D1235, T167, N631, N630, K547, K163, Q571, R1226, E372, K15, K810, H755, K557, E857, K943, K1022, K1029, K942, K949, R84, K87, K200, H206, R210, R301, R699, K705, K887, R891, K1086, K1089, R1094, R1127, R1220, Q1224, N178, N197, N204, N259, N278, N282, N519, N747, N759, N878, N889, and / or any one amino acid in the regions of 1189 - 1197, 1200 - 1208, 398 - 400, 380 - 383, 362 - 420, 1163 - 1173, 1230 - 1233, 1152 - 1148, 1076 - 1249. In a detailed embodiment, the mutation is R862A, and the Cpf1 enzyme no longer binds to RNA. In a detailed embodiment, the one or more mutations are selected from K15A, K810A, H755A, K557A, E857A, R862A, K943A, K1022A, and K1029A, and the Cpf1 enzyme no longer has RNA binding ability and / or processing ability. In a detailed embodiment, the one or more mutations are selected from K5478A, K607A, and M604A, and the TTT specificity is reduced or removed. In a detailed embodiment, the one or more mutations are selected from N631K, N613R, N630K, N630R, K547R, K163R, Q571K, Q571R, and K607R, and the non-specific DNA interaction of the Cpf1 enzyme increases. In a detailed embodiment, the one or more mutations are selected from R84A, K87A, K200A, H206A, R210A, R301A, R699A, K705A, K887A, R891A, K1086A, K1089A, R1094A, R1127A, R1220A, and Q1224A, whereby the specificity of the enzyme increases or decreases. In a detailed embodiment, one or more of D861, R862, R863, and W382 are mutated, and the RNA binding property of the Cpf1 is disrupted.In a detailed embodiment, the stability of Cpf1 is affected by one or more of the amino acids W958, K968, R951, R1226, D1253, and T167. In a detailed embodiment, one or more of K968 and R951 are mutated, disrupting the DNA binding of said Cpf1. In a detailed embodiment, one or more of N631 and N630 are mutated, increasing the interaction with the phosphate of the DNA backbone. In a detailed embodiment, the following amino acids: based on the amino acid numbering of AsCpf1 (Acidaminococcus sp. BV3L6), L117, T118, D119, T150, T151, T152, R341, N342, E343, T398, G399, K400, D451, Q452, P453, L454, P455, T456, T457, L458, K459, V486, D487, E488, S489, N490, E491, V492, D493, P494, E506, M507, E508, Q571, K572, G573, R574, Y575, T621, E649, K650, E651, D665, T737, D749, F750, K815, N848, V1108, K1109, T1110, G1111, S1124, A1195, A1196, A1197, N1198, L1244, N1245 and / or G1246 are mutated, whereby the stability and / or activity of the Cpf1 enzyme is not substantially affected.
[0129] In certain of the above Cpf1 enzymes, the enzyme is modified by mutation of one or more residues (within the RuvC domain), including but not limited to positions 884 - 1307, such as positions 993, 1263 and / or 980, based on the amino acid numbering of AsCpf1 (Acidaminococcus sp. BV3L6).
[0130] Modification in regions with B-factors higher than average Regions with low order in a polymer crystal structure (including, but not limited to, disordered regions or unstructured regions), particularly regions with low order within the solvent-exposed regions of a protein (including, but not limited to, loops), represent regions that can be modified without destabilizing them to the extent that their structure or function is intolerable. B factors, temperature factors, thermal factors, Debye-Waller factors, atomic displacement parameters, and similar terms relate to values that are indicators of the displacement of an atom from its average position in a crystal structure (e.g., as a result of temperature-dependent atomic vibrations or static disorder in the crystal lattice). Thus, a B factor higher than the average of the backbone atoms in the solvent-exposed region of a protein is an indicator of a region with relatively high local mobility or a region that can be modified without destabilizing it to the extent that the protein structure or function is intolerable. Accordingly, in certain specific Cpf1 enzymes described herein, the Cpf1 enzyme is modified by one or more substitutions, insertions, deletions, or other modifications in a solvent-exposed region having one or more backbone atoms with a B factor higher than the average compared to the entire protein or a protein domain including the solvent-exposed region. In certain specific Cpf1 enzymes, the enzyme is modified with one or more residues having a Cα atom with a B factor that is 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or more than 200% higher than the average B factor of the protein containing the one or more residues. In certain specific Cpf1 enzymes, the enzyme is modified with the residues having a Cα atom with a B factor that is 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or more than 200% higher than the average B factor of a protein domain containing the one or more residues (e.g., the C-terminal RuvC-like domain, the N-terminal α-helix region, or the α- and β-mixed region between the N-terminal domain and the C-terminal domain).In certain Cpf1 enzymes, the enzyme is modified by one or more substitutions, insertions, deletions or other modifications at L117, T118, D119, T150, T151, T152, R341, N342, E343, T398, G399, K400, D451, Q452, P453, L454, P455, T456, T457, L458, K459, V486, D487, E488, S489, N490, E491, V492, D493, P494, E506, M507, E508, Q571, K572, G573, R574, Y575, T621, E649, K650, E651, D665, T737, D749, F750, K815, N848, V1108, K1109, T1110, G1111, S1124, A1195, A1196, A1197, N1198, L1244, N1245 and / or G1246, based on the amino acid numbering of AsCpf1 (Acidaminococcus sp. BV3L6).
[0131] Inactivated / inactivation Cpf1 protein When the Cpf1 protein has nuclease activity, the Cpf1 protein can be modified to have reduced nuclease activity, e.g., at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or 100% nuclease inactivation compared to the wild-type enzyme; or put another way, the Cpf1 enzyme preferably has about 0% nuclease activity of the non-mutated or wild-type Cpf1 enzyme or CRISPR enzyme, or about 3% or about 5% or about 10% or less nuclease activity of the non-mutated or wild-type Cpf1 enzyme, e.g., the non-mutated or wild-type Acidaminococcus sp. BV3L6 (AsCpf1) Cpf1 enzyme or CRISPR enzyme. This can be achieved by introducing mutations into the nuclease domain of Cpf1 and its orthologs.
[0132] More specifically, the inactivated Cpf1 enzyme includes an enzyme mutated at an amino acid position identified in AsCpf1 as directly or indirectly contributing to the nuclease activity of AsCpf1 or at the corresponding position in a Cpf1 ortholog.
[0133] The inactivated Cpf1 CRISPR enzyme may include, consist essentially of, or consist of one or more functional domains including, for example, methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, DNA cleavage activity, nucleic acid binding activity, and molecular switch (e.g., light-inducible) domains, and one or more such functional domains may be associated (e.g., via a fusion protein). Preferred domains are Fok1, VP64, P65, HSF1, MyoD1. When Fok1 is provided, it is advantageous to provide multiple Fok1 functional domains to achieve a functional dimer, and the gRNA is designed to provide an appropriate spacing for functional use (Fok1) as specifically described in Tsai et al. Nature Biotechnology, Vol. 32, Number 6, June 2014). Adapter proteins may bind such functional domains using known linkers. Optionally, it is advantageous to provide at least one more NLS. Optionally, it is advantageous to place the NLS at the N-terminus. If two or more functional domains are included, they may be the same or different.
[0134] Generally, the positions of one or more functional domains on an inactivated Cpf1 enzyme are such that they allow for the correct spatial arrangement for the functional domain to exert its functional effect. For example, if the functional domain is a transcriptional activator (e.g., VP64 or p65), the transcriptional activator is placed in a spatial arrangement where it can affect the transcription of the target. Similarly, a transcriptional repressor will advantageously be arranged to affect the transcription of the target, and a nuclease (e.g., Fok1) will advantageously be arranged to cleave or partially cleave the target. This can include positions other than the N-terminus / C-terminus of the CRISPR enzyme.
[0135] The enzyme according to the present invention can be applied in an optimized functional CRISPR-Cas system that is beneficial for functional screening. Thus, it is also contemplated that a nucleic acid targeting effector protein-guide RNA complex can associate with two or more functional domains as a whole. For example, there may be two or more functional domains associated with the nucleic acid targeting effector protein, or two or more functional domains associated with the guide RNA (via one or more adapter proteins), or one or more functional domains associated with the nucleic acid targeting effector protein and one or more functional domains associated with the guide RNA (via one or more adapter proteins).
[0136] Using two different aptamers (each associated with an individual nucleic acid targeting guide RNA), it becomes possible to use activator - adapter protein fusions and repressor - adapter protein fusions with different nucleic acid targeting guide RNAs to activate the expression of one DNA or RNA while suppressing the expression of another DNA or RNA. These can be administered together, or substantially together, in a multiplexing approach with their different guide RNAs. Since a relatively small number of effector protein molecules can be used with a large number of modified guides, for example, a large number of such modified nucleic acid targeting guide RNAs, such as 10 or 20 or 30, can all be used simultaneously while delivering only one (or at least a minimum number of) effector protein molecules. The adapter protein may associate (preferably ligate or fuse) with one or more activators or one or more repressors. For example, the adapter protein may associate with a first activator and a second activator. The first and second activators may be the same, but preferably they are different activators. Three or more or even four or more activators (or repressors) may be used, but depending on the package size, the number may be limited to more than 5 different functional domains. A linker is preferably used compared to a direct fusion of the adapter protein with two or more functional domains associating with the adapter protein. Suitable linkers can include GlySer linkers.
[0137] The fusion between an adapter protein and an activator or repressor may include a linker. For example, the GlySer linker GGGS (SEQ ID NO: 18) can be used. By using these in 3 ((GGGGS)3 (SEQ ID NO: 19)) or 6 (SEQ ID NO: 20), 9 (SEQ ID NO: 21) or even 12 (SEQ ID NO: 22) or more repeats, a suitable length can be provided as needed. The linker can be used between the guide RNA and the functional domain (activator or repressor), or between the nucleic acid targeting Cas protein (Cas) and the functional domain (activator or repressor). With the linker, the user manipulates an appropriate amount of "mechanical flexibility".
[0138] The present invention encompasses a nucleic acid targeting complex comprising a nucleic acid targeting effector protein and a guide RNA, wherein the nucleic acid targeting effector protein has at least one mutation [so that the nucleic acid targeting effector protein has an activity of 5% or less of the nucleic acid targeting effector protein that does not have the at least one mutation], and optionally comprises at least one or more nuclear localization sequences; the guide RNA comprises a guide sequence capable of hybridizing to a target sequence in a target RNA in a cell; and wherein: the nucleic acid targeting effector protein associates with two or more functional domains; or at least one loop of the guide RNA is modified by the insertion of one or more individual RNA sequences that bind to one or more adapter proteins, and wherein the adapter protein associates with two or more functional domains; or the nucleic acid targeting Cas protein associates with one or more functional domains, and at least one loop of the guide RNA is modified by the insertion of one or more individual RNA sequences that bind to one or more adapter proteins, and wherein the adapter protein associates with one or more functional domains.
[0139] In one aspect, the present invention provides a non - naturally occurring or engineered composition comprising a V - type, more particularly a Cpf1 CRISPR guide RNA, comprising a guide sequence capable of hybridizing to a target sequence at a target genomic locus of a cell, wherein the guide RNA is modified by the insertion of one or more individual RNA sequences that bind to two or more adapter proteins (e.g., aptamers), and each adapter protein is associated with one or more functional domains; or, wherein the guide RNA is modified to have at least one non - coding functional loop. In a detailed embodiment, the guide RNA is modified by the insertion of one or more individual RNA sequences 5' to the direct repeat, within the direct repeat, or 3' to the guide sequence. When there are two or more functional domains, they may be the same or different, for example, they may be the same two or two different activators or repressors. In one aspect, the present invention provides a non - naturally occurring or engineered CRISPR - Cas complex composition comprising a guide RNA as contemplated herein and a Cpf1 enzyme, a CRISPR enzyme [optionally the Cpf1 enzyme comprises at least one mutation, such that the Cpf1 enzyme has a nuclease activity of 5% or less of a Cpf1 enzyme without its at least one mutation, and optionally comprises one or more including at least one or more nuclear localization sequences]. In one aspect, the present invention provides a complex comprising a Cpf1 CRISPR guide RNA or a Cpf1 CRISPR - Cas complex as contemplated herein, a non - naturally occurring or engineered composition comprising two or more adapter proteins, wherein each protein is associated with one or more functional domains, and the adapter proteins bind to one or more individual RNA sequences inserted into the guide RNA. In a detailed embodiment, the guide RNA is modified in addition to or instead to still ensure binding of the Cpf1 CRISPR complex, but prevent cleavage by the Cpf1 enzyme.
[0140] Enzyme mutations that reduce off - target effects In one aspect, the present invention provides a non-naturally occurring or engineered CRISPR enzyme as described herein having one or more mutations that result in a reduction of off-target effects, preferably a class 2 CRISPR enzyme, preferably a type V or type VI CRISPR enzyme, such as, for example and without limitation, Cpf1 as described elsewhere herein, i.e., an improved CRISPR enzyme that is used to introduce a modification at a target locus but has reduced or abolished off-target directed activity, such as when complexed with a guide RNA, and an improved CRISPR enzyme for increasing the activity of a CRISPR enzyme, such as when forming a complex with a guide RNA. It should be understood that the mutant enzymes as described below herein can be used in any of the methods according to the invention as described elsewhere herein. Any of the methods, products, compositions and uses as described elsewhere herein are similarly applicable to the mutant CRISPR enzymes as further detailed below. In the aspects and embodiments as described herein, when referring to or reading Cpf1 as a CRISPR enzyme, it should be understood that the reconstitution of a functional CRISPR-Cas system preferably does not require or is not dependent on a tracr sequence and / or the direct repeat is on the 5'(upstream) side of the guide (target or spacer) sequence.
[0141] Slaymaker et al. recently described a method for creating Cas9 orthologs with enhanced specificity (Slaymaker et al. 2015 "Rationally engineered Cas9 nucleases with improved specificity"). Using this strategy, the specificity of the Cpf1 enzyme can be enhanced. The major residues for mutagenesis are preferably all positively charged residues within the RuvC domain. Additional residues are positively charged residues conserved between different orthologs.
[0142] In certain embodiments, the enzyme is modified by mutation of one or more residues (within the RuvC domain) including, but not limited to, positions R909, R912, R930, R947, K949, R951, R955, K965, K968, K1000, K1002, R1003, K1009, K1017, K1022, K1029, K1035, K1054, K1072, K1086, R1094, K1095, K1109, K1118, K1142, K1150, K1158, K1159, R1220, R1226, R1242, and / or R1252, based on the amino acid numbering of AsCpf1 (Acidaminococcus sp. BV3L6). In certain ones of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of one or more residues (within the RAD50 domain) including, but not limited to, positions K324, K335, K337, R331, K369, K370, R386, R392, R393, K400, K404, K406, K408, K414, K429, K436, K438, K459, K460, K464, R670, K675, R681, K686, K689, R699, K705, R725, K729, K739, K748, and / or K752, based on the amino acid numbering of AsCpf1 (Acidaminococcus sp. BV3L6).
[0143] In certain embodiments, the specificity of Cpf1 can be improved by mutating residues that stabilize the non-target DNA strand.
[0144] In one aspect, the present invention also provides methods and mutations for modulating Cas (e.g., Cpf1) binding activity and / or binding specificity. In certain embodiments, a Cas (e.g., Cpf1) protein lacking nuclease activity is used. In certain embodiments, a modified guide RNA that promotes binding of a Cas (e.g., Cpf1) nuclease but does not promote nuclease activity is utilized. In such embodiments, on-target binding can be increased or decreased. Also, in such embodiments, off-target binding can be increased or decreased. Further, there can be an increase or decrease in specificity with respect to on-target binding versus off-target binding.
[0145] Methods and mutations can be used in various combinations to increase or decrease the activity and / or specificity of on-target versus off-target activity, or to increase or decrease the binding and / or specificity of on-target versus off-target binding, and to compensate for or enhance mutations or modifications added so as to promote other effects. Such mutations or modifications added so as to promote other effects include mutations or modifications to Cas (e.g., Cpf1) and / or mutations or modifications added to the guide RNA. In certain embodiments, the methods and mutations are used with chemically modified guide RNAs. Examples of chemical modifications of guide RNAs include, without limitation, incorporation of 2'-O-methyl (M), 2'-O-methyl 3'-phosphorothioate (MS), or 2'-O-methyl 3'-thioPACE (MSP) at one or more terminal nucleotides. Such chemically modified guide RNAs can include high stability and high activity when compared to unmodified guide RNAs, however, on-target versus off-target specificity is unpredictable (see Hendel, 2015, Nat Biotechnol. 33(9):985-9, doi:10.1038 / nbt.3290, online publication 29 June 2015). Chemically modified guide RNAs further include, without limitation, RNAs having phosphorothioate linkages and locked nucleic acid (LNA) nucleotides containing a methylene bridge between the 2' and 4' carbons of the ribose ring. The methods and mutations of the invention are used to modulate Cas (e.g., Cpf1) nuclease activity and / or binding by chemically modified guide RNAs.
[0146] In certain aspects, the invention provides methods and mutations for altering the binding and / or binding specificity of a Cas (e.g., Cpf1) protein according to the invention as defined herein, including functional domains such as nucleases, transcriptional activators, transcriptional repressors. For example, a Cas (e.g., Cpf1) protein can be made nuclease-null, or have altered or reduced nuclease activity, by introducing mutations such as Cpf1 mutations described elsewhere herein, e.g., including D908A, E993A, D1263A in the AsCpf1 protein or corresponding positions in orthologs. Nuclease-deficient Cas (e.g., Cpf1) proteins are useful for RNA-guided delivery of functional domains to target sequences. The invention provides methods and mutations for modulating the binding of a Cas (e.g., Cpf1) protein. In one embodiment, the functional domain includes VP64, which provides an RNA-guided transcriptional factor. In another embodiment, the functional domain includes FokI, which provides RNA-guided nuclease activity. U.S. Patent Application Publication No. 2014 / 0356959, U.S. Patent Application Publication No. 2014 / 0342456, U.S. Patent Application Publication No. 2015 / 0031132, and Mali, P. et al., 2013, Science 339(6121):823-6, doi:10.1126 / science.1232033, online publication 3 January 2013 are hereby incorporated by reference in their entirety, and the invention includes the methods and materials of these documents as applied in connection with the teachings herein. In certain embodiments, on-target binding is increased. In certain embodiments, off-target binding is decreased. In certain embodiments, on-target binding is decreased. In certain embodiments, off-target binding is increased. Accordingly, the invention also provides for increasing or decreasing the specificity of on-target binding versus off-target binding of a functional Cas (e.g., Cpf1) binding protein.
[0147] The use of Cas (e.g., Cpf1) as an RNA-guided binding protein is not limited to nuclease-null Cas (e.g., Cpf1). Cas (e.g., Cpf1) enzymes that include nuclease activity can also function as RNA-guided binding proteins when used with a specific guide RNA. For example, a short guide RNA and a guide RNA containing nucleotides mismatched to the target can promote Cas9 binding directed by the RNA to the target sequence with little or no target cleavage (see, e.g., Dahlman, 2015, Nat Biotechnol. 33(11):1159-1161, doi:10.1038 / nbt.3390, online publication 05 October 2015). In one aspect, the present invention provides methods and mutations for modulating the binding of Cas (e.g., Cpf1) proteins that include nuclease activity. In certain embodiments, on-target binding is increased. In certain embodiments, off-target binding is decreased. In certain embodiments, on-target binding is decreased. In certain embodiments, off-target binding is increased. In certain embodiments, there is an increase or decrease in the specificity of on-target binding versus off-target binding. In certain embodiments, the nuclease activity of the guide RNA-Cas (e.g., Cpf1) enzyme is also modulated.
[0148] RNA-DNA heteroduplex formation is important for cleavage activity and specificity not only in the seed region sequence closest to the PAM but also throughout the target region. Thus, truncated guide RNAs exhibit decreased cleavage activity and specificity. In one aspect, the present invention provides methods and mutations for increasing the activity and specificity of cleavage using a modified guide RNA.
[0149] In any of the non-naturally occurring CRISPR enzymes, the CRISPR enzyme can include one or more heterologous functional domains.
[0150] One or more heterologous functional domains may contain one or more nuclear localization signal (NLS) domains. One or more heterologous functional domains may contain at least two or more NLSs.
[0151] One or more heterologous functional domains contain one or more transcriptional activation domains. The transcriptional activation domain may contain VP64.
[0152] One or more heterologous functional domains contain one or more transcriptional repression domains. The transcriptional repression domain may contain a KRAB domain or a SID domain.
[0153] One or more heterologous functional domains may contain one or more nuclease domains. One or more nuclease domains may contain Fok1.
[0154] One or more heterologous functional domains may have one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity.
[0155] At least one or more heterologous functional domains may be at or near the amino terminus of the enzyme and / or at or near the carboxy terminus of the enzyme.
[0156] One or more heterologous functional domains may be fused to a CRISPR enzyme, or tethered to a CRISPR enzyme, or linked to a CRISPR enzyme by a linker moiety.
[0157] In any of the non-naturally occurring CRISPR enzymes, the CRISPR enzyme can comprise a CRISPR enzyme from an organism of a genus comprising Francisella tularensis1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis3, Prevotella disiens, or Porphyromonas macacae (e.g., a Cpf1 of one of these organisms modified as described herein), and can include additional mutations or variations, or can be a chimeric Cas (e.g., Cpf1).
[0158] In any of the non-naturally occurring CRISPR enzymes, the CRISPR enzyme can comprise a chimeric Cas (e.g., Cpf1) enzyme comprising a first fragment from a first Cas (e.g., Cpf1) ortholog and a second fragment from a second Cas (e.g., Cpf1) ortholog, wherein the first and second Cas (e.g., Cpf1) orthologs are different. At least one of the first and second Cas (e.g., Cpf1) orthologs can comprise a Cas (e.g., Cpf1) from an organism comprising Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, or Porphyromonas macacae.
[0159] In any of the non-naturally occurring CRISPR enzymes, the nucleotide sequence encoding the CRISPR enzyme may be codon-optimized for expression in eukaryotes.
[0160] In any of the non-naturally occurring CRISPR enzymes, the cell may be a eukaryotic or prokaryotic cell; wherein the CRISPR complex is operable in the cell, whereby the enzyme of the CRISPR complex has a reduced ability to modify one or more off-target loci of the cell when compared to the unmodified enzyme, and / or whereby the enzyme in the CRISPR complex has an increased ability to modify one or more target loci when compared to the unmodified enzyme.
[0161] Accordingly, in one aspect, the present invention provides a eukaryotic cell comprising an engineered CRISPR protein or system as defined herein.
[0162] In certain embodiments, the method as described herein may include providing Cpf1 transgenic cells in which one or more nucleic acids encoding one or more guide RNAs operably linked intracellularly to a regulatory element comprising a promoter of one or more target genes are provided or introduced. As used herein, the term “Cpf1 transgenic cell” refers to a cell such as a eukaryotic cell in which the Cpf1 gene is genomically integrated. The nature, type, or origin of the cell is not particularly limited according to the present invention. Also, the method of introducing the Cpf1 transgene into the cell may be various and may be any method known in the art. In certain embodiments, the Cpf1 transgenic cells are obtained by introducing the Cpf1 transgene into isolated cells. In certain other embodiments, the Cpf1 transgenic cells are obtained by isolating cells from a Cpf1 transgenic organism. By way of example and without limitation, the Cpf1 transgenic cells as referred to herein may be derived from a Cpf1 transgenic eukaryote such as a Cpf1 knock-in eukaryote. International Publication No. WO 2014 / 093622 pamphlet (PCT / US13 / 74667 specification) (incorporated herein by reference) is referred to. The methods of U.S. Patent Application Publication Nos. 20120017290 and 20110265198, assigned to Sangamo BioSciences, Inc. and relating to targeting of the Rosa locus, may be modified to utilize the CRISPR Cpf1 system of the present invention. The method of U.S. Patent Application Publication No. 20130236946, assigned to Cellectis and relating to targeting of the Rosa locus, may also be modified to utilize the CRISPR Cpf1 system of the present invention. As a further example, Platt et.al. (Cell; 159(2):440-455(2014)) (incorporated herein by reference) describing Cas9 knock-in mice is referred to, and this can be applied to the CRISPR enzyme of the present invention as defined herein.The Cpf1 transgene may further include a Lox-Stop-polyA-Lox (LSL) cassette, whereby Cpf1 expression can be made inducible by Cre recombinase. Alternatively, Cpf1 transgenic cells may be obtained by introducing the Cpf1 transgene into isolated cells. Delivery systems for transgenes are well known in the art. By way of example, the Cpf1 transgene can be delivered using, for example, vectors (e.g., AAV, adenovirus, lentivirus) and / or particles and / or nanoparticle delivery in eukaryotic cells, as described elsewhere herein as well.
[0163] One of ordinary skill in the art will understand that cells such as Cpf1 transgenic cells as referred to herein may contain additional genomic changes in addition to the integrated Cpf1 gene, as described, for example, and without limitation, in Platt et al. (2014), Chen et al., (2014) or Kumar et al.. (2009), or may contain mutations resulting from the sequence-specific action of Cpf1 when forming a complex with an RNA capable of guiding Cpf1 to a target locus, such as one or more oncogenic mutations.
[0164] The present invention also provides a composition comprising an engineered CRISPR protein as described herein, such as described in this section.
[0165] The present invention also provides an engineered composition that is not naturally occurring, comprising a CRISPR-Cas complex comprising any non-naturally occurring CRISPR enzyme described above.
[0166] In one aspect, the present invention provides a vector system comprising one or more vectors, wherein the one or more vectors are a) a first regulatory element operably linked to a nucleotide sequence encoding an engineered CRISPR protein as defined herein; and optionally, b) one or more nucleotide sequences encoding one or more nucleic acid molecules comprising a guide RNA comprising a guide sequence and a direct repeat sequence, and a second regulatory element operably linked to the one or more nucleotide sequences, optionally, components (a) and (b) are located on the same or different vectors.
[0167] The present invention also provides a delivery system operably configured to deliver to a cell a CRISPR-Cas complex component or one or more polynucleotide sequences comprising or encoding said component, wherein said CRISPR-Cas complex is operable in the cell and also provides an engineered composition that does not occur in nature and comprises a CRISPR-Cas complex component or one or more polynucleotide sequences encoding a CRISPR-Cas complex component for transcription and / or translation in a cell, (I) a non-naturally occurring CRISPR enzyme described herein (e.g., engineered Cpf1); (II) a CRISPR-Cas guide RNA comprising a guide sequence, a direct repeat sequence, wherein: the enzyme in the CRISPR complex has a reduced ability to modify one or more off-target loci when compared to the unmodified enzyme, and / or thereby the enzyme in the CRISPR complex has an increased ability to modify one or more target loci when compared to the unmodified enzyme.
[0168] In certain embodiments, the present invention also provides a system comprising an engineered CRISPR protein as described herein, such as described in this section.
[0169] In any such composition, as described anywhere herein, the delivery system can include a yeast-based, lipofection-based, microinjection-based, particle gun-based, virosome, liposome, immunoliposome, polycation, lipid:nucleic acid conjugate, or artificial virion.
[0170] In any such composition, the delivery system may include a vector system comprising one or more vectors, wherein the component (II) comprises a first regulatory element operably linked to a polynucleotide sequence optionally including a guide sequence and a direct repeat sequence, and the component (I) comprises a second regulatory element operably linked to a polynucleotide sequence encoding a CRISPR enzyme.
[0171] In any such composition, the delivery system may include a vector system comprising one or more vectors, wherein the component (II) comprises a first regulatory element operably linked to a guide sequence and a direct repeat sequence, and the component (I) comprises a second regulatory element operably linked to a polynucleotide sequence encoding a CRISPR enzyme.
[0172] In any such composition, the composition can include two or more guide RNAs, each guide RNA having a different target, thereby allowing for multiplexing.
[0173] In any such composition, one or more polynucleotide sequences may be on one vector.
[0174] The present invention also relates to a) a first regulatory element operably linked to a nucleotide sequence encoding a non-naturally occurring CRISPR enzyme of any one of the constructs of the present invention herein; and b) a second regulatory element operably linked to one or more nucleotide sequences encoding one or more guide RNAs, wherein the guide RNA includes a guide sequence and a direct repeat sequence. An engineered, non-naturally occurring clustered regularly interspaced short palindromic repeat (CRISPR)-CRISPR associated (Cas) (CRISPR-Cas) vector system is also provided that includes one or more vectors that include: Components (a) and (b) are located on the same or different vectors, a CRISPR complex is formed; a guide RNA targets a target polynucleotide locus and an enzyme changes the polynucleotide locus, and the enzyme in the CRISPR complex has a reduced ability to modify one or more off-target loci when compared to an unmodified enzyme and / or thereby the enzyme in the CRISPR complex has an increased ability to modify one or more target loci when compared to an unmodified enzyme.
[0175] In such a system, component (II) can include a first regulatory element operably linked to a polynucleotide sequence that includes a guide sequence and a direct repeat sequence, and component (II) can include a second regulatory element operably linked to a polynucleotide sequence encoding a CRISPR enzyme. In such a system, where applicable, the guide RNA can include a chimeric RNA.
[0176] In such a system, component (I) can include a first regulatory element operably linked to a guide sequence and a direct repeat sequence, and component (II) can include a second regulatory element operably linked to a polynucleotide sequence encoding a CRISPR enzyme. Such a system can include two or more guide RNAs, each having a different target, whereby multiplexing occurs. Components (a) and (b) can be on the same vector.
[0177] In any such system containing vectors, one or more of the vectors can contain one or more viral vectors, such as one or more retroviruses, lentiviruses, adenoviruses, adeno-associated viruses or herpes simplex viruses.
[0178] In any such system containing regulatory elements, at least one of the regulatory elements can include a tissue-specific promoter. The tissue-specific promoter can direct expression in mammalian blood cells, mammalian hepatocytes or the mammalian eye.
[0179] In any of the above-described compositions or systems, the direct repeat sequence can include one or more protein-interacting RNA aptamers. One or more of the aptamers can be located in a tetraloop. One or more of the aptamers can have the ability to bind to the MS2 bacteriophage coat protein.
[0180] In any of the above-described compositions or systems, the cell can be a eukaryotic cell or a prokaryotic cell; wherein the CRISPR complex is operable in the cell, whereby the enzyme of the CRISPR complex has a reduced ability to modify one or more off-target gene loci in the cell when compared to an unmodified enzyme, and / or whereby the enzyme in the CRISPR complex has an increased ability to modify one or more target gene loci when compared to an unmodified enzyme.
[0181] The present invention also provides a CRISPR complex from any of the compositions described above or from any of the systems described above.
[0182] The present invention also provides a method of modifying a target locus of a cell, the method comprising contacting the cell with any of the engineered CRISPR enzymes described herein (e.g., engineered Cpf1), any of the compositions, or any of the systems or vector systems described herein, or wherein the cell comprises any of the CRISPR complexes described herein that are present within the cell. In such a method, the cell may be a prokaryotic cell or a eukaryotic cell, preferably a eukaryotic cell. In such a method, an organism may comprise the cell. In such a method, the organism may not be a human or other animal.
[0183] Any such method may be ex vivo or in vitro.
[0184] In certain embodiments, the nucleotide sequence encoding at least one of the guide RNA or the Cas protein is operably linked in the cell with a regulatory element comprising a promoter of a gene of interest, whereby the expression of at least one CRISPR-Cas system component is driven by the promoter of the gene of interest. "Operably linked" is intended to mean that the nucleotide sequence encoding the guide RNA and / or Cas is linked to one or more regulatory elements in such a manner as to enable expression of the nucleotide sequence, as also referred to elsewhere in this specification. The term "regulatory element" is also described elsewhere in this specification. According to the present invention, the regulatory element preferably comprises a promoter of the gene of interest, such as the promoter of an endogenous gene of interest. In certain embodiments, the promoter is at its endogenous genomic location. In such embodiments, the nucleic acid encoding CRISPR and / or Cas is under the transcriptional control of the promoter of the gene of interest at its natural genomic location. In certain other embodiments, the promoter is provided on a (separate) nucleic acid molecule such as a vector or plasmid, or on other extrachromosomal nucleic acid, i.e., the promoter is not provided at its natural genomic location. In certain embodiments, the promoter is genomically integrated at a non-natural genomic location.
[0185] For any such method, the modification may include modification of gene expression. The modification of gene expression may include activating gene expression and / or suppressing gene expression. Thus, in one aspect, the present invention provides a method of regulating gene expression, the method comprising introducing into a cell an engineered CRISPR protein or system as described herein.
[0186] The present invention also provides a method of treating a disease, disorder or infection in an individual in need thereof, the method comprising administering an effective amount of any of the engineered CRISPR enzymes (e.g., engineered Cpf1), compositions, systems or CRISPR complexes described herein. The disease, disorder or infection may include viral infection. The viral infection may be HBV.
[0187] The present invention also provides the use of any of the engineered CRISPR enzymes (e.g., engineered Cpf1), compositions, systems or CRISPR complexes described above for gene or genome editing.
[0188] The present invention also provides a method of changing the expression of a genomic locus of interest in mammalian cells, the method comprising contacting the cells with an engineered CRISPR enzyme (e.g., engineered Cpf1), composition, system or CRISPR complex as described herein, thereby delivering the CRISPR-Cas (vector) and forming a CRISPR-Cas complex that binds to the target, and determining whether the expression of the genomic locus has changed, such as an increase or decrease in expression, or a modification of the gene product.
[0189] The present invention also provides any one of the engineered CRISPR enzymes (e.g., engineered Cpf1) described above, compositions, systems, or CRISPR complexes for use as a therapeutic agent. The therapeutic agent may be for gene or genome editing, or gene therapy.
[0190] In certain embodiments, the activity of the engineered CRISPR enzyme (e.g., engineered Cpf1) as described herein includes genomic DNA cleavage that optionally results in a decrease in gene transcription.
[0191] In one aspect, the present invention provides an isolated cell in which the expression of a genomic locus has been altered by any of the methods described herein, where the alteration in expression is a comparison to a cell that has not been subjected to a method of altering the expression of the genomic locus. In a related aspect, the present invention provides a cell line established from such a cell.
[0192] In one aspect, the present invention provides a method of modifying an organism or non-human organism by manipulating a target sequence at a genomic locus of interest (e.g., a hematopoietic stem cell (HSC)), where the genomic locus of interest is associated with abnormal protein expression or a mutation associated with a disease condition or state. The method comprises I. a CRISPR-Cas system guide RNA (gRNA) polynucleotide sequence comprising (a) a guide sequence capable of hybridizing to a target sequence within the HSC, (b) a direct repeat sequence, and delivering to the HSC, by contacting the HSC with a non-naturally occurring or engineered composition comprising the polynucleotide sequence and II. a CRISPR enzyme, optionally comprising at least one or more nuclear localization sequences, wherein the guide sequence directs sequence-specific binding of the CRISPR complex to the target sequence, and Here, the CRISPR complex comprises a guide sequence hybridizing to a target sequence and a CRISPR enzyme forming a complex therewith, and the method also optionally includes delivering an HDR template, for example, by contacting the HSCs with an HDR template-containing particle or via a particle that contacts the HSCs with another particle containing the HDR template, where the HDR template results in the expression of a normal or hypoabnormal protein; "normal" is with respect to the wild type, and "abnormal" may be protein expression that causes a pathological or diseased state; optionally the method may include isolating or obtaining HSCs from an organism or non-human organism, optionally expanding the HSC population, performing contact between one or more particles and the HSCs to obtain a modified HSC population, optionally expanding the population of modified HSCs, and optionally administering the modified HSCs to an organism or non-human organism.
[0193] This delivery can be, for example, the delivery of one or more polynucleotides encoding any one or more or all of the CRISPR complexes via one or more particles containing a vector containing one or more polynucleotides operably linked to one or more regulatory elements, advantageously polynucleotides linked to one or more regulatory elements for in vivo expression. The polynucleotide sequence encoding the CRISPR enzyme, the guide sequence, a part or all of the direct repeat sequence may be RNA. When a polynucleotide that is RNA and is said to "contain" the characteristics of such a direct repeat sequence is mentioned, it will be understood that the RNA sequence has that characteristic. When a polynucleotide is DNA and is said to contain the characteristics of such a direct repeat sequence, the DNA sequence is transcribed or can be transcribed into an RNA containing the characteristic in question. When the characteristic is a protein such as a CRISPR enzyme, the DNA or RNA sequence mentioned is translated or can be translated (in the case of DNA, after being transcribed first).
[0194] In certain embodiments, the present invention provides a method of modifying an organism, such as a mammal including a human or a non-human mammal or an organism, at a target genomic locus of a hematopoietic stem cell (HSC) associated with, for example, abnormal protein expression or a mutation associated with a disease condition or state, the method comprising delivering a non-naturally occurring or engineered composition into contact with the HSC, for example, by operably manipulating a target sequence at the target genomic locus of the HSC, wherein the composition comprises one or more particles comprising one or more viral, plasmid or nucleic acid molecular vectors (e.g., RNA) that functionally encode the composition for expression thereof, the composition comprising (A) a first regulatory element operably linked to an I. CRISPR-Cas system RNA polynucleotide sequence, the polynucleotide sequence comprising (a) a guide sequence capable of hybridizing to a target sequence in a eukaryotic cell, (b) a direct repeat sequence, the first regulatory element, and II. a second regulatory element operably linked to an enzyme coding sequence encoding a CRISPR enzyme comprising at least one or more nuclear localization sequences (or optionally at least one or more nuclear localization sequences since in some embodiments the NLS may not be involved) [(a), (b), and (c) are arranged in a 5' to 3' direction, components I and II are located on the same or different vectors of the system, and when transcribed, and the guide sequence induces sequence-specific binding of the CRISPR complex to the target sequence, and the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence that hybridizes to the target sequence], or (B) a non-naturally occurring or engineered composition comprising I. (a) a guide sequence capable of hybridizing to a target sequence in a eukaryotic cell, and (b) a first regulatory element operably linked to at least one or more direct repeat sequences, II.A composition comprising a vector system comprising one or more vectors comprising a second regulatory element operably linked to an enzyme coding sequence encoding a CRISPR enzyme, [wherein components I and II are located on the same or different vectors of the system and, when transcribed, and the guide sequence induces sequence-specific binding of the CRISPR complex to the target sequence, and the CRISPR complex comprises a CRISPR enzyme that forms a complex with a guide sequence that hybridizes to the target sequence]; the method can also optionally include delivering an HDR template, for example, by contacting the HSCs with particles containing the HDR template or by contacting the HSCs with another particle containing the HDR template, wherein the HDR template results in the expression of a normal or hypoabnormal protein; "normal" is with respect to the wild type, and "abnormal" can be protein expression that causes a pathological or disease state; and optionally the method can include isolating or obtaining HSCs from a biological or non-human organism, optionally expanding the HSC population, contacting one or more particles with the HSCs to obtain a modified HSC population, optionally expanding the population of modified HSCs, and optionally administering the modified HSCs to a biological or non-human organism. In some embodiments, components I, II, and III are located on the same vector. In other embodiments, components I and II are located on the same vector, while component III is located on a different vector. In other embodiments, components I and III are located on the same vector, while component II is located on a different vector. In other embodiments, components II and III are located on the same vector, while component I is located on a different vector. In other embodiments, each of components I, II, and III is located on a different vector. The invention also provides a viral or plasmid vector system as described herein.
[0195] By the manipulation of a target array, the applicants also mean the epigenetic manipulation of the target array. This may be an operation of the chromatin state of the target array, such as modification of the methylation state of the target array (i.e., methylation or addition or removal of a methylation pattern or CpG island), histone modification, increase or decrease in the accessibility to the target array, or promotion of three-dimensional folding. When a method of modifying an organism including a human or a mammal or a non-human mammal or an organism by the manipulation of a target array at a target genomic locus is mentioned, it will be understood that this may be applied to the organism (or mammal) as a whole or, if the organism is a multicellular organism, only to a single cell or cell population of the organism. For example, in the case of a human, the applicants particularly envision single cells or cell populations, which may preferably be modified ex vivo and then reintroduced. In this case, a biopsy or other tissue sample or biological fluid sample may be required. Stem cells are also particularly preferred in this regard. However, in vivo embodiments are of course also envisioned. And the present invention is particularly advantageous with respect to HSCs.
[0196] In some embodiments, the present invention is, for example, I. A first CRISPR-Cas (e.g., Cpf1) system RNA (RNA) polynucleotide sequence, comprising (a) A first guide sequence capable of hybridizing to a first target sequence, (b) A first direct repeat sequence, and A first polynucleotide sequence comprising II. A second CRISPR-Cas (e.g., Cpf1) system guide RNA polynucleotide sequence, comprising (a) A second guide sequence capable of hybridizing to a second target sequence, (b) A second direct repeat sequence, and A second polynucleotide sequence comprising, and III. A polynucleotide sequence encoding a CRISPR enzyme comprising at least one or more nuclear localization sequences and comprising one or more mutations [(a), (b) and (c) are arranged in the 5' to 3' direction]; or One or more expression products of one or more of IV.I.~III., such as the first and second direct repeat sequences, CRISPR enzymes; Upon transcription, the first and second guide sequences each induce sequence-specific binding of the first and second CRISPR complexes to the first and second target sequences, where the first CRISPR complex comprises a CRISPR enzyme complexed with a first guide sequence that hybridizes to the first target sequence, and the second CRISPR complex comprises a CRISPR enzyme complexed with a second guide sequence that hybridizes to the second target sequence, the polynucleotide sequence encoding the CRISPR enzyme is DNA or RNA, and the first guide sequence induces cleavage of one strand of the DNA duplex near the first target sequence and the second guide sequence induces cleavage of the other strand near the second target sequence to create a double-strand break, thereby modifying a biological or non-human organism; including the step of delivering by contacting a particle comprising a non-naturally occurring or engineered composition as described above to HSCs, the method of modifying a biological or non-human organism by manipulation of first and second target sequences on the reverse strand of a DNA duplex at a target genomic locus of interest in HSCs, such as a target genomic locus associated with abnormal protein expression or a mutation associated with a disease condition or state; and optionally, the method can also include the step of delivering an HDR template, for example, by contacting a particle containing the HDR template with the HSCs or by contacting the HSCs with another particle containing the HDR template, where the HDR template results in the expression of a normal or minimally abnormal protein; "normal" is with respect to the wild type, and "abnormal" can be protein expression that causes a pathological or disease state; and optionally the method includes the steps of isolating or obtaining HSCs from a biological or non-human organism, optionally expanding the population of HSCs, performing contact of the one or more particles with the HSCs to obtain a modified population of HSCs, optionally expanding the population of modified HSCs, and optionally administering the modified HSCs to a biological or non-human organism, the method comprising manipulation of target sequences within coding, non-coding, or regulatory elements of said genomic locus.In some methods of the present invention, a polynucleotide sequence encoding a CRISPR enzyme, the first and second guide sequences, and some or all of the first and second direct repeat sequences are RNA. In further embodiments of the present invention, the sequence encoding the CRISPR enzyme, the first and second guide sequences, and the polynucleotide encoding the first and second direct repeat sequences are RNA and are delivered by liposomes, nanoparticles, exosomes, microvesicles, or gene guns; however, delivery by particles is advantageous. In certain embodiments of the present invention, the first and second direct repeats share 100% identity. In some embodiments, the polynucleotide may be included within a vector system comprising one or more vectors. In a preferred embodiment, the first CRISPR enzyme has one or more mutations such that the enzyme becomes a complementary strand nicking enzyme, and the second CRISPR enzyme has one or more mutations such that the enzyme becomes a non-complementary strand nicking enzyme. Alternatively, the first enzyme may be a non-complementary strand nicking enzyme and the second enzyme may be a complementary strand nicking enzyme. In a preferred method of the present invention, a 5' overhang is generated by the first guide sequence inducing cleavage of one strand of the DNA duplex near the first target sequence and the second guide sequence inducing cleavage of the other strand near the second target sequence. In embodiments of the present invention, the 5' overhang is at most 200 base pairs, preferably at most 100 base pairs, or more preferably at most 50 base pairs. In embodiments of the present invention, the 5' overhang is at least 26 base pairs, preferably at least 30 base pairs, or more preferably 34 - 50 base pairs.
[0197] In some embodiments, the present invention is, for example, I. A first regulatory element, (a) A first guide sequence capable of hybridizing to a first target sequence, and (b) At least one or more direct repeat sequences A first regulatory element operably linked to II. A second regulatory element, (a) A second guide sequence capable of hybridizing to a second target sequence, and (b) At least one or more direct repeat sequences A second regulatory element functionally linked to III. A third regulatory element functionally linked to an enzyme coding sequence encoding a CRISPR enzyme (e.g., Cpf1), and V. One or more expression products of I. to IV., such as the first and second direct repeat sequences, CRISPR enzyme; When transcribed, components I, II, III, and IV are located on the same or different vectors of the system, and the first and second guide sequences induce sequence-specific binding of the first and second CRISPR complexes to the first and second target sequences, respectively. The first CRISPR complex comprises a CRISPR enzyme complexed with a first guide sequence that hybridizes to the first target sequence, and the second CRISPR complex comprises a CRISPR enzyme complexed with a second guide sequence that hybridizes to the second target sequence. The polynucleotide sequence encoding the CRISPR enzyme is DNA or RNA, and the first guide sequence induces cleavage of one strand of the DNA duplex near the first target sequence, and the second guide sequence induces cleavage of the other strand near the second target sequence to produce a double-strand break, thereby modifying a biological or non-human organism. A step of delivering one or more particles comprising a non-naturally occurring or engineered composition as described above to HSCs by contacting them, for example, at a target genomic locus of HSCs, a method of modifying a biological or non-human organism by manipulating the first and second target sequences located on the reverse strand of the DNA duplex at the target genomic locus associated with abnormal protein expression or a mutation associated with a disease condition or state is included; and the method can also optionally include a step of delivering an HDR template, for example, by contacting the HSCs with particles containing the HDR template or by contacting the HSCs with another particle containing the HDR template, where the HDR template results in the expression of a normal or minimally abnormal protein; "normal" is with respect to the wild type, and "abnormal" can be protein expression that causes a pathological or disease state; and optionally the method can include a step of isolating or obtaining HSCs from a biological or non-human organism, optionally a step of expanding this HSC population, a step of performing contact between one or more particles and the HSCs to obtain a modified HSC population, optionally a step of expanding the population of modified HSCs, and optionally a step of administering the modified HSCs to a biological or non-human organism.
[0198] The present invention also provides a vector system as described herein. This system can include one, two, three, or four different vectors. Thus, components I, II, III, and IV may be located on one, two, three, or four different vectors, and all possible combinations of the positions of the components are envisioned herein, for example: in all combinations of positions envisioned, components I, II, III, and IV may be located on the same vector; components I, II, III, and IV may each be located on a different vector; components I, II, III, and IV may be located on a total of two or three different vectors, etc. In some methods of the present invention, a polynucleotide sequence encoding a CRISPR enzyme, the first and second guide sequences, and some or all of the first and second direct repeat sequences are RNA. In a further embodiment of the present invention, the first and second direct repeat sequences share 100% identity. In a preferred embodiment, the first CRISPR enzyme has one or more mutations such that the enzyme is a complementary strand nicking enzyme, and the second CRISPR enzyme has one or more mutations such that the enzyme is a non-complementary strand nicking enzyme. Alternatively, the first enzyme may be a non-complementary strand nicking enzyme, and the second enzyme may be a complementary strand nicking enzyme. In a further embodiment of the present invention, one or more of the viral vectors are delivered by liposomes, nanoparticles, exosomes, microvesicles, or a gene gun; however, particle delivery is advantageous.
[0199] In a preferred method of the present invention, a 5' overhang is generated by the first guide sequence inducing cleavage of one strand of the DNA duplex in the vicinity of the first target sequence and the second guide sequence inducing cleavage of the other strand in the vicinity of the second target sequence. In an embodiment of the present invention, the 5' overhang is at most 200 base pairs, preferably at most 100 base pairs, or more preferably at most 50 base pairs. In an embodiment of the present invention, the 5' overhang is at least 26 base pairs, preferably at least 30 base pairs, or more preferably 34 to 50 base pairs.
[0200] The present invention also provides in vitro or ex vivo cells comprising any of the modified CRISPR enzymes, compositions, systems or complexes according to or by any of the methods described above. Such cells may be eukaryotic or prokaryotic cells. The present invention also provides the progeny of such cells. The present invention also provides the product of any such cell or any such progeny, which product is the product of said one or more target loci as modified by the modified CRISPR enzyme of the CRISPR complex. Such product may be a peptide, polypeptide or protein. Some such products may be modified by the modified CRISPR enzyme of the CRISPR complex. In some such modified products, the product of the target locus is physically different from the product of said target locus that has not been modified by said modified CRISPR enzyme.
[0201] The present invention also provides a polynucleotide molecule comprising a polynucleotide sequence encoding any of the non-naturally occurring CRISPR enzymes described above.
[0202] Any such polynucleotide may further comprise one or more regulatory elements operably linked to the polynucleotide sequence encoding the non-naturally occurring CRISPR enzyme.
[0203] In any such polynucleotide comprising one or more regulatory elements, the one or more regulatory elements may be operably configured for the expression of the non-naturally occurring CRISPR enzyme in eukaryotic cells.
[0204] In any such polynucleotide comprising one or more regulatory elements, the one or more regulatory elements may be operably configured for the expression of the non-naturally occurring CRISPR enzyme in prokaryotic cells.
[0205] In any such polynucleotide comprising one or more regulatory elements, the one or more regulatory elements may be operably configured for the expression of the non-naturally occurring CRISPR enzyme in an in vitro system.
[0206] The present invention also provides an expression vector comprising any of the above-described polynucleotide molecules. The present invention also provides one or more such polynucleotide molecules, such as one or more polynucleotide molecules operably configured to express one or more proteins and / or nucleic acid components, and one or more such vectors.
[0207] The present invention further provides a method of creating a mutation in Cas (e.g., Cpf1) or a mutated or modified Cas (e.g., Cpf1) that is an ortholog of the CRISPR enzyme according to the present invention described herein, which comprises determining one or more amino acids in said ortholog that may be in proximity to or in contact with a nucleic acid molecule, such as DNA, RNA, gRNA, etc., for modification and / or mutation, and / or one or more amino acids similar or corresponding to one or more amino acids identified herein in the CRISPR enzyme according to the present invention described herein, and synthesizing, or preparing, or expressing an ortholog comprising one or more modifications and / or one or more mutations, or consisting of, or essentially consisting of, or mutating a neutral amino acid to a charged amino acid, such as a positively charged amino acid, such as alanine, as discussed herein, e.g., modifying, e.g., changing, or mutating. The ortholog thus modified can be used in a CRISPR-Cas system; and one or more nucleic acid molecules expressing it can be used in a vector or other delivery system that delivers the molecule or encodes a CRISPR-Cas system component as discussed herein.
[0208] In certain embodiments, the present invention provides efficient on-target activity and minimizes off-target activity. In certain embodiments, the present invention provides efficient on-target cleavage by a CRISPR protein and minimizes off-target cleavage by the CRISPR protein. In certain embodiments, the present invention provides guide-specific binding of a CRISPR protein at a locus without DNA cleavage. In certain embodiments, the present invention provides efficient on-target binding directed by a guide of a CRISPR protein at a locus and minimizes off-target binding of the CRISPR protein. Thus, in certain embodiments, the present invention provides target-specific gene regulation. In certain embodiments, the present invention provides guide-specific binding of a CRISPR enzyme at a locus without DNA cleavage. Thus, in certain embodiments, the present invention provides cleavage at one locus using a single CRISPR enzyme and provides gene regulation at another locus. In certain embodiments, the present invention provides orthogonal activation and / or inhibition and / or cleavage of multiple targets using one or more CRISPR proteins and / or enzymes.
[0209] In another aspect, the invention provides a method for functional screening of a gene in a genome in ex vivo or in vivo cell pools, the method comprising administration or expression of a library comprising a plurality of CRISPR-Cas system guide RNAs (gRNAs), wherein the screening further comprises use of a CRISPR enzyme and the CRISPR complex is modified to comprise a heterologous functional domain. In certain aspects, the invention provides a method for screening a genome comprising administration of the library to a host or in vivo expression in a host. In certain aspects, the invention provides a method as contemplated herein further comprising an activator that is administered to the host or expressed in the host. In certain aspects, the invention provides a method as contemplated herein where the activator is added to the CRISPR protein. In certain aspects, the invention provides a method as contemplated herein where the activator is added to the N-terminus or C-terminus of the CRISPR protein. In certain aspects, the invention provides a method as contemplated herein where the activator is added to the gRNA loop. In certain aspects, the invention provides a method as contemplated herein further comprising a repressor that is administered to the host or expressed in the host. In certain aspects, the invention provides a method as contemplated herein where the screening comprises affecting and detecting gene activation, gene inhibition, or cleavage at a locus.
[0210] In certain aspects, the invention provides a method as contemplated herein comprising delivery of a CRISPR-Cas complex or one or more of its components or one or more nucleic acid molecules encoding the same, wherein the one or more nucleic acid molecules are operably linked to one or more regulatory sequences and are expressed in vivo. In certain aspects, the invention provides a method as contemplated herein where the in vivo expression is by lentivirus, adenovirus, or AAV. In certain aspects, the invention provides a method as contemplated herein where the delivery is by particle, nanoparticle, lipid, or cell permeable peptide (CPP).
[0211] In certain embodiments, it may be beneficial to target the CRISPR-Cas complex to the chloroplast. In many cases, this targeting can be achieved by the presence of an N-terminal extension called a chloroplast transit peptide (CTP) or plastid transit peptide. When an expressed polypeptide is to be compartmentalized into a plant plastid (e.g., chloroplast), a chromosomally-transgenic gene from a bacterial source must have a sequence encoding a CTP sequence fused to the sequence encoding the expressed polypeptide. Thus, in many cases, the localization of an exogenous polypeptide to the chloroplast is achieved by operably linking a polynucleotide sequence encoding a CTP sequence to the 5' region of the polynucleotide encoding the exogenous polypeptide. The CTP is removed at a processing step during translocation to the plastid. However, the processing efficiency can be affected by the amino acid sequence of the CTP and the sequence in the vicinity at the NH2 terminus of the peptide. Other described options for targeting chloroplasts are the maize cab-m7 signal sequence (U.S. Patent No. 7,022,896, WO 97 / 41228), the pea glutathione reductase signal sequence (WO 97 / 41228), and the CTP described in U.S. Patent Application Publication No. 2009029861.
[0212] In one aspect, the invention provides a library, method, or complex as contemplated herein, where the gRNA is modified to have at least one non-coding functional loop, e.g., at least one non-coding functional loop is inhibitory; e.g., at least one non-coding functional loop contains an Alu.
[0213] In one aspect, the present invention provides a method for changing or modifying the expression of a gene product. The method can include introducing into a cell containing and expressing a DNA molecule encoding a gene product an engineered, non-naturally occurring CRISPR-Cas system comprising a Cas protein and a guide RNA that targets the DNA molecule, whereby the guide RNA targets the DNA molecule encoding the gene product and the Cas protein cleaves the DNA molecule encoding the gene product, whereby the expression of the gene product is changed; and the Cas protein and the guide RNA do not naturally occur together. The present invention further encompasses a Cas protein codon-optimized for expression in eukaryotic cells. In a preferred embodiment, the eukaryotic cell is a mammalian cell, and in a more preferred embodiment, the mammalian cell is a human cell. In a further embodiment of the present invention, the expression of the gene product is decreased.
[0214] In certain aspects, the present invention provides altered cells and the progeny of those cells, as well as products made by said cells. The CRISPR-Cas (e.g., Cpf1) proteins and systems of the present invention are used to generate cells containing modified target loci. In some embodiments, the method can include binding a nucleic acid targeting complex to a target DNA or RNA to effect cleavage of the target DNA or RNA, thereby modifying the target DNA or RNA, where the nucleic acid targeting complex comprises a nucleic acid targeting effector protein complexed with a guide RNA that hybridizes to a target sequence within the target DNA or RNA. In one aspect, the present invention provides a method for repairing a locus in a cell. In another aspect, the present invention provides a method for modifying the expression of DNA or RNA in a eukaryotic cell. In some embodiments, the method includes binding a nucleic acid targeting complex to DNA or RNA such that the binding results in an increase or decrease in the expression of the DNA or RNA; where the nucleic acid targeting complex comprises a nucleic acid targeting effector protein complexed with a guide RNA. Similar considerations and conditions apply to the method for modifying a target DNA or RNA as described above. In fact, these options for sample collection, culture, and reintroduction apply to all aspects of the present invention. In certain aspects, the present invention provides a method for modifying a target DNA or RNA in a eukaryotic cell, which may be in vivo, ex vivo, or in vitro. In some embodiments, the method includes sampling a cell or cell population from a human or non-human animal and modifying one or more cells. Culture can be performed ex vivo at any stage. Such cells can be, without limitation, plant cells, animal cells, specific cell types of any organism, e.g., stem cells, immune cells, T cells, B cells, dendritic cells, cardiovascular cells, epithelial cells, stem cells, etc. When modified by the present invention, the cells can produce gene products, for example, in a controlled amount, which may be increased or decreased depending on the use, and / or mutated. In certain embodiments, the locus of the cell is repaired. The one or more cells may further be reintroduced into a non-human animal or plant.For the cells to be re-introduced, the cells may preferably be stem cells.
[0215] In one aspect, the present invention provides a cell that transiently contains a CRISPR system or a component. For example, a CRISPR protein or enzyme and a nucleic acid are transiently provided to the cell, and when the locus changes, subsequently the amount of one or more components of the CRISPR system decreases. Subsequently, the cell, the progeny of the cell, and the organism containing the cell that have undergone gene changes mediated by CRISPR contain the one or more CRISPR system components in a reduced amount or no longer contain the one or more CRISPR system components. One non-limiting example is the self-inactivating CRISPR-Cas system as further described herein. Accordingly, the present invention provides a cell, an organism, and the progeny of the cell and the organism that contain one or more loci changed by the CRISPR-Cas system but are essentially lacking one or more CRISPR system components. In certain embodiments, the CRISPR system components are substantially absent. Such cells, tissues, and organisms advantageously contain the desired or selected gene changes but have lost the CRISPR-Cas components or remnants thereof that may potentially act non-specifically, lead to safety concerns, or prevent regulatory approval. Further, the present invention provides the products made by the cells, organisms, and the progeny of the cells and organisms.
[0216] Gene editing or alteration of a target locus by Cpf1 The double-strand break point or the single-strand break point in one of the strands should preferably be close enough to the target position for the modification to occur. In certain embodiments, the distance is 50, 100, 200, 300, 350 or 400 nucleotides or less. Without wishing to be bound by theory, it is believed that the break point should be close enough to the target position and there should be no break point within the region that undergoes exonuclease-mediated removal during end resection. Since the template nucleic acid sequence can only be used for the modification of the sequence within the end resection region, if the distance between the target position and the break point is too large, the end resection may not contain mutations and thus may not be modified.
[0217] In embodiments where a guide RNA and a Cpf1 nuclease induce a double-strand break for the purpose of inducing HDR-mediated modification, the cleavage site is 0 to 200 bp (e.g., 0 to 175, 0 to 150, 0 to 125, 0 to 100, 0 to 75, 0 to 50, 0 to 25, 25 to 200, 25 to 175, 25 to 150, 25 to 125, 25 to 100, 25 to 75, 25 to 50, 50 to 200, 50 to 175, 50 to 150, 50 to 125, 50 to 100, 50 to 75, 75 to 200, 75 to 175, 75 to 150, 75 to 125, 75 to 100 bp) away from the target position. In certain embodiments, the cleavage site is 0 to 100 bp (e.g., 0 to 75, 0 to 50, 0 to 25, 25 to 100, 25 to 75, 25 to 50, 50 to 100, 50 to 75 or 75 to 100 bp) away from the target position. In a further embodiment, multiplexed cleavage can be induced using two or more guide RNAs complexed with Cpf1 or its ortholog or homolog to induce HDR-mediated modification.
[0218] The homology arm must extend at least to the extent of the region where at least terminal resection can occur, such that, for example, a resected single-stranded overhang can find a complementary region within the donor template. The full length may be restricted by parameters such as plasmid size or viral packaging limitations. In certain embodiments, the homology arm may not extend into the repetitive element. Exemplary homology arm lengths include at least 50, 100, 250, 500, 750, or 1000 nucleotides.
[0219] The target position, as used herein, refers to a site on a target nucleic acid or target gene (e.g., chromosome) that is modified by a Cpf1 molecule-dependent process. For example, the target position can be a modification, e.g., a correction, of the target nucleic acid by cleavage of the modifying Cpf1 molecule and induction of the template nucleic acid at the target position. In certain embodiments, the target position can be a site between two nucleotides on the target nucleic acid where one or more nucleotides are added, e.g., between adjacent nucleotides. The target position can include one or more nucleotides that are changed, e.g., corrected, by the template nucleic acid. In certain embodiments, the target position is within the range of the target sequence (e.g., the sequence to which the guide RNA binds). In certain embodiments, the target position is upstream or downstream of the target sequence (e.g., the sequence to which the guide RNA binds).
[0220] The template nucleic acid, as the term is used herein, refers to a nucleic acid sequence that can be used in combination with a Cpf1 molecule and a guide RNA molecule to change the structure of the target position. In certain embodiments, the target nucleic acid is modified to typically have a part or all of the sequence of the template nucleic acid at or near one or more cleavage sites. In certain embodiments, the template nucleic acid is single-stranded. In alternative embodiments, the template nucleic acid (nuceic acid) is double-stranded. In certain embodiments, the template nucleic acid is DNA, e.g., double-stranded DNA. In alternative embodiments, the template nucleic acid is single-stranded DNA.
[0221] In certain embodiments, the template nucleic acid alters the structure of the target locus by participating in homologous recombination. In certain embodiments, the template nucleic acid alters the sequence of the target locus. In certain embodiments, the template nucleic acid results in the incorporation of modified or non-naturally occurring bases into the target nucleic acid.
[0222] The template sequence can undergo recombination with the target sequence mediated or catalyzed by cleavage. In certain embodiments, the template nucleic acid can include a sequence corresponding to a site on the target sequence that is cleaved by a Cpf1-mediated cleavage event. In certain embodiments, the template nucleic acid can include sequences corresponding to both a first site on the target sequence cleaved in a first Cpf1-mediated event and a second site on the target sequence cleaved in a second Cpf1-mediated event.
[0223] In certain embodiments, the template nucleic acid can include sequences that result in changes to the coding sequence of the translated sequence, such as substitution of one amino acid for another in a protein product, e.g., conversion of a mutant allele to a wild-type allele, conversion of a wild-type allele to a mutant allele, and / or introduction of a stop codon, insertion of an amino acid residue, deletion of an amino acid residue, or nonsense mutation. In certain embodiments, the template nucleic acid can include sequences that result in changes to non-coding sequences, such as changes in exons or 5' or 3' untranslated or non-transcribed regions. Such changes include changes in regulatory elements, e.g., promoters, enhancers, and changes in cis-acting or trans-acting regulatory elements.
[0224] The structure of the target sequence may be altered using a template nucleic acid having homology to the target position of the target gene. The template sequence can be used to alter undesirable structures, such as undesirable or mutated nucleotides. The template nucleic acid, when incorporated, can result in a decrease in the activity of a positive control element; an increase in the activity of a positive control element; a decrease in the activity of a negative control element; an increase in the activity of a negative control element; a decrease in gene expression; an increase in gene expression; an increase in resistance to a disorder or disease; an increase in resistance to virus entry; correction of a mutation or change of an undesirable amino acid residue, imparting, increasing, abolishing or decreasing a biological property of a gene product, such as an increase in the enzymatic activity of an enzyme, or an increase in the ability of a gene product to interact with another molecule.
[0225] The template nucleic acid can include a sequence that results in an alteration of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 nucleotides or more of the target sequence. In certain embodiments, the template nucleic acid can be 20±10, 30±10, 40±10, 50±10, 60±10, 70±10, 80±10, 90±10, 100±10, 110±10, 120±10, 130±10, 140±10, 150±10, 160±10, 170±10, 180±10, 190±10, 200±10, 210±10, or 220±10 nucleotides in length. In certain embodiments, the template nucleic acid can be 30±20, 40±20, 50±20, 60±20, 70±20, 80±20, 90±20, 100±20, 110±20, 120±20, 130±20, 140±20, 150±20, 160±20, 170±20, 180±20, 190±20, 200±20, 210±20, or 220±20 nucleotides in length. In certain embodiments, the template nucleic acid is 10 - 1,000, 20 - 900, 30 - 800, 40 - 700, 50 - 600, 50 - 500, 50 - 400, 50 - 300, 50 - 200, or 50 - 100 nucleotides in length.
[0226] The template nucleic acid comprises the following components: [5' homology arm]-[substitution sequence]-[3' homology arm]. The homology arms result in recombination into the chromosome and thus substitution of the substitution sequence with an undesirable element, such as a mutation or signature. In certain embodiments, the homology arms flank the most distal cleavage site. In certain embodiments, the 3' end of the 5' homology arm is adjacent to the 5' end of the substitution sequence. In certain embodiments, the 5' homology arm can extend at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, or 2000 nucleotides 5' of the 5' end of the substitution sequence. In certain embodiments, the 5' end of the 3' homology arm is adjacent to the 3' end of the substitution sequence. In certain embodiments, the 3' homology arm can extend at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, or 2000 nucleotides 3' of the 3' end of the substitution sequence.
[0227] In certain embodiments, one or both of the homology arms can be shortened to avoid including certain sequence repeat elements. For example, the 5' homology arm may be shortened to avoid sequence repeat elements. In other embodiments, the 3' homology arm may be shortened to avoid sequence repeat elements. In some embodiments, both the 5' and 3' homology arms may be shortened to avoid including certain sequence repeat elements.
[0228] In certain embodiments, the template nucleic acid for correcting a mutation can be designed to be used as a single-stranded oligonucleotide. When using a single-stranded oligonucleotide, the 5' and 3' homology arms can range in length from about 200 base pairs (bp), e.g., at least 25, 50, 75, 100, 125, 150, 175, or 200 bp in length.
[0229] The Cpf1 effector protein complex can deliver a functional effector Unlike CRISPR-Cas mediated gene knockout, which permanently abolishes expression by mutating genes at the DNA level, CRISPR-Cas knockdown can transiently reduce gene expression using artificial transcription factors. Mutating key residues in both DNA cleavage domains of the Cpf1 protein, such as D908A, E993A, D1263A in the AsCpf1 protein, generates catalytically inactive Cpf1. Catalytically inactive Cpf1 forms a complex with the guide RNA and localizes to the DNA sequence specified by the targeting domain of the guide RNA, however, this does not cleave the target DNA. Fusing an inactive Cpf1 protein, such as the AsCpf1 protein, to an effector domain, such as a transcriptional repression domain, enables recruitment of the effector to any DNA site specified by the guide RNA. In certain embodiments, Cpf1 may be fused to a transcriptional repression domain and recruited to the promoter region of the gene. In particular for gene repression, it is contemplated herein that blocking the binding site of endogenous transcription factors can aid in downregulating gene expression. In another embodiment, inactive Cpf1 may be fused to a chromatin modifying protein. A change in chromatin state can result in a decrease in the expression of the target gene.
[0230] In certain embodiments, the guide RNA molecule can target known transcriptional response elements (e.g., promoters, enhancers, etc.), known upstream activation sequences, and / or sequences of unknown or known function suspected of being capable of controlling the expression of the target DNA.
[0231] In some methods, modification of expression in a cell can be effected by inactivating a target polynucleotide. For example, when a CRISPR complex binds to a target sequence in a cell, the target polynucleotide is inactivated such that the sequence is not transcribed, or the encoded protein is not produced, or the sequence does not function as a wild-type sequence. For example, a protein or microRNA coding sequence can be inactivated so that no protein is produced.
[0232] In certain embodiments, the CRISPR enzyme comprises one or more mutations selected from the group consisting of D917A, E1006A, and D1225A, and / or one or more mutations are in the RuvC domain of the CRISPR enzyme, or are otherwise mutations as discussed herein. In some embodiments, the CRISPR enzyme has one or more mutations in the catalytic domain, whereupon transcription, the direct repeat sequences form a single stem-loop, and the guide sequence directs sequence-specific binding of the CRISPR complex to the target sequence, and where the enzyme further comprises a functional domain. In some embodiments, the functional domain is a transcriptional activation domain, preferably VP64. In some embodiments, the functional domain is a transcriptional repression domain, preferably KRAB. In some embodiments, the transcriptional repression domain is SID, or a concatemer of SID (e.g., SID4X). In some embodiments, the functional domain is an epigenetically modified domain, and an epigenetically modified enzyme is provided. In some embodiments, the functional domain is an activation domain, which may be a P65 activation domain.
[0233] Delivery of a Cpf1 effector protein complex or a component thereof From the present disclosure and knowledge in the art, the CRISPR-Cas system, particularly the novel CRISPR systems described herein, or a component thereof or a nucleic acid molecule thereof (e.g., including an HDR template) or a nucleic acid molecule encoding or providing a component thereof can be delivered by the delivery systems described herein, either generally or in detail.
[0234] Vector delivery, e.g., plasmid, viral delivery: CRISPR enzymes, e.g., Cpf1, and / or any such RNA, e.g., guide RNA, can be delivered using any suitable vector, e.g., a plasmid or viral vector, e.g., adeno-associated virus (AAV), lentivirus, adenovirus, or other types of viral vectors, or combinations thereof. Cpf1 and one or more guide RNAs can be packaged into one or more vectors, e.g., plasmid or viral vectors. In some embodiments, the vector, e.g., plasmid or viral vector, is delivered to the target tissue, e.g., by intramuscular injection, while delivery can also be by intravenous, transdermal, intranasal, oral, mucosal, or other delivery methods. Such delivery can be by single administration or by multiple administrations. Those skilled in the art will understand that the actual dosage delivered herein can vary widely depending on various factors such as the choice of vector, target cell, organism, or tissue, the general condition of the subject being treated, the degree of transformation / modification sought, the route of administration, the mode of administration, the type of transformation / modification sought, etc.
[0235] Such dosages may further include carriers (such as water, physiological saline, ethanol, glycerol, lactose, sucrose, calcium phosphate, gelatin, dextran, agar, pectin, peanut oil, sesame oil, etc.), diluents, pharmaceutically acceptable carriers (such as phosphate buffered saline), pharmaceutically acceptable excipients, and / or other compounds known in the art. The dosage may further include one or more pharmaceutically acceptable salts, such as mineral salts like hydrochloride, hydrobromide, phosphate, sulfate, etc.; and organic acid salts like acetate, propionate, malonate, benzoate, etc. Additionally, auxiliary substances such as wetting agents or emulsifiers, pH buffering substances, gels or gelling substances, fragrances, colorants, microspheres, polymers, suspending agents, etc. may also be present therein. In addition, one or more other conventional pharmaceutical ingredients, such as preservatives, wetting agents, suspending agents, surfactants, antioxidants, anti-caking agents, fillers, chelating agents, coating agents, chemical stabilizers, etc. may also be present, especially when the dosage form is in a reconstitutable form. Exemplary suitable components include microcrystalline cellulose, sodium carboxymethyl cellulose, polysorbate 80, phenylethyl alcohol, chlorobutanol, potassium sorbate, sorbic acid, sulfur dioxide, propyl gallate, parabens, ethyl vanillin, glycerin, phenol, parachlorophenol, gelatin, albumin, and combinations thereof. A thorough discussion of pharmaceutically acceptable excipients is available in REMINGTON’S PHARMACEUTICAL SCIENCES (Mack Pub. Co., N.J. 1991) (incorporated herein by reference).
[0236] In certain embodiments herein, delivery is via an adenovirus, which can be a single - booster dosage containing at least 1×105 particles (also referred to as particle units, pu) of an adenovirus vector. In certain embodiments herein, the dosage is preferably at least about 1×10 6 particles (e.g., about 1×10 6 ~1×10 12 particles), more preferably at least about 1×10 7 particles, more preferably at least about 1×108 Particles (e.g., about 1×10 8 ~1×10 11 particles or about 1×10 8 ~1×10 12 particles), and most preferably at least about 1×10 0 particles (e.g., about 1×10 9 ~1×10 10 particles or about 1×10 9 ~1×10 12 particles), or even more preferably at least about 1×10 10 particles (e.g., about 1×10 10 ~1×10 12 particles) of an adenovirus vector. Alternatively, the dose is about 1×10 14 particles or less, preferably about 1×10 13 particles or less, even more preferably about 1×10 12 particles or less, even more preferably about 1×10 11 particles or less, and most preferably about 1×10 10 particles or less (e.g., about 1×10 9 articles or less). Thus, the dose is, for example, about 1×10 6 particle units (pu), about 2×10 6 pu, about 4×10 6 pu, about 1×10 7 pu, about 2×10 7 pu, about 4×10 7 pu, about 1×10 8 pu, about 2×10 8 pu, about 4×10 8 pu, about 1×10 9 pu, about 2×10 9 pu, about 4×10 9 pu, about 1×10 10 pu, about 2×10 10 pu, about 4×10 10 pu, about 1×10 11 pu, about 2×10 11 pu, about 4×10 11 pu, about 1×10 12 pu, about 2×10 12 pu, or about 4×10 12A single dose of an adenovirus vector may contain a pu adenovirus vector. For example, refer to the adenovirus vector of U.S. Patent No. 8,454,972 B2 to Nabel, et.al. granted on June 4, 2013 (incorporated herein by reference), and the dosage amounts in column 29, lines 36 - 58 thereof. In certain embodiments herein, the adenovirus is delivered in multiple doses.
[0237] In certain embodiments herein, the delivery is via AAV. A therapeutically effective dosage for in vivo delivery of AAV to humans is considered to be in the range of about 20 - about 50 ml of physiological saline containing about 1×10 10 ~ about 1×10 10 functional AAV / ml solution. The dosage can be adjusted to balance the therapeutic benefit and any side effects. In certain embodiments herein, the AAV dosage is generally in the range of about 1×10 5 ~ 1×10 50 genomic AAV, about 1×10 8 ~ 1×10 20 genomic AAV, about 1×10 10 ~ about 1×10 16 genome, or about 1×10 11 ~ about 1×10 16 genomic AAV. The human dosage may be about 1×10 13 genomic AAV. Such concentrations can be delivered in a carrier solution of about 0.001 ml - about 100 ml, about 0.05 - about 50 ml, or about 10 - about 25 ml. Other effective dosages can be readily established by one of ordinary skill in the art using routine testing to create a dose - response curve. For example, refer to U.S. Patent No. 8,404,658 B2 to Hajjar, et al. granted on March 26, 2013, column 27, lines 45 - 60.
[0238] In certain embodiments of the present specification, delivery is via a plasmid. In such plasmid compositions, the dosage must be an amount of plasmid sufficient to induce a response. For example, suitable amounts of plasmid DNA in the plasmid composition can be from about 0.1 to about 2 mg, or from about 1 μg to about 10 μg, per 70 kg of an individual. The plasmids of the present invention generally can include (i) a promoter; (ii) a sequence encoding a CRISPR enzyme operably linked to the promoter; (iii) a selectable marker; (iv) an origin of replication; and (v) a transcription terminator operably linked to (ii) downstream of (ii). The plasmid can also encode the RNA components of the CRISPR complex, although one or more of these can alternatively be encoded on a different vector.
[0239] The dosages herein are based on an individual of average 70 kg. The frequency of administration is within the discretion of a medical or veterinary practitioner (e.g., a physician, veterinarian), or a scientist in the art. It is also noted that the mice used in experiments are typically about 20 g and that scaling up from mouse experiments to a 70 kg individual is possible.
[0240] In some embodiments, the RNA molecules of the present invention are delivered in liposomal formulations or lipofectin formulations and can be prepared by methods well known to those skilled in the art. Such methods are described, for example, in U.S. Patent Nos. 5,593,972, 5,589,466, and 5,580,859, which are incorporated herein by reference. Delivery systems have been developed specifically for enhancing and improving siRNA delivery to mammalian cells (see, for example, Shen et al FEBS Let. 2003, 539:111-114; Xia et al., Nat. Biotech. 2002, 20:1006-1010; Reich et al., Mol. Vision. 2003, 9:210-216; Sorensen et al., J. Mol. Biol. 2003, 327:761-766; Lewis et al., Nat. Gen. 2002, 32:107-108, and Simeoni et al., NAR 2003, 31, 11:2717-2724), and can be applied to the present invention. siRNA has recently been successfully used to suppress gene expression in primates (see, for example, Tolentino et al., Retina 24(4):660), which can also be applied to the present invention.
[0241] Indeed, RNA delivery is a useful method for in vivo delivery. It is possible to deliver Cpf1 and gRNA (and, for example, HR repair templates) intracellularly using liposomes or nanoparticles. Thus, the delivery of CRISPR enzymes such as Cpf1, and / or the delivery of the RNA of the present invention, can be via exosomes, liposomes, or one or more particles in RNA form. For example, Cpf1 mRNA and gRNA can be packaged into liposomal particles for in vivo delivery. Liposomal transfection reagents, such as Lipofectamine from Life Technologies and other commercially available reagents, can effectively deliver RNA molecules to the liver.
[0242] Similarly preferred means of delivering RNA also include delivery by RNA particles or particles (Cho, S., Goldberg, M., Son, S., Xu, Q., Yang, F., Mei, Y., Bogatyrev, S., Langer, R. and Anderson, D., "Lipid-like nanoparticles for small interfering RNA delivery to endothelial cells", Advanced Functional Materials, 19: 3112-3118, 2010) or delivery by exosomes (Schroeder, A., Levins, C., Cortez, C., Langer, R., and Anderson, D., "Lipid-based nanotherapeutics for siRNA delivery", Journal of Internal Medicine, 267: 9-21, 2010, PMID: 20059641). In fact, exosomes have been shown to be particularly useful for the delivery of siRNA, a system that has some similarity to the CRISPR system. For example, El-Andaloussi S, et al. ("Exosome-mediated delivery of siRNA in vitro and in vivo", Nat Protoc. 2012 Dec; 7(12): 2112-26. doi: 10.1038 / nprot.2012.131. Epub 2012 Nov 15) describes how exosomes are a promising tool for drug delivery across various biological barriers and can be used for in vitro and in vivo delivery of siRNA. Their approach involves creating target exosomes containing exosome proteins fused to peptide ligands by transfection of an expression vector. The exosomes are then purified from the transfected cell supernatant, characterized, and then loaded with RNA. The delivery or administration according to the present invention can be carried out using exosomes, and specifically, although not limited to, to the brain.Vitamin E (α-tocopherol) can be conjugated to CRISPR Cas and delivered to the brain together with high-density lipoprotein (HDL) in the same manner as the delivery of small interfering RNA (siRNA) to the brain by, for example, Uno et al. (HUMAN GENE THERAPY 22:711-719 (June 2011)). Mice were injected with an osmotic minipump (model 1007D; Alzet, Cupertino, CA) filled with phosphate-buffered saline (PBS) or free TocsiBACE or Toc-siBACE / HDL and connected to a Brain Infusion Kit 3 (Alzet). For injection into the dorsal third ventricle, a brain infusion cannula was placed approximately 0.5 mm posterior to bregma on the midline. Uno et al. found that as little as 3 nmol of Toc-siRNA containing HDL could induce a similar degree of target reduction by the same ICV injection method. A similar dosage of CRISPR Cas conjugated to α-tocopherol co-administered with HDL targeting the brain can be contemplated in humans in the present invention, for example, about 3 nmol to about 3 μmol of CRISPR Cas targeting the brain can be contemplated. Zou et al. ((HUMAN GENE THERAPY 22:465-475 (April 2011)) describe a lentiviral-mediated delivery method of short hairpin RNA targeting PKCγ for in vivo gene silencing in the rat spinal cord. Zou et al. used about 10 μl of recombinant lentivirus with a titer of 9 about 1×10 transduction units (TU) / ml administered by a subarachnoid catheter. A similar dosage of CRISPR Cas expressed in a lentiviral vector targeting the brain can be contemplated in humans in the present invention, for example, about 10 to 50 ml of CRISPR Cas targeting the brain in a lentivirus with a titer of 9 about 1×10 transduction units (TU) / ml can be contemplated.
[0243] Regarding local delivery to the brain, this can be achieved in various ways. For example, a substance can be delivered into the striatum, for example, by injection. The injection can be performed stereotactically at the beginning.
[0244] Enhancing NHEJ or HR efficiency can also assist in delivery. NHEJ efficiency is preferably enhanced by co-expression of a terminal processing enzyme such as Trex2 (Dumitrache et al. Genetics. 2011 August;188(4):787 - 797). HR efficiency is preferably increased by transiently inhibiting the NHEJ machinery such as Ku70 and Ku86. HR efficiency can also be increased by co-expression of prokaryotic or eukaryotic homologous recombination enzymes such as RecBCD and RecA.
[0245] Packaging and promoter To mediate genome modification in vivo, methods for packaging a nucleic acid molecule encoding Cpf1 of the present invention, for example DNA, into a vector, for example a viral vector, include the following: · To achieve NHEJ-mediated gene knockout: · Single viral vector: · Vector containing two or more expression cassettes: · Promoter - Cpf1-encoding nucleic acid molecule - terminator · Promoter - gRNA1 - terminator · Promoter - gRNA2 - terminator · Promoter - gRNA(N) - terminator (up to the vector size limit) · Double viral vector: · Vector 1 containing one expression cassette for driving the expression of Cpf1 · Promoter - Cpf1-encoding nucleic acid molecule - terminator · Vector 2 containing another expression cassette for driving the expression of one or more guide RNAs · Promoter - gRNA1 - terminator · Promoter - gRNA(N) - Terminator (up to the vector size limit) · To mediate homology - dependent repair · In addition to the single and double viral vector approaches described above, additional vectors can be used for the delivery of homology - dependent repair templates.
[0246] Promoters used to drive the expression of Cpf1 - encoding nucleic acid molecules may include the following: - AAV ITR can serve as a promoter: This is advantageous in that it eliminates the need for additional promoter elements (which can be located within the vector). The available additional space can be used to drive the expression of additional elements (such as gRNA). Also, since ITR activity is relatively weak, it can be used to reduce potential toxicity due to overexpression of Cpf1. - For ubiquitous expression, promoters that can be used include CMV, CAG, CBh, PGK, SV40, ferritin heavy chain or light chain, etc.
[0247] For expression in the brain or other CNS, promoters such as synapsin I for all neurons, CaMKIIα for excitatory neurons, GAD67 or GAD65 or VGAT for GABAergic neurons can be used.
[0248] For expression in the liver, the albumin promoter can be used.
[0249] For expression in the lung, SP - B can be used.
[0250] For endothelial cells, ICAM can be used.
[0251] For hematopoietic cells, IFNβ or CD45 can be used.
[0252] For osteoblasts, OG - 2 can be used.
[0253] Promoters used for driving guide RNAs may include the following: - Pol III promoters such as U6 or H1 - Use of a Pol II promoter and intron cassette for expressing gRNA.
[0254] Adeno-associated virus (AAV) Cpf1 and one or more guide RNAs can be delivered using adeno-associated virus (AAV), lentivirus, adenovirus or other plasmid or viral vector types, in particular, for example, using the formulations and dosages from U.S. Patent No. 8,454,972 (formulations, dosages for adenovirus), U.S. Patent No. 8,404,658 (formulations, dosages for AAV) and U.S. Patent No. 5,846,946 (formulations, dosages for DNA plasmids) as well as from clinical trials involving lentivirus, AAV and adenovirus and publications regarding such clinical trials. For example, for AAV, the route of administration, formulation and dosage may be as in U.S. Patent No. 8,454,972 and clinical trials involving AAV. For adenovirus, the route of administration, formulation and dosage may be as in U.S. Patent No. 8,404,658 and clinical trials involving adenovirus. For plasmid delivery, the route of administration, formulation and dosage may be as in U.S. Patent No. 5,846,946 and clinical trials involving plasmids. Dosages may be based on, or applied to, an average 70 kg individual (e.g., male adult human) and can be adjusted for patients, subjects, mammals of different weights and species. The frequency of administration is within the purview of medical or veterinary practitioners (e.g., physicians, veterinarians) depending on usual factors including the age, sex, general health, other conditions of the patient or subject as well as the particular condition or symptom being addressed. The viral vector can be injected into the tissue of interest. For cell type-specific genomic modification, the expression of Cpf1 can be driven by a cell type-specific promoter. For example, an albumin promoter may be used for liver-specific expression and a synapsin I promoter may be used for neuron-specific expression (e.g., for targeting CNS disorders).
[0255] Regarding in vivo delivery, AAV is advantageous for several reasons compared to other viral vectors: Low toxicity (which may result from a purification method that does not require ultracentrifugation of cellular particles that can activate the immune response) and A low probability of causing insertional mutagenesis because it is not integrated into the host genome.
[0256] AAV has a packaging limit of 4.5 or 4.75 Kb. This means that Cpf1 as well as the promoter and transcription terminator all need to fit within the same viral vector. Constructs larger than 4.5 or 4.75 Kb can lead to a significant decrease in virus production. SpCas9 is quite large and the gene itself exceeds 4.1 Kb, so it is difficult to package into AAV. Accordingly, embodiments of the present invention include utilizing shorter Cpf1 homologs.
[0257] Regarding AAV, AAV can be AAV1, AAV2, AAV5 or any combination thereof. One can select the AAV of the AAV related to the cell to be targeted; for example, for targeting the brain or nerve cells, one can select AAV serotype 1, 2, 5 or hybrid capsid AAV1, AAV2, AAV5 or any combination thereof; and for targeting heart tissue, one can select AAV4. AAV8 is useful for delivery to the liver. The promoters and vectors herein are individually preferred. A list of specific AAV serotypes for these cells is as follows (see Grimm, D. et al, J. Virol. 82:5887 - 5911 (2008)).
[0258]
Table 1
[0259] Lentivirus Lentivirus is a complex retrovirus that has the ability to infect and express its genes in both mitotic and post-mitotic cells. Most commonly, the known lentivirus is the human immunodeficiency virus (HIV), which targets a wide range of cell types using the envelope glycoproteins of other viruses.
[0260] Lentivirus can be prepared as follows. After cloning pCasES10 (which contains the lentivirus transfer plasmid backbone), low passage (p = 5) HEK293FT cells were seeded in a T-75 flask to 50% confluence and transfected the next day in DMEM containing 10% fetal bovine serum and no antibiotics. After 20 hours, the medium was replaced with OptiMEM (serum-free) medium and transfection was performed 4 hours later. The cells were transfected with 10 μg of the lentivirus transfer plasmid (pCasES10) and the following packaging plasmids: 5 μg of pMD2.G (VSV-g pseudotype), and 7.5 μg of psPAX2 (gag / pol / rev / tat). The transfection was performed in 4 mL of OptiMEM containing a cationic lipid delivery agent (50 μL Lipofectamine 2000 and 100 μL Plus reagent). After 6 hours, the medium was replaced with antibiotic-free DMEM containing 10% fetal bovine serum. Although serum is used in these methods during cell culture, a serum-free method is preferred.
[0261] Lentivirus can be purified as follows. After 48 hours, the viral supernatant was collected. First, debris was removed from the supernatant and filtered through a 0.45 μm low protein binding (PVDF) filter. Next, it was spun in an ultracentrifuge at 24,000 rpm for 2 hours. The viral pellet was resuspended in 50 μL of DMEM at 4 °C overnight. Next, it was aliquoted and immediately frozen at -80 °C.
[0262] In another embodiment, a minimal non-primate lentiviral vector based on equine infectious anemia virus (EIAV) is also contemplated, particularly for ocular gene therapy (see, e.g., Balagaan, J Gene Med 2006;8:275-285). In another embodiment, a lentiviral gene therapy vector, RetinoStat®, based on equine infectious anemia virus that expresses the angiogenesis inhibitor proteins endostatin and angiostatin delivered by subretinal injection for the treatment of exudative (wet form) age-related macular degeneration is also contemplated (see, e.g., Binley et al., HUMAN GENE THERAPY 23:980-991 (September 2012)), and this vector can be modified for use with the CRISPR-Cas system of the present invention.
[0263] In another embodiment, a self-inactivating lentiviral vector containing siRNA targeting a common exon shared by HIV tat / rev, a nucleosome-localized TAR decoy, and an anti-CCR5 specific hammerhead ribozyme (see, e.g., DiGiusto et al. (2010) Sci Transl Med 2:36ra43) may be used and / or adapted for the CRISPR-Cas system of the present invention. At least 2.5×10 6 CD34+ cells per kilogram of patient body weight are collected and pre-stimulated for 16-20 hours at a density of 2×10 6 cells / ml in X-VIVO 15 medium (Lonza) containing 2 μmol / L -glutamine, stem cell factor (100 ng / ml), Flt-3 ligand (Flt-3L) (100 ng / ml), and thrombopoietin (10 ng / ml) (CellGenix). The pre-stimulated cells can be transduced with lentivirus at a multiplicity of infection of 5 for 16-24 hours in a 75 cm2 tissue culture flask coated with fibronectin (25 mg / cm2) (RetroNectin, Takara Bio Inc.).
[0264] Lentiviral vectors are disclosed for the treatment of Parkinson's disease, for example, see U.S. Patent Application Publication No. 20120295960, U.S. Patent No. 7303910, and U.S. Patent No. 7351585. Lentiviral vectors are also disclosed for the treatment of eye diseases, for example, see U.S. Patent Application Publication No. 20060281180, 20090007284, U.S. Patent Application Publication No. 20110117189; U.S. Patent Application Publication No. 20090017543; U.S. Patent Application Publication No. 20070054961, U.S. Patent Application Publication No. 20100317109. Lentiviral vectors are also disclosed for delivery to the brain, for example, see U.S. Patent Application Publication No. 20110293571; U.S. Patent Application Publication No. 20110293571, U.S. Patent Application Publication No. 20040013648, U.S. Patent Application Publication No. 20070025970, U.S. Patent Application Publication No. 20090111106, and U.S. Patent No. 7259015.
[0265] RNA Delivery RNA Delivery: The CRISPR enzyme, such as Cpf1, and / or either this RNA, such as the guide RNA, can also be delivered in the form of RNA. Cpf1 mRNA can be prepared using in vitro transcription. For example, Cpf1 mRNA can be synthesized using a PCR cassette containing the following elements: a T7_promoter - Kozak sequence (GCCACC) - Cpf1 - 3'UTR derived from β - globin - polyA tail (a series of 120 or more adenines). This cassette can be used for transcription by T7 polymerase. The guide RNA can also be transcribed using in vitro transcription from a cassette containing a T7_promoter - GG - guide RNA sequence.
[0266] To enhance expression and potentially reduce toxicity, the CRISPR enzyme coding sequence and / or the guide RNA can be modified to include one or more modified nucleosides, for example, using pseudouridine or 5 - methyl - C.
[0267] mRNA delivery methods are currently particularly promising for liver delivery.
[0268] Many clinical studies on RNA delivery have focused on RNAi or antisense, but these systems can be adapted for the delivery of RNA to practice the present invention. The following references on RNAi etc. should be read accordingly.
[0269] Particle delivery systems and / or formulations: Some types of particle delivery systems and / or formulations are known to be useful in a variety of biomedical applications. Generally, a particle is defined as a small object that behaves as a single unit in terms of its transport and properties. Particles are further classified based on their diameter. Coarse particles encompass the range of 2,500 to 10,000 nanometers. Fine particles have a size of 100 to 2,500 nanometers. Ultra-fine particles, or nanoparticles, generally have a size of 1 to 100 nanometers. The criterion of the 100 nm limit is the fact that new properties that distinguish particles from bulk materials typically occur at a critical length scale of less than 100 nm.
[0270] As used herein, the particle delivery system / formulation is defined as any biological delivery system / formulation that contains the particles of the invention. The particles of the invention are any entity having a maximum diameter (e.g., diameter) of less than 100 microns (μm). In some embodiments, the particles of the invention have a maximum diameter of less than 10 μm. In some embodiments, the particles of the invention have a maximum diameter of less than 2000 nanometers (nm). In some embodiments, the particles of the invention have a maximum diameter of less than 1000 nanometers (nm). In some embodiments, the particles of the invention have a maximum diameter of less than 900 nm, 800 nm, 700 nm, 600 nm, 500 nm, 400 nm, 300 nm, 200 nm, or 100 nm. Typically, the particles of the invention have a maximum diameter (e.g., diameter) of 500 nm or less. In some embodiments, the particles of the invention have a maximum diameter (e.g., diameter) of 250 nm or less. In some embodiments, the particles of the invention have a maximum diameter (e.g., diameter) of 200 nm or less. In some embodiments, the particles of the invention have a maximum diameter (e.g., diameter) of 150 nm or less. In some...
Claims
**Claim 1** An in vitro or ex vivo method for modifying eukaryotic cells, comprising: (a) a modified Cpf1 effector protein comprising one or more mutations in the Nuc domain, or a nucleic acid encoding said modified Cpf1 effector protein, wherein said modified Cpf1 effector protein is a nickase and comprises a mutation at the amino acid residue corresponding to R1226 of Acidaminococcus sp. BV3L6 Cpf1; and (b) a guide polynucleotide comprising a guide sequence linked to a direct repeat sequence, or a nucleic acid encoding said guide polynucleotide comprising the step of delivering a CRISPR-Cas system comprising the same into eukaryotic cells, wherein said modified Cpf1 effector protein forms a CRISPR complex with said guide polynucleotide, said guide sequence hybridizes to a target sequence near a protospacer adjacent motif (PAM) at a target genomic locus, leading to sequence-specific binding of said CRISPR complex to the target genomic locus in the nucleus of said eukaryotic cell, said target genomic locus is cleaved or edited to generate a modified eukaryotic cell, and said method excludes modification of human germline gene identity, a method. **Claim 2** The method according to claim 1, wherein said method comprises the step of delivering said guide polynucleotide and said modified Cpf1 effector protein into said eukaryotic cells. a method. **Claim 3** The method according to claim 1, wherein said method comprises the step of delivering mRNA encoding said guide polynucleotide and said modified Cpf1 effector protein into said eukaryotic cells. a method. **Claim 4** The method according to claim 1, wherein said method comprises the step of delivering one or more vectors encoding said guide polynucleotide and said modified Cpf1 effector protein into said eukaryotic cells. a method. **Claim 5** The method according to claim 4, wherein said one or more vectors comprise one or more viral vectors, wherein said one or more viral vectors may comprise an adenoviral vector, a lentiviral vector, or an adeno-associated viral vector. a method. **Claim 6** The method according to any one of claims 1 to 5, wherein The modified Cpf1 effector protein is a modified Acidaminococcus sp. Cpf1, a Lachnospiraceae bacterium Cpf1, or a Francisella novicida Cpf1, Method. **Claim 7** The method according to any one of claims 1 to 5, wherein the modified Cpf1 effector protein is a modified Acidaminococcus sp. BV3L6 Cpf1, Method. **Claim 8** The method according to any one of claims 1 to 5, wherein the modified Cpf1 effector protein is a modified Lachnospiraceae bacterium ND2006 Cpf1 or a Lachnospiraceae bacterium MA2020 Cpf1, Method. **Claim 9** The method according to any one of claims 1 to 8, wherein the mutation corresponds to R1226A of Acidaminococcus sp. BV3L6 Cpf1, Method. **Claim 10** The method according to any one of claims 1 to 9, wherein the modified Cpf1 effector protein is fused with at least one nuclear localization signal (NLS), Method. **Claim 11** The method according to any one of claims 1 to 10, wherein the modified Cpf1 effector protein is fused with at least two NLSs, Method. **Claim 12** The method according to any one of claims 1 to 11, wherein the modified Cpf1 effector protein is fused with at least one heterologous protein domain, Method. **Claim 13** The method according to claim 12, wherein the heterologous protein domain has one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, or nucleic acid binding activity, Method. **Claim 14** The method according to claim 12, The heterologous protein domain is a transcriptional repressor, transcriptional activator, nuclease domain, DNA methyltransferase, protein acetyltransferase, protein deacetylase, protein methyltransferase, protein deaminase, protein kinase, or protein phosphatase. Method. **Claim 15** The method according to any one of claims 1 to 14, wherein the CRISPR-Cas system further comprises a template polynucleotide overlapping with at least 5 nucleotides of the target sequence. Method. **Claim 16** The method according to any one of claims 1 to 15, wherein the eukaryotic cell is a mammalian cell or a human cell. Method.
Citation Information
Patent Citations
Novel CRISPR enzymes and systems
JP2018518183A