Novel crispr enzyme and system
CRISPR-Cas locus effector proteins, like Cpf1, form complexes with nucleic acid components to induce targeted modifications, addressing the need for efficient genome editing in diverse cell types with reduced off-target effects and improved scalability.
Patent Information
- Application Number
- JP2025135650
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2016-01-07
- Filing Date
- 2025-08-18
- Publication Date
- 2025-12-09
AI Technical Summary
There is a need for alternative and robust systems and techniques for targeting nucleic acids or polynucleotides for a wide range of applications, particularly in genome and epigenome targeting, which are inexpensive, easy to set up, and suitable for targeting multiple locations within eukaryotic genomes.
The use of non-naturally occurring or engineered CRISPR-Cas locus effector proteins, such as Cpf1, which form complexes with nucleic acid components to induce modifications, including strand breaks, at target loci, utilizing guide RNAs with optimized secondary structures and delivery systems like liposomes or viral vectors.
Enables precise and efficient genome editing and modification in various cell types, including eukaryotic cells, with reduced off-target effects and enhanced scalability and versatility.
Smart Images

Figure 2025179085000243 
Figure 2025179085000244 
Figure 2025179085000245
Abstract
Description
[Technical Field]
[0001] Related Applications and Incorporation by Reference This application claims the benefit of and priority to U.S. Provisional Patent Application No. 62 / 181,739, filed June 18, 2015; U.S. Provisional Patent Application No. 62 / 193,507, filed July 16, 2015; U.S. Provisional Patent Application No. 62 / 201,542, filed August 5, 2015; U.S. Provisional Patent Application No. 62 / 205,733, filed August 16, 2015; U.S. Provisional Patent Application No. 62 / 232,067, filed September 24, 2015; U.S. Patent Application No. 14 / 975,085, filed December 18, 2015; and European Patent Application No. 16150428.7.
[0002] Each of the foregoing applications, and all documents cited therein or during the prosecution thereof ("Application Citations"), and all documents cited or referenced in the documents cited herein, together with any manufacturer's instructions, manuals, product specifications, and product sheets for any products mentioned herein or in any document incorporated by reference herein, are hereby incorporated by reference herein and may be used in the practice of the present invention. More specifically, all referenced documents are incorporated by reference to the same extent as if each individual document were specifically and individually indicated to be incorporated by reference.
[0003] Federally Sponsored Research Statement This invention was made with federal support under Grant No. MH100706 awarded by the National Institutes of Health. The federal government has certain rights in this invention.
[0004] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format, which is hereby incorporated by reference in its entirety. Said ASCII copy, created on December 17, 2015, is named 47627.05.2123_SL.txt and is 2,467,205 bytes in size.
[0005] The present invention relates generally to systems, methods, and compositions used to control gene expression, including gene transcript perturbation or nucleic acid editing, including sequence targeting, which may use vector systems related to clustered regularly interspaced short palindromic repeats (CRISPR) and its components. [Background technology]
[0006] Recent advances in genome sequencing technologies and analytical methods have rapidly improved our ability to catalog and map genetic factors associated with various biological functions and diseases. Precise genome targeting technologies are needed to enable systematic reverse engineering of causal genetic variations by selectively perturbing individual genetic elements and advance synthetic biology, biotechnological, and medical applications. While genome editing techniques such as designer zinc fingers, transcription activator-like effectors (TALEs), or homing meganucleases are available to generate targeted genome perturbations, there remains a need for novel genome engineering technologies that utilize novel strategies and molecular mechanisms and are inexpensive, easy to set up, scalable, and suitable for targeting multiple locations within eukaryotic genomes. This will serve as a major resource for novel applications in genome engineering and biotechnology.
[0007] CRISPR-Cas systems in bacterial and archaeal adaptive immunity exhibit a high degree of diversity in terms of protein composition and genomic locus organization. CRISPR-Cas loci contain over 50 gene families, with no strictly universal genes, suggesting rapid evolution and a high degree of locus organization diversity. To date, a multidisciplinary approach has comprehensively identified approximately 395 cas gene profiles for 93 Cas proteins. The classification includes signature gene profiles and locus organization signatures. A novel classification of CRISPR-Cas systems has been proposed, broadly dividing these systems into two classes: Class 1, which has multisubunit effector complexes, and Class 2, which has single-subunit effector modules, as exemplified by the Cas9 protein. Novel effector proteins associated with Class 2 CRISPR-Cas systems can be developed as powerful genome engineering tools, and the prediction, engineering, and optimization of putative novel effector proteins are crucial.
[0008] Citation or identification of any document in this application shall not be construed as an admission that such document is available as prior art to the present invention. Summary of the Invention [Problem to be solved by the invention]
[0009] There is an urgent need for alternative and robust systems and techniques for targeting nucleic acids or polynucleotides (e.g., DNA or RNA, or any hybrid or derivative thereof) for a wide range of applications. The present invention addresses this need and provides related advantages. The addition of the present novel DNA or RNA targeting systems to the repertoire of genome and epigenome targeting technologies can transform the study and perturbation or editing of specific target sites through direct detection, analysis, and manipulation. To effectively utilize the present DNA or RNA targeting systems for genome or epigenome targeting without adverse effects, it is critically important to understand the engineering and optimization aspects of these DNA or RNA targeting tools. [Means for solving the problem]
[0010] The present invention provides a method for modifying a sequence associated with or at a target locus of interest, the method comprising delivering to the locus a non-naturally occurring or engineered composition comprising a putative type V CRISPR-Cas locus effector protein and one or more nucleic acid components, wherein the effector protein forms a complex with the one or more nucleic acid components, and upon binding of the complex to the locus of interest, the effector protein induces modification of the sequence associated with or at the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break. In a preferred embodiment, the sequence associated with or at the target locus of interest comprises DNA, and the effector protein is encoded by a subtype VA CRISPR-Cas locus or a subtype VB CRISPR-Cas locus.
[0011] The terms Cas enzyme, CRISPR enzyme, CRISPR protein, Cas protein and CRISPR Cas are generally used interchangeably and will be understood to refer by analogy to the novel CRISPR effector proteins described further herein unless otherwise clear, such as by specific reference to Cas9. The CRISPR effector protein described herein is preferably the Cpfl effector protein.
[0012] The present invention provides a method for modifying a sequence associated with or at a target locus of interest, the method comprising delivering a non-naturally occurring or engineered composition comprising a Cpf1 locus effector protein and one or more nucleic acid components to the sequence associated with or at the locus, wherein the Cpf1 effector protein forms a complex with the one or more nucleic acid components, and upon binding of the complex to the locus of interest, the effector protein induces modification of the sequence associated with or at the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break. In a preferred embodiment, the Cpf1 effector protein forms a complex with one nucleic acid component; advantageously an engineered or non-naturally occurring nucleic acid component. Induction of modification of the sequence associated with or at the target locus of interest can be under Cpf1 effector protein-nucleic acid guidance. In a preferred embodiment, one nucleic acid component is a CRISPR RNA (crRNA). In a preferred embodiment, one nucleic acid component is a mature crRNA or guide RNA, wherein the mature crRNA or guide RNA comprises a spacer sequence (or guide sequence) and a direct repeat sequence or a derivative thereof. In a preferred embodiment, the spacer sequence or a derivative thereof comprises a seed sequence, wherein the seed sequence is critical for recognition and / or hybridization with a sequence at the target locus. In a preferred embodiment, the seed sequence of the FnCpf1 guide RNA is located within the first approximately 5 nt on the 5' end of the spacer sequence (or guide sequence). In a preferred embodiment, the strand cleavage is a sticky-end cleavage with a 5' overhang. In a preferred embodiment, the sequence associated with or at the target locus of interest comprises linear DNA or supercoiled DNA.
[0013] An aspect of the present invention relates to a Cpf1 effector protein complex having one or more non-naturally occurring, engineered, modified, or optimized nucleic acid components. In a preferred embodiment, the nucleic acid component of the complex can include a guide sequence linked to a direct repeat sequence, wherein the direct repeat sequence includes one or more stem-loops or optimized secondary structures. In a preferred embodiment, the direct repeat has a minimum length of 16 nt and a single stem-loop. In a further embodiment, the direct repeat is longer than 16 nt, preferably longer than 17 nt, and includes two or more stem-loops or optimized secondary structures. In a preferred embodiment, the direct repeat can be modified to include one or more protein-binding RNA aptamers. In a preferred embodiment, the one or more aptamers can be included as part of an optimized secondary structure. Such aptamers can be capable of binding to bacteriophage coat proteins. The bacteriophage coat protein may be selected from the group including Qβ, F2, GA, fr, JP501, MS2, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s, and PRR1. In a preferred embodiment, the bacteriophage coat protein is MS2. The present invention also provides nucleic acid components of the complex that are 30 or more, 40 or more, or 50 or more nucleotides in length.
[0014] The present invention provides a genome editing method, which comprises two or more rounds of Cpf1 effector protein targeting and cleavage. In a specific embodiment, the first round comprises the Cpf1 effector protein cleaving a sequence associated with a target locus remote from a seed sequence, and the second round comprises the Cpf1 effector protein cleaving a sequence at the target locus. In a preferred embodiment of the present invention, the first targeting round with the Cpf1 effector protein generates an indel, and the second targeting round with the Cpf1 effector protein can be repaired by homology-dependent repair (HDR). In a most preferred embodiment of the present invention, one or more targeting rounds with the Cpf1 effector protein generate a sticky-end break that can be repaired by insertion of a repair template.
[0015] The present invention provides a method for genome editing or modification of a sequence associated with or located at a target locus of interest, comprising introducing a Cpf1 effector protein complex into any desired cell type, prokaryotic or eukaryotic, whereby the Cpf1 effector protein complex functions to integrate a DNA insert into the genome of the eukaryotic or prokaryotic cell. In a preferred embodiment, the cell is a eukaryotic cell and the genome is a mammalian genome. In a preferred embodiment, integration of the DNA insert is facilitated by a non-homologous end joining (NHEJ)-based gene insertion mechanism. In a preferred embodiment, the DNA insert is an exogenously introduced DNA template or repair template. In a preferred embodiment, the exogenously introduced DNA template or repair template is delivered along with a polynucleotide vector for expression of the Cpf1 effector protein complex or one of its components or components. In a more preferred embodiment, the eukaryotic cell is a non-dividing cell (e.g., a non-dividing cell in which genome editing by HDR is particularly challenging). In a preferred genome editing method in human cells, Cpf1 effector proteins may include, but are not limited to, FnCpf1, AsCpf1, and LbCpf1 effector proteins.
[0016] The present invention also provides a method of modifying a target locus of interest, the method comprising delivering to the locus a non-naturally occurring or engineered composition comprising a C2c1 locus effector protein and one or more nucleic acid components, wherein the C2c1 effector protein forms a complex with the one or more nucleic acid components, and upon binding of the complex to the locus of interest, the effector protein induces modification of the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break.
[0017] In such methods, a target locus of interest may be contained in a DNA molecule in vitro, which in a preferred embodiment is a plasmid.
[0018] In such methods, a target locus of interest can be contained in a DNA molecule within the cell. The cell can be a prokaryotic or eukaryotic cell. The cell can be a mammalian cell. The mammalian cell can be a non-human primate, bovine, porcine, rodent, or murine cell. The cell can also be a non-mammalian eukaryotic cell, such as a poultry, fish, or shrimp cell. The cell can also be a plant cell. The plant cell can be a crop plant, such as cassava, corn, sorghum, wheat, or rice. The plant cell can also be an algae, tree, or vegetable. The modification introduced into the cell by the present invention can be such that the cell and its progeny are altered for improved production of a biological product, such as an antibody, starch, alcohol, or other desired cellular product. The modification introduced into the cell by the present invention can be such that the cell and its progeny contain a change that alters the biological product produced.
[0019] The present invention provides a method of modifying a target locus of interest, the method comprising delivering to the locus a non-naturally occurring or engineered composition comprising a Type VI CRISPR-Cas locus effector protein and one or more nucleic acid components, wherein the effector protein forms a complex with the one or more nucleic acid components, and upon binding of the complex to the locus of interest, the effector protein induces modification of the target locus of interest. In a preferred embodiment, the modification is the introduction of a strand break.
[0020] In a preferred embodiment, the target locus of interest comprises DNA.
[0021] In such methods, the target locus of interest can be contained in a DNA molecule within the cell. The cell can be a prokaryotic or eukaryotic cell. The cell can be a mammalian cell. The mammalian cell can be a non-human mammalian cell, such as a primate, bovine, ovine, porcine, canine, rodent, or Leporidae cell, such as a monkey, cow, sheep, pig, dog, rabbit, rat, or mouse cell. The cell can also be a non-mammalian eukaryotic cell, such as a poultry avian (e.g., chicken), vertebrate fish (e.g., salmon), or crustacean (e.g., oyster, clam, lobster, shrimp) cell. The cell can also be a plant cell. The plant cell can be a monocotyledonous or dicotyledonous plant or a crop or cereal plant, such as cassava, corn, sorghum, soybean, wheat, oat, or rice. The plant cell may also be that of an algae, a tree or productive plant, a fruit or vegetable (e.g., a citrus tree, such as an orange, grapefruit or lemon tree; a peach or nectarine tree; an apple or pear tree; a nut tree, such as an almond, walnut or pistachio tree; a Solanaceae plant; a Brassica plant; a Lactuca plant; a Spinacia plant; a Capsicum plant; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.).
[0022] In any of the described methods, the target locus of interest may be a genomic or epigenomic locus of interest. In any of the described methods, the complex may be delivered with multiple guides for multiplexed use. In any of the described methods, more than one protein may be used.
[0023] In a preferred embodiment of the present invention, biochemical or in vitro or in vivo cleavage of a sequence associated with or located at a target locus of interest, e.g., by an FnCpf1 effector protein, occurs without a putative transactivating crRNA (tracr RNA) sequence. In other embodiments of the present invention, cleavage, e.g., by other CRISPR family effector proteins, can occur with a putative transactivating crRNA (tracr RNA) sequence; however, after evaluation of the FnCpf1 locus, Applicants concluded that target DNA cleavage by the Cpf1 effector protein complex does not require tracrRNA. Applicants determined that a Cpf1 effector protein complex containing only a Cpf1 effector protein and crRNA (a guide RNA comprising a direct repeat sequence and a guide sequence) was sufficient for target DNA cleavage. Thus, the present invention provides a method for modifying a target locus of interest as described herein above, wherein the effector protein is a Cpf1 protein, and the effector protein forms a complex with the target sequence without the presence of tracr.
[0024] In any of the described methods, the effector protein (e.g., Cpf1) and nucleic acid components may be provided by one or more polynucleotide molecules encoding the protein and / or one or more nucleic acid components, and wherein the one or more polynucleotide molecules are operably configured to express the protein and / or one or more nucleic acid components. The one or more polynucleotide molecules may comprise one or more regulatory elements operably configured to express the protein and / or one or more nucleic acid components. The one or more polynucleotide molecules may be contained within one or more vectors. The present invention encompasses such one or more polynucleotide molecules, e.g., such polynucleotide molecules operably configured to express the protein and / or one or more nucleic acid components, and such one or more vectors.
[0025] In any of the methods described, the strand break may be a single-strand break or a double-strand break.
[0026] The regulatory element may comprise an inducible promoter. The polynucleotide and / or vector system may comprise an inducible system.
[0027] In any of the methods described, one or more polynucleotide molecules may be included in the delivery system, or one or more vectors may be included in the delivery system.
[0028] In any of the methods described, the non-naturally occurring or engineered composition may be delivered by liposomes, particles (e.g., nanoparticles), exosomes, microvesicles, gene guns, or one or more vectors, such as nucleic acid molecules or viral vectors.
[0029] The present invention also provides non-naturally occurring or engineered compositions that have the characteristics as discussed herein or are compositions defined in any of the methods described herein.
[0030] The present invention also provides vector systems comprising one or more vectors, wherein the one or more vectors comprise one or more polynucleotide molecules encoding components of a non-naturally occurring or engineered composition having the characteristics as discussed herein or being a composition defined in any of the methods described herein.
[0031] The present invention also provides delivery systems comprising one or more vectors or one or more polynucleotide molecules, wherein the one or more vectors or polynucleotide molecules comprise one or more polynucleotide molecules encoding components of a non-naturally occurring or engineered composition having the characteristics as discussed herein or being a composition defined by any of the methods described herein.
[0032] The present invention also provides non-naturally occurring or engineered compositions, or one or more polynucleotides encoding components of said compositions, or vectors or delivery systems comprising one or more polynucleotides encoding components of said compositions, for use in therapeutic treatment methods, which may include gene or genome editing or gene therapy.
[0033] The present invention also encompasses computational methods and algorithms for predicting novel Class 2 CRISPR-Cas systems and identifying components therein.
[0034] The present invention also provides methods and compositions in which one or more amino acid residues of an effector protein, e.g., an engineered or non-naturally occurring effector protein or Cpf1, may be modified. In certain embodiments, the modification may include a mutation of one or more amino acid residues of the effector protein. The one or more mutations may be in one or more catalytically active domains of the effector protein. The effector protein may have reduced or eliminated nuclease activity compared to an effector protein lacking the one or more mutations. The effector protein may not induce cleavage of either DNA or RNA strand at a target locus of interest. The effector protein may not induce cleavage of either DNA or RNA strand at a target locus of interest. In a preferred embodiment, the one or more mutations may include two mutations. In a preferred embodiment, one or more amino acid residues are modified in a Cpf1 effector protein, e.g., an engineered or non-naturally occurring effector protein or Cpf1. In a preferred embodiment, the Cpf1 effector protein is an FnCpf1 effector protein. In a preferred embodiment, the one or more modified or mutated amino acid residues are D917A, E1006A, or D1255A relative to the amino acid position numbering of the FnCpf1 effector protein. In a further preferred embodiment, the one or more mutated amino acid residues are D908A, E993A, or D1263A relative to the amino acid position in AsCpf1, or LbD832A, E925A, D947A, or D1180A relative to the amino acid position in LbCpf1.
[0035] The present invention also provides one or more mutations or two or more mutations to be present in the catalytically active domain of an effector protein, including a RuvC domain. In some embodiments of the present invention, the RuvC domain may comprise a RuvCI, RuvCII, or RuvCIII domain, or a catalytically active domain that is homologous to a RuvCI, RuvCII, or RuvCIII domain, or any related domain as described in any of the methods described herein. The effector protein may comprise one or more heterologous functional domains. The one or more heterologous functional domains may comprise one or more nuclear localization signal (NLS) domains. The one or more heterologous functional domains may comprise at least two or more NLS domains. The one or more NLS domains may be located at, near, or adjacent to the terminus of the effector protein (e.g., Cpf1), and if there are two or more NLSs, each of the two may be located at, near, or adjacent to the terminus of the effector protein (e.g., Cpf1). The one or more heterologous functional domains may comprise one or more transcription activation domains. In a preferred embodiment, the transcription activation domain may comprise VP64. The one or more heterologous functional domains may comprise one or more transcription repression domains. In a preferred embodiment, the transcription repression domain comprises a KRAB domain or a SID domain (e.g., SID4X). The one or more heterologous functional domains may comprise one or more nuclease domains. In a preferred embodiment, the nuclease domain comprises Fok1.
[0036] The present invention also provides one or more heterologous functional domains having one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription deactivator activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity. At least one or more heterologous functional domains may be at or near the amino terminus of the effector protein, and / or wherein at least one or more heterologous functional domains are at or near the carboxy terminus of the effector protein. One or more heterologous functional domains may be fused to the effector protein. One or more heterologous functional domains may be tethered to the effector protein. One or more heterologous functional domains may be linked to the effector protein via a linker moiety.
[0037] The present invention also provides a method for preventing the development of bacteria of the genera Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaete, and the like. haeta, Lactobacillus, Eubacterium, Corynebacter, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Clostridiaridium , Leptotrichia, Francisella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus, Letospira, Desulfovibrio Also provided are effector proteins (e.g., Cpf1), including effector proteins (e.g., Cpf1) from organisms of genera including Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium, or Acidaminococcus.
[0038] The present invention also provides a method for the prevention and treatment of S. mutans, S. agalactiae, S. equisimilis, S. sanguinis, S. pneumoniae, C. jejuni, C. coli, N. salsuginis, N. tergarcus, S. auricularis, S. carnosus, S. ca Also provided are effector proteins (e.g., Cpf1) including effector proteins (e.g., Cpf1) from organisms derived from: N. rnosus; N. meningitidis, N. gonorrhoeae; L. monocytogenes, L. ivanovii; C. botulinum, C. difficile, C. tetani, C. sordellii.
[0039] The effector protein can include a chimeric effector protein comprising a first fragment derived from a first effector protein (e.g., Cpf1) ortholog and a second fragment derived from a second effector (e.g., Cpf1) protein ortholog, wherein the first and second effector protein orthologs are different. At least one of the first and second effector protein (e.g., Cpf1) orthologs is / are from a species of the genera Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus, Eubacterium, Corynebacter, Carnobacterium, Rhodobacter, Listeria, or the like. teria), Paludibacter, Clostridium, Lachnospiraceae, Clostridiaridium, Leptotrichia, Francisella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus, Letospira, Desulfovibrio, Desulfonatronum, Opitutaceae,The effector protein may include an effector protein (e.g., Cpf1) derived from an organism including Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium, or Acidaminococcus; for example, a chimeric effector protein comprising a first fragment and a second fragment, wherein each of the first and second fragments is derived from an organism of the genus Streptococcus, Campylobacter (C ampylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus, Eubacterium m), Corynebacter, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Clostridiaridium, Leptotrichia, Francisella, Legionella nella), Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus, Letospira, Desulfovibrio, Desulfonatronum, Opitutaceae,Cpf1 from an organism including Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium or Acidaminococcus, wherein the first and second fragments are not from the same bacterium; for example, a chimeric effector protein comprising a first fragment and a second fragment, wherein each of the first and second fragments is selected from Cpf1 from S. mutans, S. agalactiae, S. equisimilis, S. sanguinis, Streptococcus pneumoniae (S. pneumonia; C. jejuni, C. coli; N. salsuginis, N. tergarcus; S. auricularis, S. carnosus; N. meningitides, N. gonorrhoeae; L. monocytogenes, L. ivanovii; C. botulinum, C. difficile, C. tetani, C. sordellii; Francisella tularensis tularensis 1, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020,selected from Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis, Prevotella disiens, and Cpf1 of Porphyromonas macacae, wherein the first and second fragments are not derived from the same bacterium.
[0040] In a preferred embodiment of the invention, the effector protein is derived from the Cpfl locus (such effector proteins are also referred to herein as "Cpflp"), e.g., a Cpfl protein (and such effector proteins or Cpfl proteins or proteins derived from the Cpfl locus are also referred to as "CRISPR enzymes"). Cpfl loci include, but are not limited to, the Cpfl loci of bacterial species listed in Figure 64. In a more preferred embodiment, Cpflp is selected from Francisella tularensis 1, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae The bacterial species may be selected from the group consisting of Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis, Prevotella disiens, and Porphyromonas macacae.In certain embodiments, Cpflp is derived from a bacterial species selected from Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020. In certain embodiments, the effector protein is derived from a subspecies of Francisella tularensis 1, including but not limited to Francisella tularensis subsp. Novicida.
[0041] In further embodiments of the invention, a protospacer adjacent motif (PAM) or PAM-like motif directs binding of an effector protein complex to a target locus of interest. In a preferred embodiment of the invention, the PAM is 5'TTN, where N is A / C / G or T, and the effector protein is FnCpflp. In another preferred embodiment of the invention, the PAM is 5'TTTV, where V is A / C or G, and the effector protein is AsCpfl, LbCpfl, or PaCpflp. In a specific embodiment, the PAM is 5'TTN, where N is A / C / G or T, the effector protein is FnCpflp, and the PAM is located upstream of the 5' end of the protospacer. In a specific embodiment of the invention, the PAM is 5'CTA, where the effector protein is FnCpflp, and the PAM is located upstream of the 5' end of the protospacer or target locus. In a preferred embodiment, the present invention provides an expanded targeting range for RNA-guided genome editing nucleases, where T-rich PAMs of the Cpf1 family enable targeting and editing of AT-rich genomes.
[0042] In certain embodiments, CRISPR enzymes can be engineered to contain one or more mutations that reduce or eliminate nuclease activity. Amino acid positions in the FnCpflp RuvC domain include, but are not limited to, D917A, E1006A, E1028A, D1227A, D1255A, N1257A, D917A, E1006A, E1028A, D1227A, D1255A, and N1257A. Applicants have also identified a putative second nuclease domain that is most similar to the PD-(D / E)XK nuclease superfamily and HincII endonuclease-like. Point mutations made in this putative nuclease domain to substantially reduce nuclease activity include, but are not limited to, N580A, N584A, T587A, W609A, D610A, K613A, E614A, D616A, K624A, D625A, K627A, and Y629A. In a preferred embodiment, the mutation in the FnCpflp RuvC domain is D917A or E1006A, wherein the D917A or E1006A mutation completely inactivates the DNA cleavage activity of the FnCpfl effector protein. In another embodiment, the mutation in the FnCpflp RuvC domain is D1255A, wherein the mutant FnCpfl effector protein has significantly reduced nucleolytic activity.
[0043] Amino acid positions in the AsCpf1p RuvC domain include, but are not limited to, 908, 993, and 1263. In a preferred embodiment, the mutations in the AsCpf1p RuvC domain are D908A, E993A, and D1263A, where the D908A, E993A, and D1263A mutations completely inactivate the DNA cleavage activity of the AsCpf1 effector protein. Amino acid positions in the LbCpf1p RuvC domain include, but are not limited to, 832, 947, or 1180. In a preferred embodiment, the mutations in the LbCpf1p RuvC domain are LbD832A, E925A, D947A, or D1180A, where the LbD832A, E925A, D947A, or D1180A mutations completely inactivate the DNA cleavage activity of the LbCpf1 effector protein.
[0044] Mutations can also be made in adjacent residues, e.g., amino acids near those indicated above that are involved in nuclease activity. In some embodiments, only the RuvC domain is inactivated, and in other embodiments, another putative nuclease domain is inactivated, where the effector protein complex functions as a nickase to cleave only one DNA strand. In preferred embodiments, the other putative nuclease domain is a HincII-like endonuclease domain. In some embodiments, two FnCpf1, AsCpf1, or LbCpf1 mutants (each with a different nickase) are used to increase specificity, and two nickase mutants are used to cleave DNA at the target (where both nickases cleave the DNA strand with minimal or no off-target modifications, where only one DNA strand is cleaved and subsequently repaired). In a preferred embodiment, the Cpf1 effector protein cleaves a sequence associated with or at a target locus of interest as a homodimer comprising two Cpf1 effector protein molecules. In a preferred embodiment, the homodimer may comprise two Cpf1 effector protein molecules that contain different mutations in their respective RuvC domains.
[0045] The present invention contemplates methods using two or more nickases, particularly dual or double nickase approaches. In some aspects and embodiments, a single type of FnCpf1, AsCpf1, or LbCpf1 nickase, such as a modified FnCpf1, AsCpf1, or LbCpf1 or a modified FnCpf1, AsCpf1, or LbCpf1 nickase as described herein, may be delivered. This results in two FnCpf1 nickases binding to the target DNA. In addition, it is also contemplated that different orthologs may be used, such as an FnCpf1, AsCpf1, or LbCpf1 nickase for one strand of DNA (e.g., the coding strand) and an ortholog for the non-coding strand or the opposite DNA strand. The ortholog may be a Cas9 nickase, such as, but not limited to, SaCas9 nickase or SpCas9 nickase. It may be advantageous to use two different orthologs, which may require different PAMs and may also have different guide requirements, thus allowing a greater degree of user control. In certain embodiments, DNA cleavage involves at least four types of nickases, each type guided to a different sequence in the target DNA, where each pair introduces a first nick in one DNA strand and a second pair introduces a nick in the second DNA strand. In such methods, at least two single-strand break pairs are introduced into the target DNA, where introduction of the first and second single-strand break pairs results in excision of the target sequence between the first and second single-strand break pairs. In certain embodiments, one or both of the orthologs is controllable, i.e., inducible.
[0046] In certain embodiments of the present invention, the guide RNA or mature crRNA comprises, consists essentially of, or consists of a direct repeat sequence and a guide sequence or spacer sequence. In certain embodiments, the guide RNA or mature crRNA comprises, consists essentially of, or consists of a direct repeat sequence linked to a guide sequence or spacer sequence. In certain embodiments, the guide RNA or mature crRNA comprises a 19-nt partial direct repeat followed by a 20-30-nt, preferably about 20-nt, 23-25-nt, or 24-nt guide sequence or spacer sequence. In certain embodiments, the effector protein is an FnCpf1, AsCpf1, or LbCpf1 effector protein, which requires at least 16-nt of guide sequence to achieve detectable DNA cleavage and a minimum of 17-nt of guide sequence to achieve efficient DNA cleavage in vitro. In certain embodiments, the direct repeat sequence is located upstream (i.e., 5') of the guide sequence or spacer sequence. In a preferred embodiment, the seed sequence of the FnCpf1, AsCpf1 or LbCpf1 guide RNA (i.e., the critical sequence essential for recognizing and / or hybridizing to a sequence at the target locus) is within about the first 5 nt on the 5' end of the guide sequence or spacer sequence.
[0047] In preferred embodiments of the present invention, the mature crRNA comprises a stem-loop or optimized stem-loop structure or optimized secondary structure. In preferred embodiments, the mature crRNA comprises a stem-loop or optimized stem-loop structure in a direct repeat sequence, where the stem-loop or optimized stem-loop structure is important for cleavage activity. In certain embodiments, the mature crRNA preferably comprises a single stem-loop. In certain embodiments, the direct repeat sequence preferably comprises a single stem-loop. In certain embodiments, the cleavage activity of the effector protein complex is modified by introducing mutations that affect the stem-loop RNA duplex structure. In preferred embodiments, mutations that maintain the stem-loop RNA duplex may be introduced, thereby maintaining the cleavage activity of the effector protein complex. In other preferred embodiments, mutations that disrupt the stem-loop RNA duplex structure may be introduced, thereby completely abolishing the cleavage activity of the effector protein complex.
[0048] The present invention also provides nucleotide sequences encoding effector proteins that are codon-optimized for expression in eukaryotic organisms or cells in any of the methods or compositions described herein. In certain embodiments of the invention, the codon-optimized effector protein is FnCpflp, AsCpfl, or LbCpfl, and is codon-optimized for operation in a eukaryotic cell or organism, such as a cell or organism listed elsewhere herein, including, without limitation, a yeast cell, or a mammalian cell or organism, such as a mouse cell, a rat cell, and a human cell, or a non-human eukaryotic organism, such as a plant.
[0049] In certain embodiments of the present invention, at least one nuclear localization signal (NLS) is added to the nucleic acid sequence encoding the Cpf1 effector protein. In preferred embodiments, at least one or more C- or N-terminal NLSs are added (thus, one or more nucleic acid molecules encoding the Cpf1 effector protein can contain coding sequences for one or more NLSs, such that the expressed product has one or more NLSs attached or connected thereto). In preferred embodiments, a C-terminal NLS is added for optimal expression and nuclear targeting in eukaryotic cells, preferably human cells. In preferred embodiments, the codon-optimized effector protein is FnCpf1p, AsCpf1, or LbCpf1, and the spacer length of the guide RNA is 15 to 35 nt. In certain embodiments, the spacer length of the guide RNA is at least 16 nucleotides, e.g., at least 17 nucleotides. In certain embodiments, the spacer length is 15-17 nt, 17-20 nt, 20-24 nt, e.g., 20, 21, 22, 23, or 24 nt, 23-25 nt, e.g., 23, 24, or 25 nt, 24-27 nt, 27-30 nt, 30-35 nt, or 35 nt or more. In certain embodiments of the present invention, the codon-optimized effector protein is FnCpflp, and the direct repeat length of the guide RNA is at least 16 nucleotides. In certain embodiments, the codon-optimized effector protein is FnCpflp, and the direct repeat length of the guide RNA is 16-20 nt, e.g., 16, 17, 18, 19, or 20 nucleotides. In certain preferred embodiments, the direct repeat length of the guide RNA is 19 nucleotides.
[0050] The present invention also encompasses methods for delivering multiple nucleic acid components, each specific for a different target locus of interest, thereby modifying multiple target loci of interest. The nucleic acid components of the complex may comprise one or more protein-binding RNA aptamers. The one or more aptamers may be capable of binding to a bacteriophage coat protein. The bacteriophage coat protein may be selected from the group including Qβ, F2, GA, fr, JP501, MS2, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s, and PRR1. In a preferred embodiment, the bacteriophage coat protein is MS2. The present invention also provides nucleic acid components of the complex that are 30 or more, 40 or more, or 50 or more nucleotides in length.
[0051] The present invention also encompasses cells, components and / or systems of the present invention in which trace amounts of cations are present in the cells, components and / or systems. Advantageously, the cation is magnesium, e.g., Mg 2+ The cation may be present in trace amounts. A preferred range may be from about 1 mM to about 15 mM of cation, which is advantageously Mg 2+ A preferred concentration can be about 1 mM for human-based cells, components, and / or systems, and about 10 mM to about 15 mM for bacterial-based cells, components, and / or systems. See, e.g., Gasiunas et al., PNAS, published online September 4, 2012, www.pnas.org / cgi / doi / 10.1073 / pnas.1208507109.
[0052] Accordingly, it is an object of the present invention not to include within its scope any previously known product, process for making a product, or method for using a product, to which the applicants reserve their rights and which hereby disclose a disclaimer of any previously known product, process, or method. It is further noted that the present invention is not intended to include within its scope any product, process for making a product, or method for using a product that does not meet the description and enablement requirements of the United States Patent and Trademark Office (USPTO) (35 U.S.C. § 112, first paragraph) or the European Patent Office (EPO) (Article 83 EPC), to which the applicants reserve their rights and which hereby disclose a disclaimer of any previously described product, process for making a product, or method for using a product. Compliance with Article 53(c) EPC and Rule 28(b) and (c) EPC may be advantageous in the practice of the present invention. Nothing herein should be construed as a promise.
[0053] It is noted that in this disclosure, and particularly in the claims and / or paragraphs, terms such as "comprises," "comprised," "comprising," and the like may have the meaning ascribed to them in U.S. patent law; for example, they may mean "includes," "included," "including," and the like; and that terms such as "consisting essentially of" and "consists essentially of" have the meaning ascribed to them in U.S. patent law.
[0054] These and other embodiments are disclosed or are apparent from and encompassed by the following detailed description.
[0055] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which: [Brief explanation of the drawings]
[0056] [Figure 1] Figures 1A-1B show a novel classification of CRISPR-Cas systems: Class 1 contains multi-subunit crRNA-effector complexes (Cascade), and Class 2 contains single-subunit crRNA-effector complexes (Cas9-like). [Figure 2] Figure 2 provides the molecular organization of CRISPR-Cas. [Figure 3] Figures 3A-3D provide the structures of type I and type III effector complexes: common architecture / common ancestry despite extensive sequence diversity. [Figure 4] Figure 4 shows CRISPR-Cas as an RNA recognition motif (RRM)-centric system. [Figure 5] Figures 5A-5D show the Cas1 phylogeny, in which adaptation and recombination of the crRNA-effector module represent key aspects of CRISPR-Cas evolution. [Figure 6] Figure 6 shows the CRISPR-Cas census, specifically the distribution of CRISPR-Cas types / subtypes among archaea and bacteria. [Figure 7] Figure 7 shows the pipeline for identifying Cas candidates. [Figure 8] Figures 8A-8D show the complete locus organization of the class 2 lineage. [Figure 9] 9A-9B show the C2c1 neighborhood. [Figure 10] Figures 10A-10C show the Cas1 tree. [Figure 11] 11A and 11B show the domain structure of the class 2 family. [Figure 12] 12A to 12B show the TnpB homology regions of class 2 proteins (SEQ ID NOs: 246 to 428, respectively, in the order listed). [Figure 13] 13A-13B show the C2c2 neighborhood. [Figure 14] Figures 14A to 14E show HEPN RxxxxH motifs of the C2c2 family (SEQ ID NOS: 429 to 1032, respectively, in order of appearance). [Figure 15] FIG. 15 shows C2C1:1. Alicyclobacillus acidoterrestris ATCC 49025 (SEQ ID NOs: 1034 to 1037, respectively, in order of appearance). [Figure 16] FIG. 16 shows C2C1:4. Desulfonatronum thiodismutans strain MLF-1 (SEQ ID NOs: 1038-1041, respectively, in order of appearance). [Figure 17] FIG. 17 shows C2C1:5. Opitutaceae bacterium TAV5 (SEQ ID NOs: 1042 to 1045, respectively, in the order listed). [Figure 18] FIG. 18 shows C2C1:7. Bacillus thermoamylovorans strain B4166 (SEQ ID NOs: 1046-1049, respectively, in order of appearance). [Figure 19] FIG. 19 shows C2C1:9. Bacillus sp. NSP2.1 (SEQ ID NOs: 1050 to 1053, respectively, in order of appearance). [Figure 20] FIG. 20 shows C2C2:1. Lachnospiraceae bacterium MA2020 (SEQ ID NOs: 1054 to 1057, respectively, in order of appearance). [Figure 21] FIG. 21 shows C2C2:2. Lachnospiraceae bacterium NK4A179 (SEQ ID NOs: 1058 to 1064, respectively, in order of appearance). [Figure 22] FIG. 22 shows C2C2:3. [Clostridium] aminophilum DSM 10710 (SEQ ID NOs: 1065 to 1068, respectively, in order of appearance). [Figure 23] FIG. 23 shows C2C2:4. Lachnospiraceae bacterium NK4A144 (SEQ ID NOs: 1069 and 1070, respectively, in order of appearance). [Figure 24] FIG. 24 shows C2C2:5. Carnobacterium gallinarum DSM 4847 (SEQ ID NOs: 1071 to 1074, respectively, in order of appearance). [Figure 25] Figure 25 shows C2C2:6. Carnobacterium gallinarum DSM 4847 (SEQ ID NOs: 1075 to 1081, respectively, in order of appearance). [Figure 26] Figure 26 shows C2C2:7. Paludibacter propionicigenes WB4 (SEQ ID NO: 1082). [Figure 27] Figure 27 shows C2C2:8. Listeria seeligeri serotype 1 / 2b (SEQ ID NOs: 1083 to 1086, respectively, in order of appearance). [Figure 28] Figure 28 shows C2C2:9. Listeria weihenstephanensis FSL R9-0317 (SEQ ID NO: 1087). [Figure 29] FIG. 29 shows C2C2:10. Listeria bacterium FSL M6-0635 (SEQ ID NOs: 1088 and 1091, respectively, in order of appearance). [Figure 30] Figure 30 shows C2C2:11. Leptotrichia wadei F0279 (SEQ ID NO: 1092). [Figure 31]FIG. 31 shows C2C2:12. Leptotrichia wadei F0279 (SEQ ID NOs: 1093 to 1099, respectively, in order of appearance). [Figure 32] Figure 32 shows C2C2:14. Leptotrichia shahii DSM 19757 (SEQ ID NOs: 1100 to 1103, respectively, in order of appearance). [Figure 33] Figure 33 shows C2C2:15. Rhodobacter capsulatus SB 1003 (SEQ ID NOs: 1104 and 1105, respectively, in order of appearance). [Figure 34] Figure 34 shows C2C2:16. Rhodobacter capsulatus R121 (SEQ ID NOs: 1106 and 1107, respectively, in order of appearance). [Figure 35] Figure 35 shows C2C2:17. Rhodobacter capsulatus DE442 (SEQ ID NOs: 1108 and 1109, respectively, in order of appearance). [Figure 36] FIG. 36 shows the tree of DR. [Figure 37] Figure 37 shows the C2C2 tree. [Figure 38] Figures 38A to 38BB show sequence alignments of Cas-Cpf1 orthologs (SEQ ID NOs: 1033 and 1110 to 1166, respectively, in order of appearance). [Figure 39] Figures 39A-B show an overview of the Cpf1 locus alignment. [Figure 40] Figures 40A to 40X show the PACYC184 FnCpf1 (PY001) vector constructs (SEQ ID NO: 1167 and SEQ ID NOs: 1168 to 1189, respectively, in the order listed). [Figure 41] 41A-41I show the sequence of humanized PaCpf1, which has a nucleotide sequence as set forth in SEQ ID NO:1190 and a protein sequence as set forth in SEQ ID NO:1191. [Figure 42]Figure 42 shows the PAM challenge assay. [Figure 43] Figure 43 shows a schematic of the endogenous FnCpf1 locus. pY0001 is a pACY184 backbone (derived from NEB) containing a partial FnCpf1 locus. The FnCpf1 locus was PCR amplified in three pieces using Gibson assembly and cloned into Xba1- and Hind3-cleaved pACYC184. pY0001 contained the endogenous FnCpf1 locus from 255 bp of the acetyltransferase 3' sequence through the fourth spacer sequence. Because spacer 4 is no longer flanked by direct repeats, only spacers 1 through 3 may be active. [Figure 44] Figure 44 shows the PAM libraries and discloses SEQ ID NOs: 1192-1195, respectively, in the order listed. Both PAM libraries (left and right) are in pUC19. The complexity of the left PAM library is 48-65k and the complexity of the right PAM library is 47-16k. Both libraries were prepared at representations of >500. [Figure 45] Figures 45A-4E show the FnCpf1 PAM screen computational analysis. After sequencing the screen DNA, regions corresponding to either the left or right PAM were extracted. For each sample, the number of PAMs present in the sequenced library was compared to the expected number of PAMs in the library (4^8 for the left library and 4^7 for the right). Figure 44A shows that the left library showed PAM depletion. To quantify this depletion, an enrichment ratio was calculated. For both conditions (control pACYC or pACYC containing FnCpf1), the ratio was calculated for each PAM in the library as follows:
number
number
[0057] The figures herein are for illustrative purposes only and are not necessarily drawn to scale.
[0058] This application describes novel RNA-guided endonucleases (e.g., Cpf1 effector proteins) that are functionally distinct from previously described CRISPR-Cas9 systems, and therefore the terminology of elements associated with these novel endonucleases is amended accordingly herein. The Cpf1-associated CRISPR arrays described herein are processed into mature crRNAs without the need for an additional tracrRNA. The crRNAs described herein contain a spacer sequence (or guide sequence) and a direct repeat sequence, and the Cpf1p-crRNA complex alone is sufficient for efficient cleavage of target DNA. The seed sequence described herein, e.g., the seed sequence of the FnCpf1 guide RNA, is located within approximately the first 5 nt on the 5' end of the spacer sequence (or guide sequence), and mutations in the seed sequence adversely affect the cleavage activity of the Cpf1 effector protein complex.
[0059] Generally, CRISPR systems are characterized by elements that promote the formation of a CRISPR complex at the site of a target sequence (also referred to as a protospacer in the context of endogenous CRISPR systems). In the context of CRISPR complex formation, a "target sequence" refers to a sequence to which a guide sequence is designed to target, e.g., be complementary to, where hybridization between the target sequence and the guide sequence promotes the formation of a CRISPR complex. The section of a guide sequence whose complementarity with the target sequence is important for cleavage activity is referred to herein as a seed sequence. A target sequence can comprise any polynucleotide, such as a DNA or RNA polynucleotide, and is contained within a target locus of interest. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell. The invention described herein encompasses novel effector proteins for Class 2 CRISPR-Cas systems, in which Cas9 is an exemplary effector protein; therefore, the terms used herein to describe novel effector proteins can be related to the terms used to describe CRISPR-Cas9 systems.
[0060] CRISPR-Cas loci contain over 50 gene families, and no gene is strictly universal. Therefore, a single phylogenetic tree is not feasible, and identifying new families requires a multi-pronged approach. To date, 395 profiles of cas genes have been comprehensively identified for 93 Cas proteins. The classification includes signature gene profiles plus locus organization signatures. Figure 1 proposes a novel classification of CRISPR-Cas systems. Class 1 includes multisubunit crRNA-effector complexes (Cascade), while Class 2 includes single-subunit crRNA-effector complexes (Cas9-like). Figure 2 provides the molecular organization of CRISPR-Cas. Figure 3 provides the structures of type I and type III effector complexes: common architecture / common ancestry despite extensive sequence diversity. Figure 4 illustrates CRISPR-Cas as an RNA recognition motif (RRM)-centered system. Figure 5 shows the Cas1 phylogeny, where adaptation and recombination of the crRNA-effector module represent key aspects of CRISPR-Cas evolution. Figure 6 shows the CRISPR-Cas census, specifically the distribution of CRISPR-Cas types / subtypes among archaea and bacteria.
[0061] The action of CRISPR-Cas systems is generally divided into three stages: (1) adaptation or spacer integration, (2) processing of the primary transcript of the CRISPR locus (pre-crRNA) and maturation of the crRNA, which contains variable regions corresponding to the spacer and the 5' and 3' segments of the CRISPR repeats, and (3) DNA (or RNA) interference. Two proteins, Cas1 and Cas2, present in the majority of known CRISPR-Cas systems, are sufficient for the insertion of the spacer into the CRISPR cassette. These two proteins form a complex required for this adaptation process; the endonuclease activity of Cas1 is required for spacer integration, while Cas2 appears to perform a non-enzymatic function. The Cas1-Cas2 complex represents a highly conserved "information processing" module of CRISPR-Cas that appears to be quasi-autonomous from the rest of the system (see Annotation and Classification of CRISPR-Cas Systems. Makarova KS, Koonin EV. Methods Mol Biol. 2015;1311:47-75).
[0062] The class 2 systems described so far, namely types II and putative types V, consist of only three or four genes in the cas operon: the cas1 and cas2 genes containing the adaptation module (the cas1-cas2 gene pair is not involved in interference), a single multidomain effector protein involved in interference but also contributing to pre-crRNA processing and adaptation, and, in many cases, a fourth gene of uncharacterized function that is dispensable in at least some type II systems. (In some cases, the fourth gene is cas4 (biochemical and in silico evidence indicates that Cas4 is a PD-(DE)xK superfamily nuclease containing a three-cysteine C-terminal cluster; it has 5'-ssDNA exonuclease activity) or csn2, which encodes an inactivating ATPase.) In most cases, the class 2 cas operon is flanked by genes for a CRISPR array and a distinct RNA species known as tracrRNA, i.e., the trans-encoded small CRISPR RNA. The tracrRNA is partially homologous to repeats within each CRISPR array and is essential for pre-crRNA processing, which is catalyzed by RNase III, a ubiquitous bacterial enzyme that is not associated with CRISPR-Cas loci.
[0063] Cas1 is the most conserved protein present in most CRISPR-Cas systems and evolves at a slower rate than other Cas proteins. Therefore, Cas1 phylogeny has been used as a guide for classifying CRISPR-Cas systems. Biochemical and in silico evidence indicates that Cas1 is a metal-dependent deoxyribonuclease. Deletion of Cas1 in E. coli increases sensitivity to DNA damage and impairs chromosome segregation, as described in "A dual function of the CRISPR-Cas system in bacterial antivirus immunity and DNA repair," Babu M et al. Mol Microbiol 79:484-502 (2011). Biochemical and in silico evidence indicates that Cas2 is a U-rich region-specific RNase and a double-stranded DNase.
[0064] Aspects of the present invention relate to the identification and engineering of novel effector proteins associated with Class 2 CRISPR-Cas systems. In preferred embodiments, the effector proteins comprise single-subunit effector modules. In further embodiments, the effector proteins function in prokaryotic or eukaryotic cells in in vitro, in vivo, or ex vivo applications. Certain aspects of the present invention encompass computational methods and algorithms for predicting novel Class 2 CRISPR-Cas systems and identifying components therein.
[0065] In one embodiment, a computational method for identifying novel class 2 CRISPR-Cas loci includes the following steps: detecting all contigs encoding the Cas1 protein; identifying all predicted protein-coding genes within 20 kB of the Cas1 gene; comparing the identified genes with a Cas protein-specific profile to predict CRISPR arrays; selecting unclassified candidate CRISPR-Cas loci containing proteins greater than 500 amino acids (>500 aa); and analyzing the selected candidates using PSI-BLAST and HHPred, thereby isolating and identifying novel class 2 CRISPR-Cas loci. In addition to the above steps, further analysis of the candidates can be performed by searching metagenomics databases for additional homologs.
[0066] In one embodiment, the step of detecting all contigs encoding Cas1 proteins is performed by GenemarkS, a gene prediction program as further described in "GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions," John Besemer, Alexandre Lomsadze and Mark Borodovsky, Nucleic Acids Research (2001) 29, pp 2607-2618, incorporated herein by reference.
[0067] In one embodiment, the step of identifying all predicted protein-coding genes is performed by comparing the identified genes to the Cas protein-specific profile and annotating them according to the NCBI Conserved Domain Database (CDD), a protein annotation resource consisting of a collection of fully annotated multiple sequence alignment models of ancient domains and full-length proteins. These are available as position-specific score matrices (PSSMs) for rapid identification of conserved domains in protein sequences via RPS-BLAST. The CDD content includes NCBI-curated domains that explicitly define domain boundaries using three-dimensional structural information and provide insight into sequence / structure / function relationships, as well as domain models imported from several external database sources (Pfam, SMART, COG, PRK, TIGRFAM). In a further embodiment, CRISPR arrays were predicted using the PILER-CR program, a publicly available domain software for finding CRISPR repeats, as described in "PILER-CR: fast and accurate identification of CRISPR repeats," Edgar, RC, BMC Bioinformatics, Jan 20;8:18 (2007), incorporated herein by reference.
[0068] In a further embodiment, a case-by-case analysis is performed using PSI-BLAST (Position-Specific Iterative Basic Local Alignment Search Tool). PSI-BLAST derives a position-specific scoring matrix (PSSM) or profile from a multiple sequence alignment of sequences found above a given score threshold using protein-protein BLAST. This PSSM is used to further search the database for new matches and is subsequently updated to iterate with these newly found sequences. Thus, PSI-BLAST provides a means to detect distant relationships between proteins.
[0069] In another embodiment, case-by-case analysis is performed using HHpred, a sequence database search and structure prediction method that is as easy to use as BLAST or PSI-BLAST, yet much more sensitive in finding distant homologs. In fact, HHpred's sensitivity rivals that of the most powerful structure prediction servers currently available. HHpred is the first server based on pairwise comparison of profile hidden Markov models (HMMs). While most traditional sequence search methods search sequence databases such as UniProt or NR, HHpred searches alignment databases such as Pfam or SMART. This significantly simplifies the list of hits to a few sequence families rather than a cluttered pile of single sequences. All major publicly available profile and alignment databases are accessible in HHpred. HHpred accepts a single query sequence or multiple alignments as input. HHpred returns search results in an easy-to-read format similar to PSI-BLAST within just a few minutes. Search options include local or global alignments and secondary structure similarity scoring. HHpred can generate pairwise query-template sequence alignments, synthetic query-template multiple alignments (eg, for transitive searches), as well as three-dimensional structural models calculated by the MODELLER software from HHpred alignments.
[0070] The term "nucleic acid targeting system," in which the nucleic acid is DNA or RNA and in some embodiments can also refer to a DNA-RNA hybrid or derivative thereof, collectively refers to transcripts and other elements involved in the expression or activity of a DNA- or RNA-targeting CRISPR-associated ("Cas") gene, which may include sequences encoding a DNA- or RNA-targeting Cas protein and a DNA- or RNA-targeting guide RNA, including a CRISPR RNA (crRNA) sequence and (in CRISPR-Cas9 systems, but not all systems) a trans-activating CRISPR-Cas system RNA (tracrRNA) sequence, or other sequences and transcripts from a DNA- or RNA-targeting CRISPR locus. In the Cpf1 DNA-targeting RNA-guided endonuclease system described herein, the tracrRNA sequence is not required. In general, RNA-targeting systems are characterized by elements that promote the formation of an RNA-targeting complex at the site of the target RNA sequence. In the context of forming a DNA or RNA targeting complex, "target sequence" refers to a DNA or RNA sequence to which a DNA or RNA targeting guide RNA is designed to have complementarity, where hybridization between the target sequence and the RNA targeting guide RNA promotes the formation of an RNA targeting complex. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell.
[0071] In certain embodiments of the present invention, the novel DNA targeting system, also referred to as DNA-targeting CRISPR-Cas or CRISPR-Cas DNA targeting system, described herein does not require the creation of customized proteins to target specific DNA sequences, but rather is based on identified type V (e.g., subtype VA and subtype VB) Cas proteins that allow a single effector protein or enzyme to be programmed to recognize a specific DNA target via an RNA molecule, which in turn can be used to recruit the enzyme to a specific DNA target. Aspects of the present invention particularly relate to a DNA-targeting RNA-guided Cpf1 CRISPR system.
[0072] In one embodiment of the present invention, the novel RNA targeting system, also referred to as RNA- or RNA-targeting CRISPR-Cas or CRISPR-Cas-based RNA targeting system, described herein does not require the creation of customized proteins to target specific RNA sequences, but rather is based on identified Type VI Cas proteins that allow a single enzyme to be programmed to recognize a specific RNA target by an RNA molecule, which in turn can be used to recruit the enzyme to a specific RNA target.
[0073] The nucleic acid targeting systems, vector systems, vectors and compositions described herein can be used in a variety of nucleic acid targeting applications, altering or modifying the synthesis of gene products such as proteins, nucleic acid cleavage, nucleic acid editing, nucleic acid splicing; target nucleic acid transport, target nucleic acid tracking, target nucleic acid isolation, target nucleic acid visualization, etc.
[0074] As used herein, Cas protein or CRISPR enzyme refers to any of the proteins represented in a novel classification of CRISPR-Cas systems. In an advantageous embodiment, the present invention encompasses effector proteins identified in type V CRISPR-Cas loci, e.g., the Cpf1-encoding locus designated subtype VA. Currently, subtype VA loci encompass distinct genes designated cas1, cas2, cpf1, and CRISPR arrays. Cpf1 (CRISPR-associated protein Cpf1, subtype PREFRAN) is a large protein (approximately 1300 amino acids) containing a RuvC-like nuclease domain homologous to the corresponding domain in Cas9, along with a corresponding characteristic arginine-rich cluster in Cas9. However, Cpf1 lacks the HNH nuclease domain present in all Cas9 proteins, and in contrast to Cas9, which contains long inserts containing the HNH domain, the RuvC-like domain is continuous in the Cpf1 sequence. Thus, in particular embodiments, the CRISPR-Cas enzyme comprises only a RuvC-like nuclease domain.
[0075] The Cpf1 gene is found in several diverse bacterial genomes, typically at the same locus as the cas1, cas2, and cas4 genes and CRISPR cassettes (e.g., FNFX1_1431-FNFX1_1428 in Francisella cf. novicida Fx1). Therefore, the layout of this putative novel CRISPR-Cas system appears to resemble type II-B. Furthermore, like Cas9, the Cpf1 protein contains a readily identifiable C-terminal region homologous to transposon ORF-B, including an active RuvC-like nuclease, an arginine-rich region, and a zinc finger (absent in Cas9). However, unlike Cas9, Cpf1 is also present in some genomes without a CRISPR-Cas context, and its relatively high similarity to ORF-B suggests that it may be a transposon component. It has been suggested that if this were a bona fide CRISPR-Cas system, Cpf1 would be a functional analog of Cas9 and a novel CRISPR-Cas type, namely, type V (see Annotation and Classification of CRISPR-Cas Systems. Makarova KS, Koonin EV. Methods Mol Biol. 2015;1311:47-75). However, as described herein, Cpf1 is designated subtype VA to distinguish it from C2c1p, which does not have the same domain structure and is therefore designated subtype VB.
[0076] In advantageous embodiments, the present invention encompasses compositions and systems comprising effector proteins identified in the Cpf1 locus designated subtype VA.
[0077] Aspects of the present invention also include methods and uses of the compositions and systems described herein for, for example, altering or manipulating the expression of one or more genes or one or more gene products in genome engineering in vitro, in vivo or ex vivo in prokaryotic or eukaryotic cells.
[0078] In embodiments of the present invention, the terms mature crRNA, guide RNA, and single guide RNA are used interchangeably as in the aforementioned references, such as WO2014 / 093622 (PCT / US2013 / 074667). Generally, a guide sequence is any polynucleotide sequence that has sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more, when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any algorithm suitable for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., the Burrows-Wheeler aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequence is about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length. In some embodiments, the guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12 nucleotides in length, or shorter. Preferably, the guide sequence is 10-30 nucleotides in length. The ability of a guide sequence to direct sequence-specific binding of a CRISPR complex to a target sequence can be assessed by any suitable assay.For example, sufficient components of a CRISPR system to form a CRISPR complex, including the guide sequence to be tested, can be provided to a host cell having a corresponding target sequence, such as by transfection of a vector encoding the components of the CRISPR sequence, followed by evaluation of preferential cleavage within the target sequence, such as by a Surveyor assay as described herein. Similarly, cleavage of a target polynucleotide sequence can be determined in vitro by providing the target sequence, the components of a CRISPR complex, including the guide sequence to be tested, and a control guide sequence that is different from the test guide sequence, and comparing the binding or cleavage rate at the target sequence between reactions with the test guide sequence and the control guide sequence. Other assays are possible and will occur to those skilled in the art. The guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the genome of a cell. Exemplary target sequences include those that are unique to the target genome.
[0079] Generally, and throughout this specification, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules containing one or more free ends or no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other types of polynucleotides known in the art. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, in which the vector contains viral-derived DNA or RNA sequences for packaging into a virus (e.g., retrovirus, replication-deficient retrovirus, adenovirus, replication-deficient adenovirus, and adeno-associated virus). Viral vectors also include polynucleotides carried by viruses for transfection into host cells. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors." Vectors that are used for and result in expression in eukaryotic cells may be referred to herein as "eukaryotic expression vectors." Common expression vectors of utility in recombinant DNA techniques are often in the form of plasmids.
[0080] A recombinant expression vector can comprise a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, meaning that the recombinant expression vector comprises one or more regulatory elements (which may be selected based on the host cell used for expression) operably linked to the nucleic acid sequence to be expressed. Within the scope of a recombinant expression vector, "operably linked" is intended to mean that the nucleotide sequence of interest is linked to one or more regulatory elements in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).
[0081] The term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and polyU sequences). Such regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in specific host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in a desired target tissue, such as muscle, neuron, bone, skin, blood, a specific organ (e.g., liver, pancreas), or a specific cell type (e.g., lymphocyte). Regulatory elements may also direct expression in a time-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which expression may or may not also be tissue- or cell-type-specific. In some embodiments, the vector comprises one or more pol III promoters (e.g., 1, 2, 3, 4, 5, or more pol III promoters), one or more pol II promoters (e.g., 1, 2, 3, 4, 5, or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5, or more pol I promoters), or a combination thereof. Examples of pol III promoters include, but are not limited to, U6 and H1 promoters.Examples of pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer) [see, e.g., Boshart et al., Cell, 41:521-530 (1985)], the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. The term "regulatory element" also encompasses enhancer elements such as the WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), p. 466-472, 1988); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), p. 1527-31, 1981). Those skilled in the art will understand that the design of the expression vector may depend on factors such as the choice of host cell to be transformed and the desired expression level. The vector can be introduced into a host cell, thereby producing transcripts, proteins, or peptides encoded by nucleic acids as described herein, including fusion proteins or peptides (e.g., clustered regularly interspaced short palindromic repeats (CRISPR) transcripts, proteins, enzymes, mutants thereof, fusion proteins thereof, etc.).
[0082] Advantageous vectors include lentiviruses and adeno-associated viruses, and the type of vector can also be selected to target specific cell types.
[0083] As used herein, the terms "crRNA" or "guide RNA" or "single guide RNA" or "sgRNA" or "one or more nucleic acid components" of a Type V CRISPR-Cas locus effector protein include any polynucleotide sequence that has sufficient complementarity with a target nucleic acid sequence to hybridize with the target nucleic acid sequence and direct sequence-specific binding of a nucleic acid targeting complex to the target nucleic acid sequence. In some embodiments, the degree of complementarity is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., the Burrows-Wheeler aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). The ability of a guide sequence (within a nucleic acid targeting guide RNA) to direct sequence-specific binding of a nucleic acid targeting complex to a target nucleic acid sequence can be assessed by any suitable assay. For example, sufficient components of a nucleic acid targeting CRISPR system to form a nucleic acid targeting complex may be provided to a host cell having a corresponding target nucleic acid sequence, including a guide sequence to be tested, such as by transfection of a vector encoding the components of the nucleic acid targeting complex, followed by assessment of preferential targeting (e.g., cleavage) within the target nucleic acid sequence, such as by a Surveyor assay as described herein.Similarly, cleavage of a target nucleic acid sequence can be determined in vitro by providing the target nucleic acid sequence, components of a nucleic acid targeting complex, including the guide sequence to be tested, and a control guide sequence different from the test guide sequence, and comparing the binding or cleavage rate at the target sequence between reactions with the test guide sequence and the control guide sequence. Other assays are also possible and will occur to those skilled in the art. The guide sequence, and therefore the nucleic acid targeting guide RNA, can be selected to target any target nucleic acid sequence. The target sequence may be DNA. The target sequence may be any RNA sequence. In some embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of messenger RNA (mRNA), pre-mRNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), double-stranded RNA (dsRNA), non-coding RNA (ncRNA), long non-coding RNA (lncRNA), and small cytoplasmic RNA (scRNA). In some preferred embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of mRNA, pre-mRNA, and rRNA. In some preferred embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of ncRNA and lncRNA. In some more preferred embodiments, the target sequence may be a sequence within an mRNA molecule or a pre-mRNA molecule.
[0084] In some embodiments, the nucleic acid targeting guide RNA is selected to reduce the degree of secondary structure within the RNA targeting guide RNA. In some embodiments, when optimally folded, no more than about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1%, or even less of the nucleotides of the nucleic acid targeting guide RNA are involved in self-complementary base pairing. Optimal folding can be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculations of minimum Gibbs free energy. An example of such an algorithm is mFold, as described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another exemplary folding algorithm is the online web server RNAfold, developed at the Institute for Theoretical Chemistry at the University of Vienna, which uses a centroid structure prediction algorithm (see, e.g., A.R. Gruber et al., 2008, Cell 106(1):23-24; and P.A. Carr and G.M. Church, 2009, Nature Biotechnology 27(12):1151-62).
[0085] The "tracrRNA" sequence or similar term includes any polynucleotide sequence having sufficient complementarity to hybridize with the crRNA sequence. As noted hereinabove, in embodiments of the present invention, tracrRNA is not required for the cleavage activity of the Cpf1 effector protein complex.
[0086] We also conduct challenge experiments to verify the DNA targeting and cleavage capabilities of Type V / Type VI proteins, such as Cpf1 / C2c1 / C2c2. This experiment is very similar to a similar study on heterologous expression of StCas9 in E. coli (Sapranauskas, R. et al. Nucleic Acids Res 39, 9275-9282 (2011)). We introduce a plasmid containing both a PAM and a resistance gene into heterologous E. coli and then plate it on the corresponding antibiotic. If there is DNA cleavage of the plasmid, we will not observe any surviving colonies.
[0087] In further detail, the assay for DNA targets is as follows: Two E. coli strains are used in this assay. One strain carries a plasmid encoding the endogenous effector protein locus from the bacterial strain. The other strain carries an empty plasmid (e.g., pACYC184, a control strain). All possible 7- or 8-bp PAM sequences are provided on an antibiotic resistance plasmid (pUC19 carrying an ampicillin resistance gene). The PAMs are located adjacent to the sequence of protospacer 1 (the DNA target for the first spacer in the endogenous effector protein locus). Two PAM libraries are cloned: one has 8 random bp 5' to the protospacer (e.g., a total of 65,536 different PAM sequences = complexity); the other library has 7 random bp 3' to the protospacer (e.g., a total complexity of 16,384 different PAMs). Cloning both libraries resulted in an average of 500 plasmids per possible PAM. The 5'PAM and 3'PAM libraries were transformed into test and control strains in separate transformations, and the transformed cells were plated separately on ampicillin plates. Recognition by the plasmid and subsequent cleavage / interference rendered the cells vulnerable to ampicillin, preventing growth. Approximately 12 hours after transformation, all colonies formed by the test and control strains were harvested, and plasmid DNA was isolated. The plasmid DNA was used as a template for PCR amplification, followed by deep sequencing. Expression of all PAMs in the untransformed library indicated the expected expression of PAMs in the transformed cells. Expression of all PAMs found in the control strain indicated actual expression. Expression of all PAMs in the test strain indicated which PAMs were not recognized by the enzyme, and comparison with the control strain allowed for the extraction of depleted PAM sequences.
[0088] In some embodiments of the CRISPR-Cas9 system, the degree of complementarity between the tracrRNA and crRNA sequences is along the length of the shorter of the two when optimally aligned. As described herein, in embodiments of the invention, a tracrRNA is not required. In some embodiments of previously described CRISPR-Cas systems (e.g., CRISPR-Cas9 systems), the design of a chimeric synthetic guide RNA (sgRNA) may incorporate at least a 12-bp duplex structure between the crRNA and tracrRNA; however, in the Cpfl CRISPR system described herein, such a chimeric RNA (chi-RNA) is not possible because this system does not utilize a tracrRNA.
[0089] To minimize toxicity and off-target effects, it may be important to control the concentration of the delivered nucleic acid targeting guide RNA. The optimal concentration of the nucleic acid targeting guide RNA can be determined by testing various concentrations in cell models or non-human eukaryotic animal models and analyzing the degree of modification at potential off-target genomic loci using deep sequencing. The concentration that produces the highest level of on-target modification while minimizing the level of off-target modification should be selected for in vivo delivery. The nucleic acid targeting system is advantageously derived from a type V / type VI CRISPR system. In some embodiments, one or more elements of the nucleic acid targeting system are derived from a specific organism that contains an endogenous RNA targeting system. In a preferred embodiment of the present invention, the RNA targeting system is a type V / type VI CRISPR system. In a specific embodiment, the type V / type VI RNA targeting Cas enzyme is Cpf1 / C2c1 / C2c2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Csel, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof. In embodiments, Type V / Type VI proteins such as Cpf1 / C2c1 / C2c2 as referred to herein also encompass homologs or orthologs of Type V / Type VI proteins such as Cpf1 / C2c1 / C2c2. The terms "orthologue" (also referred to herein as "ortholog") and "homologue" (also referred to herein as "homolog") are well known in the art.As further guidance, a "homolog" of a protein, as used herein, is a protein of the same species that performs the same or similar function as the protein of which it is a homolog. Homologous proteins, however, may be structurally unrelated, or may be only partially related. An "ortholog" of a protein, as used herein, is a protein of a different species that performs the same or similar function as the protein of which it is an ortholog. Orthologous proteins, however, may be structurally unrelated, or may be only partially related. Homologs and orthologs can be identified by homology modeling (e.g., Greer, Science vol. 228 (1985) 1055, and Blundell et al. Eur J Biochem vol 172 (1988), 513) or "structural BLAST" (see Dey F, Cliff Zhang Q, Petrey D, Honig B. "Toward a "structural BLAST": using structural relationships to infer function". Protein Sci. 2013 Apr;22(4):359-66. doi:10.1002 / pro.2225). See also Shmakov et al. (2015) for applications in the field of CRISPR-Cas loci. Homologous proteins, however, may be structurally unrelated or only partially related. In particular embodiments, a homologue or orthologue of Cpfl as referred to herein has at least 80%, more preferably at least 85%, even more preferably at least 90%, such as at least 95%, sequence homology or identity with Cpfl 1. In further embodiments, a homologue or orthologue of Cpfl as referred to herein has at least 80%, more preferably at least 85%, even more preferably at least 90%, such as at least 95% sequence identity with wild-type Cpfl.When Cpf1 has one or more mutations (mutated form), a homologue or orthologue of said Cpf1 as referred to herein has at least 80%, more preferably at least 85%, even more preferably at least 90%, such as at least 95%, sequence identity with the mutant Cpf1.
[0090] In certain embodiments, the V-type Cas protein can be an ortholog of an organism of a genera including, but not limited to, Acidaminococcus sp., Lachnospiraceae bacterium, or Moraxella bovoculi; in particular embodiments, the V-type Cas protein can be an ortholog of an organism of a species including, but not limited to, Acidaminococcus sp. BV3L6; Lachnospiraceae bacterium ND2006 (LbCpf1), or Moraxella bovoculi 237. In particular embodiments, a homologue or orthologue of Cpf1 as referred to herein has at least 80%, more preferably at least 85%, even more preferably at least 90%, such as at least 95% sequence homology or identity to one or more of the Cpf1 sequences disclosed herein. In further embodiments, a homologue or orthologue of Cpf as referred to herein has at least 80%, more preferably at least 85%, even more preferably at least 90%, such as at least 95% sequence identity to wild-type FnCpf1, AsCpf1 or LbCpf1.
[0091] In a particular embodiment, the Cpfl protein of the present invention has at least 60%, more particularly at least 70%, such as at least 80%, more particularly at least 85%, even more particularly at least 90%, such as at least 95% sequence homology or identity to FnCpfl, AsCpfl, or LbCpfl. In a further embodiment, the Cpfl protein as referred to herein has at least 60%, such as at least 70%, more particularly at least 80%, more particularly at least 85%, even more preferably at least 90%, such as at least 95% sequence identity to wild-type AsCpfl or LbCpfl. In a particular embodiment, the Cpfl protein of the present invention has less than 60% sequence identity to FnCpfl. Those skilled in the art will understand that this includes truncated forms of the Cpfl protein, and therefore the sequence identity is determined over the length of the truncated form.
[0092] Some methods for identifying orthologs of CRISPR-Cas system enzymes can include identifying tracr sequences in a genome of interest. Identifying tracr sequences can involve the following steps: searching databases for direct repeat or tracr mate sequences to identify CRISPR regions containing CRISPR enzymes; searching for homologous sequences in the CRISPR region flanking the CRISPR enzyme in both the sense and antisense directions; examining transcription terminators and secondary structures; identifying any sequence that is not a direct repeat or tracr mate sequence but has greater than 50% identity to a direct repeat or tracr mate sequence as a potential tracr sequence; and analyzing the potential tracr sequence for associated transcription termination sequences. In this system, RNA sequencing data revealed that the computationally identified potential tracrRNA was very lowly expressed, suggesting that tracrRNA may not be required for the function of the system. After further evaluation of the FnCpf1 locus and combining the in vitro cleavage results, Applicants concluded that tracrRNA is not required for target DNA cleavage by the Cpf1 effector protein complex. Applicants determined that a Cpf1 effector protein complex containing only the Cpf1 effector protein and crRNA (a guide RNA comprising a direct repeat sequence and a guide sequence) was sufficient to cleave the target DNA.
[0093] It will be understood that any of the functions described herein can be engineered into CRISPR enzymes from other orthologs, including chimeric enzymes comprising fragments from multiple orthologs. Examples of such orthologs are described elsewhere herein. Thus, chimeric enzymes can be engineered into CRISPR enzymes from, but not limited to, the genera Corynebacter, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus, Lactobacillus, Mycoplasma, Bacteroides, Flaviivola, Flavobacterium, and the like. The chimeric enzyme may comprise a fragment of a CRISPR enzyme orthologue from an organism of the genera including: Terium, Sphaerochaeta, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus, Nitratifractor, Mycoplasma, and Campylobacter. The chimeric enzyme may comprise a first fragment and a second fragment, which may be from a CRISPR enzyme orthologue from an organism of a genera listed herein or a species listed herein; advantageously, the fragments are from a CRISPR enzyme orthologue from a different species.
[0094] In embodiments, the Type V / Type VI RNA-targeting effector protein, particularly the Cpf1 / C2c1 / C2c2 protein, as referred to herein also encompasses functional variants of Cpf1 / C2c1 / C2c2 or its homologs or orthologs. A "functional variant" of a protein, as used herein, refers to a variant of such a protein that retains at least part of the activity of the protein. Functional variants may include mutations (which may be insertion, deletion, or substitution mutants), including polymorphisms, etc. Also included within the scope of functional variants are fusion products of such proteins with another, normally unrelated, nucleic acid, protein, polypeptide, or peptide. Functional variants may be naturally occurring or artificial. Advantageous embodiments may include engineered or non-naturally occurring Type V / Type VI RNA-targeting effector proteins, such as Cpf1 / C2c1 / C2c2 or its orthologs or homologs.
[0095] In certain embodiments, one or more nucleic acid molecules encoding a Type V / Type VI RNA-targeting effector protein, particularly Cpf1 / C2c1 / C2c2 or an ortholog or homolog thereof, may be codon-optimized for expression in a eukaryotic cell. The eukaryotic organism may be as discussed herein. The one or more nucleic acid molecules may be engineered or non-naturally occurring.
[0096] In certain embodiments, a Type V / Type VI RNA-targeting effector protein, particularly Cpf1 / C2c1 / C2c2 or its ortholog or homolog, can contain one or more mutations (and thus one or more nucleic acid molecules encoding it can have one or more mutations). Mutations may be artificially introduced and can include, but are not limited to, one or more mutations in the catalytic domain. Examples of catalytic domains associated with the Cas9 enzyme include, but are not limited to, RuvC I, RuvC II, RuvC III, and HNH domains.
[0097] In certain embodiments, Type V / Type VI proteins, such as Cpf1 / C2c1 / C2c2 or their orthologs or homologs, can be used as universal nucleic acid binding proteins fused to or operably linked to functional domains, including, but not limited to, translation initiation factors, translation activators, translation repressors, nucleases, particularly ribonucleases, spliceosomes, beads, light-inducible / regulatory domains, or chemical-inducible / regulatory domains.
[0098] In some embodiments, unmodified nucleic acid targeting effector proteins can have cleavage activity. In some embodiments, RNA targeting effector proteins can direct cleavage of one or both nucleic acid (DNA or RNA) strands at or near the target sequence, such as within the target sequence and / or within the complement of the target sequence or in a sequence associated with the target sequence. In some embodiments, nucleic acid targeting effector proteins can direct cleavage of one or both DNA or RNA strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, cleavage can be cohesive, i.e., result in a cohesive end. In some embodiments, cleavage is a cohesive end cleavage with a 5' overhang of 1 to 5 nucleotides, preferably 4 or 5 nucleotides. In some embodiments, the cleavage site is distant from the PAM; for example, cleavage occurs after the 18th nucleotide on the non-target strand and after the 23rd nucleotide on the target strand (Figure 97A). In some embodiments, the cleavage site occurs after the 18th nucleotide (counting from the PAM) on the non-target strand and after the 23rd nucleotide (counting from the PAM) on the target strand (Figure 97A). In some embodiments, the vector encodes a nucleic acid targeting effector protein that may be mutated relative to the corresponding wild-type enzyme, such that the mutant nucleic acid targeting effector protein lacks the ability to cleave one or both DNA or RNA strands of a target polynucleotide containing a target sequence. As a further example, mutant Cas proteins that lack substantially all DNA cleavage activity may be generated by mutating two or more catalytic domains of a Cas protein (e.g., the RuvC I, RuvC II, and RuvC III or HNH domains of a Cas9 protein).As described herein, the corresponding catalytic domain of the Cpf1 effector protein can also be mutated to create a mutant Cpf1 effector protein that lacks all DNA cleavage activity or has substantially reduced DNA cleavage activity. In some embodiments, a nucleic acid targeting effector protein can be considered to lack substantially all RNA cleavage activity when the RNA cleavage activity of the mutant enzyme is about 25%, 10%, 5%, 1%, 0.1%, 0.01%, or less of the nucleic acid cleavage activity of the non-mutated enzyme; an example would be when the nucleic acid cleavage activity of the mutant is zero or negligible compared to the non-mutated form. Effector proteins can be identified by reference to a general class of enzymes that share homology with the largest nucleases with multiple nuclease domains from Type V / Type VI CRISPR systems. Most preferably, the effector protein is a Type V / Type VI protein, such as Cpf1 / C2c1 / C2c2. In further embodiments, the effector protein is a Type V protein. By derived, Applicants mean that the derived enzyme is generally based on a wild-type enzyme in the sense that it has a high degree of sequence homology to the wild-type enzyme, but that it has been mutated (modified) in some way as known in the art or as described herein.
[0099] Again, the terms Cas and CRISPR enzymes and CRISPR proteins and Cas proteins are generally used interchangeably, and unless otherwise clear, such as by specific reference to Cas9, any reference herein will be understood by analogy to refer to the novel CRISPR effector proteins described further herein. As noted above, many of the residue numbering used herein refers to effector proteins from the Type V / Type VI CRISPR loci. However, it will be understood that the present invention encompasses many more effector proteins from other microbial species. In certain embodiments, the effector protein may be constitutively present, inducibly present, or conditionally present, or may be administered or delivered. Effector protein optimization may be used to enhance function or develop novel functions, and chimeric effector proteins may be created. And as described herein, effector proteins may be engineered for use as universal nucleic acid-binding proteins.
[0100] Typically, in the context of a nucleic acid targeting system, formation of a nucleic acid targeting complex (comprising a guide RNA hybridized to the target sequence and complexed with one or more nucleic acid targeting effector proteins) results in cleavage of one or both DNA or RNA strands within or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50 or more base pairs therefrom). As used herein, the term "one or more sequences associated with a target locus of interest" refers to sequences that are in the vicinity of the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50 or more base pairs therefrom, where the target sequence is contained within the target locus of interest).
[0101] An example of a codon-optimized sequence is one that is optimized for expression in a eukaryote, e.g., a human (i.e., optimized for human expression), or another eukaryote, animal, or mammalian-optimized sequence as discussed herein; see, e.g., the SaCas9 human codon-optimized sequence in WO 2014 / 093622 (PCT / US2013 / 074667) for an example of a codon-optimized sequence. (Given knowledge in the art and this disclosure, codon optimization of one or more encoding nucleic acid molecules, particularly with respect to effector proteins (e.g., Cpf1), is within the skill of the art.) While this is preferred, it is understood that other examples are possible, and codon optimization for host species other than humans, or codon optimization for specific organs, are known. In some embodiments, an enzyme-coding sequence encoding a DNA / RNA-targeting Cas protein is codon-optimized for expression in a specific cell, e.g., a eukaryotic cell. Eukaryotic cells may be of or derived from specific organisms, such as plants or mammals, including but not limited to humans, or non-human eukaryotic organisms or animals or mammals as discussed herein, such as mice, rats, rabbits, dogs, livestock, or non-human mammals or primates. In some embodiments, methods of modifying human germline genetic identity and / or methods of modifying animal genetic identity that may cause suffering to humans or animals without any substantial medical benefit to them, as well as animals resulting from such methods, may be excluded. Generally, codon optimization refers to a method of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) of the native sequence with a codon that is more frequently or most frequently used in the genes of that host cell while maintaining the native amino acid sequence. Different species exhibit specific biases for certain codons for specific amino acids.Codon bias (differences in codon usage among organisms) is often correlated with the efficiency of messenger RNA (mRNA) translation, which in turn is thought to depend, among other things, on the properties of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell generally reflects the codons most frequently used in peptide synthesis. Therefore, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the "Codon Usage Database," available at www.kazusa.orjp / codon / , and these tables can be adapted in several ways. See Nakamura, Y., et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000," Nucl. Acids Res. 28:292 (2000). Computer algorithms are also available for codon-optimizing a particular sequence for expression in a particular host cell, such as Gene Forge (Aptagen; Jacobus, PA). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more, or all codons) in a sequence encoding a DNA / RNA-targeting Cas protein correspond to the most frequently used codon for a particular amino acid. For codon usage in yeast, see the online yeast genome database available at http: / / www.yeastgenome.org / community / codon_usage.shtml, or "Codon selection in yeast," Bennetzen and Hall, J Biol Chem. 1982 Mar 25;257(6):3026-31.For codon usage in plants, including algae, see "Codon usage in higher plants, green algae, and cyanobacteria," Campbell and Gowri, Plant Physiol. 1990 Jan;92(1):1-11; and "Codon usage in plant genes," Murray et al, Nucleic Acids Res. 1989 Jan 25;17(2):477-98; or "Selection on the codon bias of chloroplast and cyanelle genes in different plant and algal lineages," Morton BR, J Mol Evol. 1998 Apr;46(4):449-59.
[0102] In some embodiments, the vector encodes a nucleic acid targeting effector protein, e.g., a Type V / Type VI RNA targeting effector protein, particularly Cpf1 / C2c1 / C2c2 or an ortholog or homolog thereof, that comprises one or more nuclear localization sequences (NLSs), such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs. In some embodiments, the RNA targeting effector protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs at or near the amino terminus, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs at or near the carboxy terminus, or a combination thereof (e.g., zero or at least one NLS at the amino terminus and zero or one or more NLSs at the carboxy terminus). When more than one NLS is present, each may be selected independently of the others, and thus a single NLS may be present in two or more copies and / or in combination with one or more other NLSs present in one or more copies. In some embodiments, an NLS is considered to be near the N- or C-terminus when the nearest amino acid of the NLS is within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the N- or C-terminus.Non-limiting examples of NLSs include the NLS of the SV40 virus large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO: 2); the NLS of nucleoplasmin (e.g., the nucleoplasmin bipartite NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 3)); the c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 4) or RQRRNELKRSP (SEQ ID NO: 5); the hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 6); the IBB domain of importin-α having the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 7); the sequences VSRKRPRP (SEQ ID NO: 8) and PPKKARED (SEQ ID NO: 9) of the fibroid T protein; the sequence PQPKKKPL (SEQ ID NO: 10) of human p53; and the mouse c-abl Examples of NLS sequences include those derived from the sequence SALIKKKKKMAP (SEQ ID NO: 11) of IV; the sequences DRLRR (SEQ ID NO: 12) and PKQKKRK (SEQ ID NO: 13) of influenza virus NS1; the sequence RKLKKKIKKL (SEQ ID NO: 14) of hepatitis virus delta antigen; the sequence REKKKFLKRR (SEQ ID NO: 15) of mouse Mx1 protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 16) of human poly(ADP-ribose) polymerase; and the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 17) of steroid hormone receptor (human) glucocorticoid. Generally, one or more NLSs are strong enough to drive the accumulation of detectable amounts of DNA / RNA-targeting Cas proteins in the nucleus of eukaryotic cells. Generally, the strength of nuclear localization activity can be derived from the number of NLSs in the nucleic acid-targeting effector protein, the specific NLSs used, or a combination of these factors. Detection of nuclear accumulation can be performed by any suitable technique. For example, a detectable marker may be fused to the nucleic acid targeting protein, thereby visualizing its location within the cell, for example, by combining it with a means of detecting nuclear location (e.g., a nuclear-specific stain such as DAPI).Cell nuclei may also be isolated from cells, and their contents may then be analyzed by any suitable protein detection method, such as immunohistochemistry, Western blot, or enzyme activity assay. Nuclear accumulation may also be determined indirectly, such as by assaying for the effect of nucleic acid-targeting complex formation (e.g., assaying for DNA or RNA cleavage or mutation at the target sequence, or assaying for altered gene expression activity affected by DNA or RNA-targeting complex formation and / or DNA or RNA-targeting Cas protein activity), compared to a control not exposed to the nucleic acid-targeting Cas protein or nucleic acid-targeting complex, or a control exposed to a nucleic acid-targeting Cas protein lacking one or more NLSs. In preferred embodiments of the Cpfl effector protein complexes and systems described herein, the codon-optimized Cpfl effector protein comprises an NLS added to the C-terminus of the protein. In certain embodiments, other localization tags may be fused to the Cas protein, such as to localize Cas to specific sites within the cell, such as, without limitation, organelles, e.g., mitochondria, plastids, chloroplasts, vesicles, Golgi, (nuclear or cell) membranes, ribosomes, nucleoli, ER, cytoskeleton, vacuoles, centrosomes, nucleosomes, granules, centrioles, etc.
[0103] In some embodiments, one or more vectors that drive the expression of one or more elements of nucleic acid targeting system are introduced into host cells, and the expression of the elements of nucleic acid targeting system leads to the formation of nucleic acid targeting complexes at one or more target sites.For example, the nucleic acid targeting effector enzyme and the nucleic acid targeting guide RNA can each be operably linked to separate regulatory elements on separate vectors.One or more RNAs of the nucleic acid targeting system can be delivered to transgenic nucleic acid targeting effector protein animals or mammals, for example, animals or mammals that constitutively, inducibly, or conditionally express nucleic acid targeting effector protein; or animals or mammals that naturally express nucleic acid targeting effector protein or have cells that contain nucleic acid targeting effector protein, by pre-administering one or more vectors that encode and express nucleic acid targeting effector protein in vivo. Alternatively, two or more of the elements expressed from the same or different regulatory elements may be combined into a single vector, with one or more additional vectors providing any components of the nucleic acid targeting system not included in the first vector. Nucleic acid targeting system elements combined into a single vector may be arranged in any suitable orientation, such as an element being 5' ("upstream") or 3' ("downstream") relative to the second element. The coding sequence of an element may be located on the same or opposite strand as the coding sequence of the second element, and may be oriented in the same or opposite direction. In some embodiments, a single promoter drives the expression of transcripts encoding nucleic acid targeting effector proteins and nucleic acid targeting guide RNAs integrated within one or more intron sequences (e.g., each in a different intron, two or more in at least one intron, or all in a single intron). In some embodiments, the nucleic acid targeting effector protein and nucleic acid targeting guide RNA may be operably linked to and expressed from the same promoter.Delivery vehicles, vectors, particles, nanoparticles, formulations, and components thereof for expressing one or more elements of a nucleic acid targeting system are as described in the aforementioned documents, such as International Publication No. WO 2014 / 093622 (PCT / US2013 / 074667). In some embodiments, a vector contains one or more insertion sites (also referred to as "cloning sites"), such as restriction endonuclease recognition sequences. In some embodiments, one or more insertion sites (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more insertion sites) are located upstream and / or downstream of one or more sequence elements of one or more vectors. The use of multiple different guide sequences allows a single expression construct to target nucleic acid targeting activity to multiple different corresponding target sequences within a cell. For example, a single vector can contain about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more guide sequences. In some embodiments, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more such guide sequence-containing vectors may be provided and optionally delivered to cells. In some embodiments, the vector comprises a regulatory element operably linked to an enzyme-coding sequence encoding a nucleic acid targeting effector protein. The nucleic acid targeting effector protein or one or more nucleic acid targeting guide RNAs may be delivered separately; and advantageously, at least one of these is delivered via a particle complex. The nucleic acid targeting effector protein mRNA may be delivered prior to the nucleic acid targeting guide RNA to allow time for the nucleic acid targeting effector protein to be expressed. The nucleic acid targeting effector protein mRNA may be administered 1 to 12 hours (preferably about 2 to 6 hours) before administration of the nucleic acid targeting guide RNA. Alternatively, the nucleic acid targeting effector protein mRNA and the nucleic acid targeting guide RNA may be administered together.Advantageously, a second booster dose of guide RNA may be administered 1-12 hours (preferably about 2-6 hours) after the initial administration of nucleic acid targeting effector protein mRNA + guide RNA. Additional administrations of nucleic acid targeting effector protein mRNA and / or guide RNA may be useful to achieve the most efficient level of genome modification.
[0104] In one aspect, the present invention provides a method for using one or more elements of a nucleic acid targeting system. The nucleic acid targeting complexes of the present invention provide an effective means for modifying target DNA or RNA (single-stranded or double-stranded, linear or supercoiled). The nucleic acid targeting complexes of the present invention have a wide range of uses, including modifying (e.g., deleting, inserting, translocating, inactivating, or activating) target DNA or RNA in numerous cell types. Thus, the nucleic acid targeting complexes of the present invention have broad applicability, for example, in gene therapy, drug screening, disease diagnosis, and prognosis. An exemplary nucleic acid targeting complex comprises a DNA or RNA targeting effector protein complexed with a guide RNA that hybridizes to a target sequence within a target locus of interest.
[0105] In one embodiment, the present invention provides a method for cleaving a target RNA. The method may include modifying the target RNA using a nucleic acid targeting complex that binds to the target RNA and causes cleavage of the target DNA. In some embodiments, the nucleic acid targeting complex of the present invention, when introduced into a cell, can create a break (e.g., a single- or double-strand break) in an RNA sequence. For example, the method can be used to cleave a disease RNA in a cell. For example, an exogenous RNA template flanked by upstream and downstream sequences from a sequence to be integrated can be introduced into a cell. These upstream and downstream sequences share sequence similarity with both sides of the integration site in the RNA. Optionally, the donor RNA can be mRNA. The exogenous RNA template includes a sequence to be integrated (e.g., a mutant RNA). The integration sequence can be a sequence endogenous or exogenous to the cell. Examples of sequences to be integrated include protein-coding RNA or non-coding RNA (e.g., microRNA). Thus, the integration sequence can be operably linked to one or more appropriate regulatory sequences. Alternatively, the integration sequence can provide a regulatory function. The upstream and downstream sequences in the exogenous RNA template are selected to promote recombination between the RNA sequence of interest and the donor RNA. The upstream sequence is an RNA sequence that shares sequence similarity with the RNA sequence upstream of the target integration site. Similarly, the downstream sequence is an RNA sequence that shares sequence similarity with the RNA sequence downstream of the target integration site. The upstream and downstream sequences in the exogenous RNA template can have 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with the target RNA sequence. Preferably, the upstream and downstream sequences in the exogenous RNA template have about 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the target RNA sequence. In some methods, the upstream and downstream sequences in the exogenous RNA template have about 99% or 100% sequence identity with the target RNA sequence.The upstream or downstream sequence can comprise from about 20 bp to about 2500 bp, e.g., about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, or 2500 bp. In some methods, exemplary upstream or downstream sequences have a length of from about 200 bp to about 2000 bp, from about 600 bp to about 1000 bp, or more specifically, from about 700 bp to about 1000 bp. In some methods, the exogenous RNA template can further comprise a marker. Such a marker can facilitate screening for targeted integration. Examples of suitable markers include restriction sites, fluorescent proteins, or selectable markers. The exogenous RNA templates of the present invention can be constructed using recombinant techniques (see, e.g., Sambrook et al., 2001 and Ausubel et al., 1996). In methods for modifying target RNA by incorporating an exogenous RNA template, a nucleic acid targeting complex introduces a break in the DNA or RNA sequence (e.g., a double- or single-stranded break in double- or single-stranded DNA or RNA), and upon repair of the break by homologous recombination with the exogenous RNA template, the template is integrated into the RNA target. The presence of the double-stranded break promotes integration of the template. In another embodiment, the present invention provides a method for modifying RNA expression in a eukaryotic cell. The method includes increasing or decreasing expression of a target polynucleotide using a nucleic acid targeting complex that binds to DNA or RNA (e.g., mRNA or pre-mRNA). In some methods, the target RNA can be inactivated, resulting in modified expression in the cell. For example, when the RNA targeting complex binds to a target sequence in a cell, the target RNA is inactivated, so that the sequence is not translated, the encoded protein is not produced, or the sequence does not function like the wild-type sequence.For example, the sequence encoding a protein or microRNA can be inactivated so that the protein, microRNA, or pre-microRNA transcript is not produced.The target RNA of the RNA targeting complex can be any RNA that is endogenous or exogenous to eukaryotic cells.For example, the target RNA may be RNA present in the nucleus of a eukaryotic cell. The target RNA may be a sequence (e.g., mRNA or pre-mRNA) that encodes a gene product (e.g., a protein) or a non-coding sequence (e.g., ncRNA, lncRNA, tRNA, or rRNA). Examples of target RNA include sequences associated with biochemical signaling pathways, such as biochemical signaling pathway-associated RNA. Examples of target RNA include disease-associated RNA. "Disease-associated" RNA refers to any RNA whose translation product is produced at an abnormal level or in an abnormal form in cells derived from diseased tissue compared to non-diseased control tissues or cells. It may be RNA transcribed from a gene that becomes expressed at an abnormally high level; it may be RNA transcribed from a gene that becomes expressed at an abnormally low level, where the change in expression correlates with the incidence and / or progression of the disease. Disease-associated RNA also refers to RNA transcribed from a gene that has one or more mutations or genetic variations that are directly involved in the pathogenesis of the disease or that are in linkage disequilibrium with one or more genes involved in the pathogenesis of the disease. The translation product may be known or unknown, and may be at a normal or abnormal level. The target RNA of the RNA targeting complex may be any RNA endogenous or exogenous to a eukaryotic cell. For example, the target RNA may be RNA present in the nucleus of a eukaryotic cell. The target RNA may be a sequence (e.g., mRNA or pre-mRNA) that encodes a gene product (e.g., a protein) or a non-coding sequence (e.g., ncRNA, lncRNA, tRNA, or rRNA).
[0106] In some embodiments, the method can include binding a nucleic acid targeting complex to a target DNA or RNA to cause cleavage of the target DNA or RNA, thereby modifying the target DNA or RNA, wherein the nucleic acid targeting complex comprises a nucleic acid targeting effector protein complexed with a guide RNA that hybridizes to a target sequence within the target DNA or RNA. In one aspect, the present invention provides a method for modifying DNA or RNA expression in a eukaryotic cell. In some embodiments, the method includes binding a nucleic acid targeting complex to DNA or RNA, wherein the binding causes an increase or decrease in expression of the DNA or RNA; wherein the nucleic acid targeting complex comprises a nucleic acid targeting effector protein complexed with a guide RNA. Similar considerations and conditions apply to methods of modifying target DNA or RNA, as described above. Indeed, these sampling, culturing, and reintroduction options apply to all aspects of the present invention. In one aspect, the present invention provides a method for modifying target DNA or RNA in a eukaryotic cell, which may be in vivo, ex vivo, or in vitro. In some embodiments, the method comprises sampling a cell or cell population from a human or non-human animal and modifying one or more cells. Culturing can be performed ex vivo at any stage. The one or more cells may then be reintroduced into the non-human animal or plant. For reintroduced cells, it is particularly preferred that the cells are stem cells.
[0107] Indeed, in any embodiment of the present invention, the nucleic acid targeting complex may comprise a nucleic acid targeting effector protein complexed with a guide RNA that hybridizes to a target sequence.
[0108] The present invention relates to the engineering and optimization of systems, methods, and compositions used to control gene expression involving DNA or RNA sequence targeting, including nucleic acid targeting systems and their components. In an advantageous embodiment, the effector enzyme is a type V / type VI protein, such as Cpf1 / C2c1 / C2c2. An advantage of this method is that the CRISPR system minimizes or avoids off-target binding and the resulting side effects. This is achieved by using a system configured to have high sequence specificity for target DNA or RNA.
[0109] For nucleic acid targeting complexes or systems, preferably, the crRNA sequence has one or more stem-loops or hairpins and is 30 nucleotides or more, 40 nucleotides or more, or 50 nucleotides or more in length; the crRNA sequence is 10-30 nucleotides in length, and the nucleic acid targeting effector protein is a Type V / Type VI Cas enzyme. In a specific embodiment, the crRNA sequence is 42-44 nucleotides in length, and the nucleic acid targeting Cas protein is Cpf1 of Francisella tularensis subsp. novocida U112. In a specific embodiment, the crRNA comprises, consists essentially of, or consists of a 19-nucleotide direct repeat and a 23-25 nucleotide spacer sequence, and the nucleic acid targeting Cas protein is Cpf1 of Francisella tularensis subsp. novocida U112.
[0110] The use of two different aptamers (each associated with a distinct nucleic acid-targeting guide RNA) allows for the use of an activator-adapter protein fusion and a repressor-adapter protein fusion with different nucleic acid-targeting guide RNAs to activate expression of one DNA or RNA while suppressing expression of another DNA or RNA. These, along with their different guide RNAs, can be administered together, or substantially together, in a multiplexed manner. Because a relatively small number of effector protein molecules can be used with multiple modified guides, it is sufficient to deliver only one (or at least a minimal number) effector protein molecule while many such modified nucleic acid-targeting guide RNAs, e.g., 10, 20, or 30, can all be used simultaneously. The adapter protein may be associated (preferably linked or fused) with one or more activators or one or more repressors. For example, the adapter protein may be associated with a first activator and a second activator. The first and second activators may be the same, but are preferably different activators. More than two or even more than three activators (or repressors) may be used, although package size may limit the number to more than five different functional domains. As compared to direct fusion with an adaptor protein, preferably a linker is used, in which two or more functional domains are associated with the adaptor protein. Suitable linkers may include a GlySer linker.
[0111] It is also envisioned that the nucleic acid targeting effector protein-guide RNA complex as a whole may be associated with more than one functional domain. For example, there may be more than one functional domain associated with the nucleic acid targeting effector protein, or there may be more than one functional domain associated with the guide RNA (via one or more adaptor proteins), or there may be one or more functional domains associated with the nucleic acid targeting effector protein and one or more functional domains associated with the guide RNA (via one or more adaptor proteins).
[0112] The fusion between the adaptor protein and the activator or repressor may include a linker. For example, the GlySer linker GGGS (SEQ ID NO: 18) can be used. These can be used in repeats of 3 ((GGGGS)3 (SEQ ID NO: 19)), 6 (SEQ ID NO: 20), 9 (SEQ ID NO: 21), or even 12 (SEQ ID NO: 22) or more to provide the appropriate length, as needed. Linkers can be used between the guide RNA and the functional domain (activator or repressor) or between the nucleic acid-targeting Cas protein (Cas) and the functional domain (activator or repressor). With linkers, the user can engineer the appropriate amount of "mechanical flexibility."
[0113] The present invention encompasses nucleic acid targeting complexes comprising a nucleic acid targeting effector protein and a guide RNA, wherein the nucleic acid targeting effector protein comprises at least one mutation (such that the nucleic acid targeting effector protein has 5% or less of the activity of a nucleic acid targeting effector protein without the at least one mutation), and optionally at least one or more nuclear localization sequences; the guide RNA comprises a guide sequence capable of hybridizing to a target sequence in an RNA of interest in a cell; and wherein: the nucleic acid targeting effector protein is associated with two or more functional domains; or at least one loop of the guide RNA is modified by the insertion of one or more distinct RNA sequences that bind to one or more adaptor proteins, and wherein the adaptor proteins are associated with two or more functional domains; or the nucleic acid targeting Cas protein is associated with one or more functional domains, and at least one loop of the guide RNA is modified by the insertion of one or more distinct RNA sequences that bind to one or more adaptor proteins, and wherein the adaptor proteins are associated with one or more functional domains.
[0114] In one aspect, the present invention provides a method for generating a model eukaryotic cell containing a mutated disease gene. In some embodiments, the disease gene is any gene associated with an increased risk of developing or contracting a disease. In some embodiments, the method includes the steps of (a) introducing one or more vectors into a eukaryotic cell, where the one or more vectors drive expression of one or more of a Cpf1 enzyme and a protected guide RNA comprising a guide sequence linked to a direct repeat sequence; and (b) binding a CRISPR complex to a target polynucleotide to cause cleavage of the target polynucleotide within the disease gene, where the CRISPR complex comprises a Cpf1 enzyme complexed with a guide RNA comprising a sequence that hybridizes to a target sequence within the target polynucleotide, thereby generating a model eukaryotic cell containing a mutated disease gene. In some embodiments, the cleavage comprises cleavage of one or both strands at the location of the target sequence by the Cpf1 enzyme. In some embodiments, the cleavage results in decreased transcription of the target gene. In some embodiments, the method further comprises repairing the cleaved target polynucleotide with an exogenous template polynucleotide by non-homologous end joining (NHEJ)-based gene insertion, wherein the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides in the target polynucleotide. In some embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence.
[0115] In some embodiments, the present invention provides a method as discussed herein, wherein the host is a eukaryotic cell. In some embodiments, the present invention provides a method as discussed herein, wherein the host is a mammalian cell. In some embodiments, the present invention provides a method as discussed herein, wherein the host is a non-human eukaryotic cell. In some embodiments, the present invention provides a method as discussed herein, wherein the non-human eukaryotic cell is a non-human mammalian cell. In some embodiments, the present invention provides a method as discussed herein, wherein the non-human mammalian cell may include, but is not limited to, a primate, bovine, ovine, porcine, canine, rodent, lagomorph, such as a monkey, cow, sheep, pig, dog, rabbit, rat, or mouse cell. In some embodiments, the present invention provides a method as discussed herein, wherein the cell may be a non-mammalian eukaryotic cell, such as a poultry (e.g., chicken), vertebrate fish (e.g., salmon), or crustacean (e.g., oyster, clam, lobster, shrimp) cell. In an aspect, the invention provides a method as discussed herein, wherein the non-human eukaryotic cell is a plant cell. The plant cell may be of a monocotyledonous or dicotyledonous plant or of a crop or cereal plant, such as cassava, maize, sorghum, soybean, wheat, oat, or rice. The plant cell may also be that of an algae, a tree or a productive plant, a fruit or vegetable (e.g., a citrus tree, such as an orange, grapefruit or lemon tree; a peach or nectarine tree; an apple or pear tree; a nut tree, such as an almond, walnut or pistachio tree; a Solanaceae plant; a Brassica plant; a Lactuca plant; a Spinacia plant; a Capsicum plant; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.).
[0116] In one aspect, the present invention provides a method for developing a biologically active agent that modulates a cell signaling event associated with a disease gene. In some embodiments, the disease gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the method comprises the steps of: (a) contacting a model cell according to any one of the above embodiments with a test compound; and (b) detecting a readout change indicative of a decrease or increase in a cell signaling event associated with the mutation in the disease gene, thereby developing the biologically active agent that modulates the cell signaling event associated with the disease gene.
[0117] In one aspect, the present invention provides a method for selecting one or more cells by introducing one or more mutations into a gene in the one or more cells, the method comprising the steps of: introducing into the one or more cells one or more vectors, wherein the one or more vectors drive expression of one or more of Cpf1, a guide sequence linked to a direct repeat sequence, and an editing template; wherein the editing template comprises one or more mutations that abolish Cpf1 cleavage; allowing the editing template to homologously recombine with a target polynucleotide in the one or more cells to be selected; and allowing a Cpf1 CRISPR-Cas complex to bind to the target polynucleotide, resulting in cleavage of the target polynucleotide within the gene, wherein the Cpf1 CRISPR-Cas complex comprises (1) a guide sequence that hybridizes to a target sequence within the target polynucleotide, and (2) Cpf1 complexed with a direct repeat sequence, wherein Cpf1 Binding of the CRISPR-Cas complex to the target polynucleotide induces cell death, thereby allowing for the selection of one or more cells into which one or more mutations have been introduced; which contain the split Cpf1. In another preferred embodiment of the invention, the selected cell may be a eukaryotic cell. Aspects of the invention allow for the selection of specific cells without the need for a two-step process that may include a selection marker or counter-selection system. In particular embodiments, the model eukaryotic cell is contained within a model eukaryotic organism.
[0118] In one aspect, the present invention provides a recombinant polynucleotide comprising a guide sequence downstream of a direct repeat sequence, wherein the guide sequence, upon expression, directs sequence-specific binding of a Cpf1 CRISPR-Cas complex to a corresponding target sequence present in a eukaryotic cell. In some embodiments, the target sequence is a viral sequence present in a eukaryotic cell. In some embodiments, the target sequence is a proto-oncogene or an oncogene.
[0119] In one aspect, the invention provides a vector system or eukaryotic host cell comprising: (a) a first regulatory element operably linked to a direct repeat sequence and one or more insertion sites for inserting one or more guide sequences (including any of the modified guide sequences described herein) downstream of a DR sequence, where the guide sequence, upon expression, directs sequence-specific binding of a Cpf1 CRISPR-Cas complex to a target sequence in a eukaryotic cell, wherein the Cpf1 CRISPR-Cas complex comprises Cpf1 (including any of the modifying enzymes described herein) complexed with the guide sequence that hybridizes to the target sequence (and optionally the DR sequence); and / or (b) a second regulatory element operably linked to an enzyme-coding sequence encoding the Cpf1 enzyme that comprises a nuclear localization sequence and / or an NES. In some embodiments, the host cell comprises components (a) and (b). In some embodiments, components (a), (b), or (a) and (b) are stably integrated into the genome of the host eukaryotic cell. In some embodiments, component (a) further comprises two or more guide sequences operably linked to the first regulatory element, wherein each of the two or more guide sequences, upon expression, directs sequence-specific binding of the Cpfl CRISPR-Cas complex to a different target sequence in a eukaryotic cell. In some embodiments, the CRISPR enzyme comprises one or more nuclear localization sequences and / or nuclear export sequences or NESs of sufficient strength to drive accumulation of said CRISPR enzyme in detectable amounts in and / or outside the nucleus of a eukaryotic cell.In some embodiments, the Cpf1 enzyme is selected from the group consisting of Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, or Porphyromonas macacae macacae Cpf1 (including any of the modified enzymes as described herein), may contain additional Cpf1 changes or mutations, and may be chimeric Cpf1. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs one or two strand cleavages at the target sequence. In preferred embodiments, the strand cleavage is a sticky-end cleavage with a 5' overhang.In some embodiments, Cpf1 lacks DNA strand cleavage activity (e.g., 5% or less nuclease activity compared to a wild-type enzyme or an enzyme without a mutation or alteration that reduces nuclease activity). In some embodiments, the first regulatory element is a polymerase III promoter. In some embodiments, the second regulatory element is a polymerase II promoter. In some embodiments, the direct repeat has a minimum length of 16 nt and a single stem-loop. In further embodiments, the direct repeat has a length greater than 16 nt, preferably greater than 17 nt, and has two or more stem-loops or optimized secondary structures. In some embodiments, the guide sequence is at least 16, 17, 18, 19, 20, 25 nucleotides, or 16-30, or 16-25, or 16-20 nucleotides in length.
[0120] In one aspect, the invention provides kits comprising one or more of the components described herein. In some embodiments, the kits comprise a vector system or host cell as described herein and instructions for use of the kit.
[0121] Modified Cpf1 enzyme Computational analysis of the primary structure of Cpf1 nuclease reveals three distinct regions (Fig. 1): a C-terminal RuvC-like domain, which is the only functionally characterized domain; an N-terminal α-helical region; and a mixed α and β region located between the RuvC-like domain and the α-helical region.
[0122] Several small stretches of unstructured regions are predicted within the Cpf1 primary structure. Unstructured regions that are solvent-exposed and not conserved among different Cpf1 orthologs are favorable sides for splitting and inserting small protein sequences (Figures 2 and 3). Additionally, these sides can be used to create chimeric proteins between Cpf1 orthologs.
[0123] Based on the above information, mutations can be created that lead to inactivation of the enzyme or that modify the double-stranded nuclease to nickase activity. In alternative embodiments, this information is used to develop enzymes with reduced off-target effects (as described elsewhere herein).
[0124] In certain of the above-mentioned Cpfl enzymes, the enzyme is modified by mutation of one or more residues, including but not limited to, positions D917, E1006, E1028, D1227, D1255A, N1257, based on the FnCpfl protein or any corresponding orthologue. In certain aspects, the invention provides a composition as discussed herein, wherein the Cpfl enzyme is an inactivated enzyme comprising one or more mutations selected from the group consisting of D917A, E1006A, E1028A, D1227A, D1255A, N1257A, D917A, E1006A, E1028A, D1227A, D1255A, and N1257A, based on the corresponding positions in the FnCpfl protein or Cpfl orthologue. In certain aspects, the invention provides a composition as discussed herein, wherein the CRISPR enzyme comprises D917, or E1006 and D917, or D917 and D1255, based on the corresponding positions in the FnCpf1 protein or Cpf1 ortholog.
[0125] In particular of the above-mentioned Cpf1 enzymes, the enzyme may be, but is not limited to, AsCpf1 (Acidaminococcus sp. The RuvC domain is modified by mutation of one or more residues, including positions R909, R912, R930, R947, K949, R951, R955, K965, K968, K1000, K1002, R1003, K1009, K1017, K1022, K1029, K1035, K1054, K1072, K1086, R1094, K1095, K1109, K1118, K1142, K1150, K1158, K1159, R1220, R1226, R1242, and / or R1252, based on the amino acid position numbering of RuvC domain (BV3L6).
[0126] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of one or more residues (within the RAD50 domain), including, but not limited to, positions K324, K335, K337, R331, K369, K370, R386, R392, R393, K400, K404, K406, K408, K414, K429, K436, K438, K459, K460, K464, R670, K675, R681, K686, K689, R699, K705, R725, K729, K739, K748, and / or K752 relative to the amino acid position numbering of AsCpf1 (Acidaminococcus sp. BV3L6).
[0127] In certain Cpf1 enzymes, the enzyme is modified by mutation of one or more residues including, but not limited to, positions R912, T923, R947, K949, R951, R955, K965, K968, K1000, R1003, K1009, K1017, K1022, K1029, K1072, K1086, F1103, R1226, and / or R1252 based on the amino acid position numbering of AsCpf1 (Acidaminococcus sp. BV3L6).
[0128] In certain embodiments, the Cpf1 enzyme is modified by mutation of one or more residues, including but not limited to, positions R833, R836, K847, K879, K881, R883, R887, K897, K900, K932, R935, K940, K948, K953, K960, K984, K1003, K1017, R1033, R1138, R1165, and / or R1252, based on the amino acid position numbering of LbCpf1 (Lachnospiraceae bacterium ND2006).
[0129] In certain embodiments, the Cpf1 enzyme is selected from the group consisting of, but not limited to, positions K15, R18, K26, Q34, R43, K48, K51, R56, R84, K85, K87, N93, R103, N104, T118, K123, K134, R176, K177, R192, K200, K226, K273, K280, K290, K300, K310, K320, K330, K340, K350, K360, K370, K380, K390, K410, K420, K430, K440, K450, K460, K470, K480, K510, R56, R84, K85, K87, N93, R103, N104, T118, K123, K134, R176, K177, R192, K200, K226, K273, K280, K29 ...00, K420, K430, K440, K450, K460, K470, K480, K510, R56, R84, K85, K8 75, T291, R301, K307, K369, S404, V409, K414, K436, K438, K468, D482, K516, R51 8, K524, K530, K532, K548, K559, K570, R574, K592, D596, K603, K607, K613, C647, R681, K686, H720, K739, K748, K757, T766, K780, R790, P791, K796, K809, K815, T 816, K860, R862, R863, K868, K897, R909, R912, T923, R947, K949, R951, R955, K9 65, K968, K1000, R1003, K1009, K1017, K1022, K1029, A1053, K1072, K1086, F1103, S1209, R1226, R1252, K1273, K1282, and / or K1288.
[0130] In certain embodiments, the enzyme may be, but is not limited to, FnCpf1 (Francisella novicida) Based on the amino acid numbering of the genus U112), the following positions are used: K15, R18, K26, R34, R43, K48, K51, K56, K87, K88, D90, K96, K106, K107, K120, Q125, K143, R186, K187, R202, K210, K235, K296, K298, K314, K320, K326, K397, K444, K449, E454, A483, E491, K527, K541, K581, R583, K589, K595, K597, K613, K624, K635, K639, K656, K660, K671, K672, K673, K674, K675, K676, K677, K678, K679, K680, K681, K682, K683, K684, K685, K686, K687, K688, K689, K690, K691, K692, K693, K694, K695, K696, K697, K698, K699, K700, K701, K702, K703, K704, K705, K706, K707, K710, K711, K712, K713, K714, K715, K716, K717, K718, K720, K721, K722, K723, K724, K725, K730, K731, and / or K1098.
[0131] In certain embodiments, the enzyme may comprise, but is not limited to, the amino acid sequence of positions K15, R18, K26, K34, R43, K48, K51, R56, K83, K84, R86, K92, R102, K103, K116, K121, R158, R159, R174, R182, K206, K251, R264, R274, R282, R294, R306, R310, R320, R332, R340, R352, R360, R374, R382, R394, R406, R410, R420, R430, R440, R452, R460, R470, R480, R490, R510, R520, R530, R540, R550, R560, R570, R580, R590, R600, R610, R620, R630, R640, R650, R660, R670, R680, R690, R710, R720, R730, R740, R750, R760, R770, R780, R790, R800, R810, R820, R830, R840, R860, R920, R930, R940, R950, R960, R970, R980, R990, R990, R991, R992, R993, R994, R995, R996, R997, R998, R999, R999, R999, R9 , K253, K269, K271, K278, P342, K380, R385, K390, K415, K421, K457, K471, A506 , R508, K514, K520, K522, K538, Y548, K560, K564, K580, K584, K591, K595, K601, K634, K640, R645, K679, K689, K707, T716, K725, R737, R747, R748, K753, K768, K774, K775, K785, K787, R788, Q793, K821, R833, R836, K847, K879, K881, R883, R and / or K1208.
[0132] In certain embodiments, the enzyme may comprise, but is not limited to, the following amino acid positions relative to the amino acid position numbering of MbCpf1 (Moraxella bovoculi): K14, R17, R25, K33, M42, Q47, K50, D55, K85, N86, K88, K94, R104, K105, K118, K123, K131, R174, K175, R190, R198, I221, K267, Q228, I231, I232, I233, I234, I235, I236, I237, I238, I239, I240, I241, I252, I253, I254, I255, I266, I267, I270, I271, I272, I273, I274, I275, I276, I277, I278, I279, I280, I281, I282, I283, I284, I285, I286, I287, I288, I289, I290, I291, I292, I293, I300, I301, I302, I303, I314, I315, I316, I317, I318, I321, I322, I323, I334, I335, I346, I347, I350, I351, I352, I353, I365, I366, I377, I378, I381, I382, I383, I384, I385, I386, I38 69, K285, K291, K297, K357, K403, K409, K414, K448, K460, K501, K515, K550, R552 , K558, K564, K566, K582, K593, K604, K608, K623, K627, K633, K637, E643, K780, Y7 87, K792, K830, Q846, K858, K867, K876, K890, R900, K901, M906, K921, K927, K928 , K937, K939, R940, K945, Q975, R987, R990, K1001, R1034, I1036, R1038, R1042, K1 and / or K1356.
[0133] Non-activated / inactivated Cpf1 protein Where the Cpf1 protein has nuclease activity, the Cpf1 protein can be modified to have reduced nuclease activity, for example, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or 100% nuclease inactivation when compared to the wild-type enzyme; or stated differently, the Cpf1 enzyme advantageously has about 0% of the nuclease activity of a non-mutated or wild-type Cpf1 enzyme or CRISPR enzyme, or a non-mutated or wild-type Cpf1 enzyme, such as a non-mutated or wild-type Francisella novicida U112 (FnCpf1), Acidaminococcus sp. BV3L6 (AsCpf1), Lachnospiraceae The nuclease activity of Cpf1 enzymes or CRISPR enzymes is approximately 3%, 5%, or 10% of the nuclease activity of Cpf1 and its orthologs.
[0134] More specifically, inactivated Cpf1 enzymes include enzymes mutated at amino acid positions As908, As993, and As1263 of AsCpf1, or corresponding positions in Cpf1 orthologs. Additionally, inactivated Cpf1 enzymes include enzymes mutated at amino acid positions Lb832, 925, 947, and 1180 of LbCpf1, or corresponding positions in Cpf1 orthologs. More specifically, inactivated Cpf1 enzymes include enzymes containing one or more of the AsCpf1 mutations AsD908A, AsE993A, and AsD1263A, or corresponding mutations in Cpf1 orthologs. Additionally, inactivated Cpf1 enzymes include enzymes containing one or more of the LbCpf1 mutations LbD832A, E925A, D947A, and D1180A, or corresponding mutations in Cpf1 orthologs.
[0135] The inactivated Cpf1 CRISPR enzyme may include one or more domains from the group consisting of, consisting essentially of, or a molecular switch (e.g., light-inducible), such as methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription deactivation activity, histone modification activity, RNA cleavage activity, DNA cleavage activity, nucleic acid binding activity, and light-inducible activity, and may be associated with one or more functional domains (e.g., via a fusion protein). Preferred domains are Fok1, VP64, P65, HSF1, and MyoD1. When Fok1 is provided, it is advantageous to provide multiple Fok1 functional domains to achieve functional dimerization, and to design the gRNA to provide appropriate spacing for functional use (Fok1), as specifically described in Tsai et al. Nature Biotechnology, Vol. 32, Number 6, June 2014. The adaptor protein may link such functional domains using a known linker. In some cases, it is advantageous to further provide at least one NLS. In some cases, it is advantageous to place the NLS at the N-terminus. When two or more functional domains are included, the functional domains may be the same or different.
[0136] Generally, the location of one or more functional domains on the inactivated Cpf1 enzyme allows the functional domain to be in the correct spatial arrangement to affect the target with the ascribed functional effect. For example, if the functional domain is a transcriptional activator (e.g., VP64 or p65), the transcriptional activator is placed in a spatial arrangement that allows it to affect the transcription of the target. Similarly, a transcriptional repressor will be advantageously positioned to affect the transcription of the target, and a nuclease (e.g., Fok1) will be advantageously positioned to cleave or partially cleave the target. This may include locations other than the N-terminus / C-terminus of the CRISPR enzyme.
[0137] Destabilized Cpf1 In certain embodiments, the effector protein of the present invention (CRISPR enzyme; Cpf1) as described herein is associated with or fused to a destabilization domain (DD). In some embodiments, the DD is ER50. The corresponding stabilizing ligand of this DD is, in some embodiments, 4HT. Therefore, in some embodiments, one of at least one DD is ER50, and thus the stabilizing ligand is 4HT or CMP8. In some embodiments, the DD is DHFR50. The corresponding stabilizing ligand of this DD is, in some embodiments, TMP. Therefore, in some embodiments, one of at least one DD is DHFR50, and thus the stabilizing ligand is TMP. In some embodiments, the DD is ER50. The corresponding stabilizing ligand of this DD is, in some embodiments, CMP8. Therefore, CMP8 can be an alternative stabilizing ligand to 4HT in the ER50 system. It is possible that CMP8 and 4HT can / should be used competitively, but some cell types may be susceptible to either one of these two ligands, and based on this disclosure and knowledge in the art, one skilled in the art can use CMP8 and / or 4HT.
[0138] In some embodiments, one or two DDs may be fused to the N-terminal end of the CRISPR enzyme, and one or two DDs may be fused to the C-terminal end of the CRISPR enzyme. In some embodiments, at least two DDs are associated with the CRISPR enzyme, and the DDs are the same DD, i.e., the DDs are homologous. Thus, both (or two or more) of the DDs may be ER50 DDs. This is preferred in some embodiments. Alternatively, both (or two or more) of the DDs may be DHFR50 DDs. This is also preferred in some embodiments. In some embodiments, at least two DDs are associated with the CRISPR enzyme, and the DDs are different DDs, i.e., the DDs are heterologous. Thus, one of the DDs may be ER50, while one or more of the DDs or any other DD may be DHFR50. Having two or more heterologous DDs may be advantageous, as it may result in a higher level of degradation control. Tandem fusion of two or more DDs at the N- or C-terminus can enhance degradation; and such tandem fusions can be, for example, ER50-ER50-C2c2 or DHFR-DHFR-Cpf1. It is envisioned that high levels of degradation occur in the absence of either stabilizing ligand, intermediate levels of degradation can occur in the absence of one stabilizing ligand and the presence of the other (or another) stabilizing ligand, while low levels of degradation can occur in the presence of both (or more) stabilizing ligands. Control can also be provided by having an N-terminal ER50 DD and a C-terminal DHFR50 DD.
[0139] In some embodiments, the fusion of the CRISPR enzyme and the DD comprises a linker between the DD and the CRISPR enzyme. In some embodiments, the linker is a GlySer linker. In some embodiments, the DD-CRISPR enzyme further comprises at least one nuclear export signal (NES). In some embodiments, the DD-CRISPR enzyme comprises two or more NESs. In some embodiments, the DD-CRISPR enzyme comprises at least one nuclear localization signal (NLS), which may be in addition to an NES. In some embodiments, the CRISPR enzyme comprises, consists essentially of, or consists of a localization (nuclear import or nuclear export) signal as, or as part of, the linker between the CRISPR enzyme and the DD. HA or Flag tags are also within the scope of the present invention as linkers. Applicants use NLSs and / or NESs as linkers, as well as glycine-serine linkers as short as GS up to (GGGGS)3.
[0140] Destabilizing domains have general utility in conferring instability to a wide range of proteins; see, e.g., Miyazaki, J Am Chem Soc. Mar 7, 2012;134(9):3942-3945 (incorporated herein by reference). CMP8 or 4-hydroxytamoxifen can be destabilizing domains. More generally, temperature-sensitive mutants of mammalian DHFR (DHFRts), destabilizing residues due to the N-end rule, were found to be stable at permissive temperatures but unstable at 37°C. Addition of methotrexate, a high-affinity ligand for mammalian DHFR, to cells expressing DHFRts partially inhibited protein degradation. This was an important demonstration that small molecule ligands can stabilize proteins that are normally targeted for degradation in cells. Using a rapamycin derivative, a destabilizing mutant of the FRB domain of mTOR (FRB *) stabilized, restoring the function of the fused kinase GSK-3β. 6,7 This system demonstrated that ligand-dependent stability represents an attractive strategy for modulating the function of specific proteins in complex biological environments. A system for controlling protein activity may involve rapamycin-induced dimerization of FK506-binding protein with FKBP12, resulting in ubiquitin complementation, rendering the DD functional. Mutants of human FKBP12 or ecDHFR proteins can be engineered to be metabolically unstable in the absence of their high-affinity ligands, Shield-1 or trimethoprim (TMP), respectively. These mutants are some of the potential destabilization domains (DDs) useful in the practice of the present invention, and the instability of the DD as a fusion with a CRISPR enzyme leads to proteasomal degradation of the entire fusion protein. Shield-1 and TMP bind to and stabilize the DD in a dose-dependent manner. The estrogen receptor ligand-binding domain (ERLBD, residues 305–549 of ERS1) can also be engineered as a destabilization domain. Because the estrogen receptor signaling pathway is involved in various diseases, such as breast cancer, this pathway has been extensively studied, and numerous estrogen receptor agonists and antagonists have been developed. Therefore, compatible pairs of ERLBDs and drugs are known. There are ligands that bind to mutant ERLBDs but not to wild-type ERLBDs. By using one of these mutant domains, encoding three mutations (L384M, M421G, and G521R),12 it is possible to modulate the stability of ERLBD-derived DDs with ligands that do not disrupt the endogenous estrogen-sensitive network. An additional mutation (Y537S) can be introduced to further destabilize the ERLBD, making it a potential DD candidate. This quadruple mutant is an advantageous DD expansion. This mutant ERLBD can be fused to a CRISPR enzyme, and its stability can be modulated or disrupted using a ligand, thereby enabling the CRISPR enzyme to possess a DD.Another DD can be a 12 kDa (107 amino acids) tag based on a mutant FKBP protein stabilized by a Shield1 ligand; see, for example, Nature Methods 5, (2008).For example, the DD can be a modified FK506-binding protein 12 (FKBP12) that binds to and is reversibly stabilized by a synthetic, biologically inert small molecule, Shield-1; see, e.g., Banaszynski LA, Chen LC, Maynard-Smith LA, Ooi AG, Wandless TJ. "A rapid, reversible, and tunable method to regulate protein function in living cells using synthetic small molecules." Cell. 2006;126:995-1004; Banaszynski LA, Sellmyer MA, Contag CH, Wandless TJ, Thorne SH. "Chemical control of protein stability and function in living mice." Nat Med. 2008;14:1123-1127; Maynard-Smith LA, Chen See LC, Banaszynski LA, Ooi AG, Wandless TJ. "A directed approach for engineering conditional protein stability using biologically silent small molecules." The Journal of biological chemistry. 2007;282:24866-24872; and Rodriguez, Chem Biol. Mar 23, 2012;19(3):391-398, all of which are incorporated herein by reference, and may be used in the practice of the invention in selecting a DD to associate with a CRISPR enzyme.As can be seen, the knowledge in the art includes several DDs, and the DDs can be associated with, for example, fused to, a CRISPR enzyme, preferably with a linker, such that the DD can be stabilized in the presence of a ligand and destabilized in its absence, thereby destabilizing the CRISPR enzyme as a whole, or the DD can be stabilized in the absence of a ligand and destabilized in the presence of a ligand; the DD allows for the regulation or control of the CRISPR enzyme and thus the CRISPR-Cas complex or system—in other words, turning it on or off, thereby providing a means for regulating or controlling the system, for example, in vivo or in vitro. For example, expressing a protein of interest as a fusion with a DD tag destabilizes it in cells and is rapidly degraded, for example, by the proteasome. Thus, the absence of a stabilizing ligand leads to degradation of the D-associated Cas. Fusing a novel DD to a protein of interest confers its instability to the protein of interest, resulting in rapid degradation of the entire fusion protein. Peak Cas activity is sometimes beneficial for reducing off-target effects. Therefore, a short burst of high activity is preferred. The present invention can provide such peaks. In one sense, the system is inducible. In another sense, the system is inhibited in the absence of a stabilizing ligand and de-inhibited in the presence of a stabilizing ligand.
[0141] Enzyme mutations that reduce off-target effects In one aspect, the present invention provides a non-naturally occurring or engineered CRISPR enzyme as described herein, preferably a Class 2 CRISPR enzyme, preferably a Type V or Type VI CRISPR enzyme, such as, but not limited to, Cpfl as described elsewhere herein, that has one or more mutations that result in reduced off-target effects, i.e., an improved CRISPR enzyme that is used to generate modifications at a target locus but has reduced or eliminated off-target activity, such as when complexed with a guide RNA, as well as an improved CRISPR enzyme for increasing the activity of a CRISPR enzyme, such as when complexed with a guide RNA. It should be understood that a mutant enzyme as described herein below can be used in any of the methods of the present invention as described elsewhere herein. Any of the methods, products, compositions, and uses as described elsewhere herein are also applicable to mutant CRISPR enzymes as further detailed below. In aspects and embodiments as described herein, when Cpf1 is referred to or read as a CRISPR enzyme, it should be understood that reconstitution of a functional CRISPR-Cas system preferably does not require or depend on a tracr sequence and / or the direct repeat is 5' (upstream) of the guide (target or spacer) sequence.
[0142] The following detailed aspects and embodiments are provided as further guidance.
[0143] The inventors have unexpectedly determined that modifications can be made to CRISPR enzymes that confer reduced off-target activity compared to unmodified CRISPR enzymes and / or increased on-target activity compared to unmodified CRISPR enzymes.Accordingly, in certain aspects of the present invention, improved CRISPR enzymes are provided herein that may be useful in a wide range of gene modification applications.Also provided herein are CRISPR complexes, compositions and systems, as well as methods and uses, all of which include the modified CRISPR enzymes disclosed herein.
[0144] In the present disclosure, the term "Cas" can refer to "Cpf1" or a CRISPR enzyme. In the context of this aspect of the invention, the Cpf1 or CRISPR enzyme is mutated or modified "whereby the enzyme in the CRISPR complex has a reduced ability to modify one or more off-target loci compared to the unmodified enzyme" (or similar phrase); and, when read herein, the terms "Cpf1" or "Cas" or "CRISPR enzyme" etc. are intended to include the mutated or modified Cpf1 or Cas or CRISPR enzyme of the invention, i.e., "whereby the enzyme in the CRISPR complex has a reduced ability to modify one or more off-target loci compared to the unmodified enzyme" (or similar phrase).
[0145] In one embodiment, an engineered Cpfl protein, e.g., Cpfl, as defined herein is provided, which protein complexes with a nucleic acid molecule comprising RNA to form a CRISPR complex, wherein when in the CRISPR complex, the nucleic acid molecule targets one or more target polynucleotide loci, and wherein the protein comprises at least one modification compared to an unmodified Cpfl protein, and wherein a CRISPR complex comprising the modified protein has altered activity compared to a complex comprising an unmodified Cpfl protein. When referred to herein as a CRISPR "protein," the Cpfl protein is preferably a modified CRISPR enzyme (with increased or decreased (or no) enzymatic activity), for example, without limitation, Cpfl. It should be understood that the term "CRISPR protein" can be used synonymously with "CRISPR enzyme," regardless of whether the CRISPR protein has been altered, such as having increased or decreased (or no) enzymatic activity compared to a wild-type CRISPR protein.
[0146] In some embodiments, the altered activity of the engineered CRISPR protein includes an altered binding property for a nucleic acid molecule comprising RNA or a target polynucleotide locus, an altered binding kinetics for a nucleic acid molecule comprising RNA or a target polynucleotide locus, or an altered binding specificity for a nucleic acid molecule comprising RNA or a target polynucleotide locus compared to an off-target polynucleotide locus.
[0147] In some embodiments, the unmodified Cas has DNA cleavage activity, such as Cpfl. In some embodiments, the Cas directs cleavage of one or both strands at the location of the target sequence, such as within the target sequence and / or within the complement of the target sequence. In some embodiments, the Cas directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the vector encodes a Cas that is mutated relative to the corresponding wild-type enzyme such that the mutant Cas lacks the ability to cleave one or both strands of a target polynucleotide containing the target sequence. In some embodiments, a Cas is considered to be substantially devoid of any DNA cleavage activity if the DNA cleavage activity of the mutated enzyme is about 25%, 10%, 5%, 1%, 0.1%, 0.01%, or less of the DNA cleavage activity of the unmutated enzyme; an example would be when the DNA cleavage activity of the mutant is nonexistent or negligible compared to the unmutated form. Thus, a Cas may contain one or more mutations and be used as a general-purpose DNA-binding protein, with or without fusion to a functional domain. The mutations may be artificially introduced or may be gain-of-function or loss-of-function mutations. In one aspect of the present invention, a Cas enzyme may be fused to a protein, e.g., a TAG, and / or an inducible / controllable domain, such as a chemically inducible / controllable domain. The Cas of the present invention may be a chimeric Cas protein; for example, a Cas with enhanced function due to its chimeric nature. A chimeric Cas protein may be a novel Cas containing fragments from two or more naturally occurring Cass. These may include fusions of one or more N-terminal fragments of one Cas9 homolog with one or more C-terminal fragments of another Cas homolog. Cas can be delivered to cells in the form of mRNA. Cas expression may be under the control of an inducible promoter. Avoiding read-through to known mutations is an explicit objective of the present invention.Indeed, the phrase "which reduces the ability of the enzyme in the CRISPR complex to modify one or more off-target loci compared to the unmodified enzyme, and / or which increases the ability of the enzyme in the CRISPR complex to modify one or more target loci compared to the unmodified enzyme" (or similar phrase) is not intended to refer to mutations that only result in nickases or dead Cas or known Cas9 mutations. However, this is not to say that one or more modifications or one or more mutations of the present invention that "which reduces the ability of the enzyme in the CRISPR complex to modify one or more off-target loci compared to the unmodified enzyme, and / or which increases the ability of the enzyme in the CRISPR complex to modify one or more target loci compared to the unmodified enzyme" (or similar phrase) cannot be combined with a mutation that results in a nickase or dead enzyme. Such a dead enzyme can be an enhanced nucleic acid molecule binder. And such a nickase can be an enhanced nickase. For example, changing one or more neutral amino acids in and / or near the groove and / or other charged residues at other positions in the Cas in close proximity to the nucleic acid (e.g., DNA, cDNA, RNA, gRNA) to one or more positively charged amino acids "may thereby decrease the ability of the enzyme in the CRISPR complex to modify one or more off-target loci compared to the unmodified enzyme, and / or may thereby increase the ability of the enzyme in the CRISPR complex to modify one or more target loci compared to the unmodified enzyme," e.g., resulting in additional cleavage.This can be both enhanced on-target and off-target cleavage (super-cutting Cpf1), so using it with what is known in the art as a tru-guide or tru-sgRNA (see, e.g., Fu et al., "Improving CRISPR-Cas nuclease specificity using truncated guide RNAs," Nature Biotechnology 32, 279-284 (2014) doi:10.1038 / nbt.2808 Received 17 November 2013 Accepted 06 January 2014 Published online 26 January 2014 Corrected online 29 January 2014) enhances on-target activity without increased off-target cleavage, or creates a super-cutting nickase, or combines with mutations that render Cas dead for super-binders.
[0148] In certain embodiments, the altered activity of the engineered Cpf1 protein includes increased targeting efficiency or reduced off-target binding, hi certain embodiments, the altered activity of the engineered Cpf1 protein includes altered cleavage activity.
[0149] In certain embodiments, the altered activity includes an altered binding characteristic for a nucleic acid molecule comprising RNA or a target polynucleotide locus, an altered binding kinetics for a nucleic acid molecule comprising RNA or a target polynucleotide locus, or an altered binding specificity for a nucleic acid molecule comprising RNA or a target polynucleotide locus compared to an off-target polynucleotide locus.
[0150] In certain embodiments, the activity change includes increased targeting efficiency or decreased off-target binding. In certain embodiments, the activity change includes altered cleavage activity. In certain embodiments, the activity change includes increased cleavage activity for the target polynucleotide locus. In certain embodiments, the activity change includes decreased cleavage activity for the target polynucleotide locus. In certain embodiments, the activity change includes decreased cleavage activity for off-target polynucleotide loci. In certain embodiments, the activity change includes increased cleavage activity for off-target polynucleotide loci.
[0151] Thus, in certain embodiments, the specificity for the target polynucleotide locus is increased relative to the off-target polynucleotide loci, while in other embodiments, the specificity for the target polynucleotide locus is decreased relative to the off-target polynucleotide loci.
[0152] In some embodiments of the present invention, the altered activity of the engineered Cpf1 protein comprises an alteration in the helicase kinetics.
[0153] In some embodiments of the present invention, the engineered Cpf1 protein comprises a modification that alters the association of the protein with a nucleic acid molecule, including RNA, or a target polynucleotide locus strand, or an off-target polynucleotide locus strand. In some embodiments of the present invention, the engineered Cpf1 protein comprises a modification that alters the formation of a CRISPR complex.
[0154] In certain embodiments, the modified Cpf1 protein comprises a modification that alters targeting of a nucleic acid molecule to a polynucleotide locus. In certain embodiments, the modification comprises a mutation in a region of the protein that associates with a nucleic acid molecule. In certain embodiments, the modification comprises a mutation in a region of the protein that associates with a strand of a target polynucleotide locus. In certain embodiments, the modification comprises a mutation in a region of the protein that associates with a strand of an off-target polynucleotide locus. In certain embodiments, the modification or mutation comprises a decrease in positive charge in a region of the protein that associates with a nucleic acid molecule comprising RNA, or a strand of a target polynucleotide locus, or a strand of an off-target polynucleotide locus. In certain embodiments, the modification or mutation comprises a decrease in negative charge in a region of the protein that associates with a nucleic acid molecule comprising RNA, or a strand of a target polynucleotide locus, or a strand of an off-target polynucleotide locus. In certain embodiments, the modification or mutation comprises an increase in positive charge in a region of the protein that associates with a nucleic acid molecule comprising RNA, or a strand of a target polynucleotide locus, or a strand of an off-target polynucleotide locus. In certain embodiments, the modification or mutation comprises an increase in negative charge in a region of the protein that associates with a nucleic acid molecule comprising RNA, or a strand of a target polynucleotide locus, or a strand of an off-target polynucleotide locus. In certain embodiments, the modification or mutation increases steric hindrance between the protein and a nucleic acid molecule comprising RNA, or a strand of a target polynucleotide locus, or a strand of an off-target polynucleotide locus. In certain embodiments, the modification or mutation comprises a substitution of Lys, His, Arg, Glu, Asp, Ser, Gly, or Thr. In certain embodiments, the modification or mutation comprises a substitution with Gly, Ala, Ile, Glu, or Asp. In certain embodiments, the modification or mutation comprises an amino acid substitution within the binding groove.
[0155] In one aspect, the present invention provides a method for producing a pharmaceutical composition comprising: a non-naturally occurring CRISPR enzyme as defined herein, such as Cpf1 Provided here: This enzyme complexes with the guide RNA to form the CRISPR complex, When in the CRISPR complex, the guide RNA targets one or more target polynucleotide loci, the enzyme alters the polynucleotide loci, and The enzyme comprises at least one modification, This reduces the ability of the enzyme in the CRISPR complex to modify one or more off-target loci compared to the unmodified enzyme, and / or this increases the ability of the enzyme in the CRISPR complex to modify one or more target loci compared to the unmodified enzyme.
[0156] In any such non-naturally occurring CRISPR enzyme, the modification can include modifying one or more amino acid residues of the enzyme.
[0157] In any such non-naturally occurring CRISPR enzyme, the modification can include modifying one or more amino acid residues located in a region that in the unmodified enzyme includes a positively charged residue.
[0158] In any such non-naturally occurring CRISPR enzyme, the modification can include modifying one or more amino acid residues that are positively charged in the unmodified enzyme.
[0159] In any such non-naturally occurring CRISPR enzyme, the modification can include modifying one or more amino acid residues that are not positively charged in the unmodified enzyme.
[0160] The modification may include modifying one or more amino acid residues that are uncharged in the unmodified enzyme.
[0161] The modification may include modifying one or more amino acid residues that are negatively charged in the unmodified enzyme.
[0162] The modification may include altering one or more amino acid residues that are hydrophobic in the unmodified enzyme.
[0163] The modification may include altering one or more amino acid residues that are polar in the unmodified enzyme.
[0164] In certain of the non-naturally occurring CRISPR enzymes described above, the modification can include modification of one or more residues located in the groove.
[0165] In certain of the non-naturally occurring CRISPR enzymes described above, the modification can include modifying one or more residues located outside the groove.
[0166] In certain of the non-naturally occurring CRISPR enzymes described above, the modification includes modification of one or more residues, wherein the one or more residues include arginine, histidine, or lysine.
[0167] In any of the non-naturally occurring CRISPR enzymes described above, the enzyme can be modified by mutation of one or more of said residues.
[0168] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, where the mutation comprises substitution of a residue in the unmodified enzyme with an alanine residue.
[0169] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, where the mutation comprises substitution of the residue in the unmodified enzyme with aspartic acid or glutamic acid.
[0170] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, where the mutation comprises substitution of a residue in the unmodified enzyme with a serine, threonine, asparagine, or glutamine.
[0171] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, where the mutation comprises substitution of a residue in the unmodified enzyme with alanine, glycine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine, or valine.
[0172] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, wherein the mutation comprises substitution of a residue in the unmodified enzyme with a polar amino acid residue.
[0173] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, wherein the mutation comprises substitution of the residue in the unmodified enzyme with an amino acid residue that is not a polar amino acid residue.
[0174] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, wherein the mutation comprises substitution of a residue in the unmodified enzyme with a negatively charged amino acid residue.
[0175] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, wherein the mutation comprises substitution of the residue in the unmodified enzyme with an amino acid residue that is not a negatively charged amino acid residue.
[0176] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, wherein the mutation comprises substitution of the residue in the unmodified enzyme with an uncharged amino acid residue.
[0177] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, wherein the mutation comprises substitution of the residue in the unmodified enzyme with an amino acid residue that is not an uncharged amino acid residue.
[0178] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, wherein the mutation comprises substitution of a residue in the unmodified enzyme with a hydrophobic amino acid residue.
[0179] In certain of the above-described non-naturally occurring CRISPR enzymes, the enzyme is modified by mutation of said one or more residues, wherein the mutation comprises substitution of the residue in the unmodified enzyme with an amino acid residue that is not a hydrophobic amino acid residue.
[0180] In some embodiments, the CRISPR enzyme, preferably such as the Cpf1 enzyme, is selected from the group consisting of Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, or Porphyromonas macacae macacae) Cpf1 (e.g., Cpf1 of one of these organisms modified as described herein) and may include additional mutations or changes, or may be a chimeric Cpf1.
[0181] In certain embodiments, the Cpf1 protein comprises one or more nuclear localization signal (NLS) domains. In certain embodiments, the Cpf1 protein comprises at least two or more NLSs.
[0182] In certain embodiments, the Cpf1 protein comprises a chimeric CRISPR protein comprising a first fragment from a first CRISPR ortholog and a second fragment from a second CRISPR ortholog, and the first and second CRISPR orthologs are different.
[0183] In certain embodiments, the enzyme is modified by or comprises a modification, e.g., comprises, consists essentially of, or consists of a modification by mutation of any one of the residues listed herein or the corresponding residues in their respective orthologs; or the enzyme comprises, consists essentially of, or consists of a modification at any one (single), two (double), three (triple), four (quadruple) or more positions in this disclosure throughout this application, or at the corresponding residues or positions in a CRISPR enzyme ortholog, e.g., the enzyme comprises, consists essentially of, or consists of a modification at any one of the Cpfl residues listed herein, or at the corresponding residues or positions in a CRISPR enzyme ortholog. In such enzymes, each residue may be modified by substitution with an alanine residue.
[0184] Applicants recently described a method for generating Cas9 orthologs with enhanced specificity (Slaymaker et al. 2015, "Rationally engineered Cas9 nucleases with improved specificity"). This strategy can be used to enhance the specificity of Cpfl orthologs. The primary residues for mutagenesis are preferably all positively charged residues within the RuvC domain. The additional residues are positively charged residues that are conserved among different orthologs.
[0185] In certain embodiments, the specificity of Cpf1 may be improved by mutating residues that stabilize the non-target DNA strand.
[0186] In certain of the above-described non-naturally occurring Cpf1 enzymes, the enzyme may be, but is not limited to, AsCpf1 (Acidaminococcus sp. The RuvC domain is modified by mutation of one or more residues (within the RuvC domain), including positions R909, R912, R930, R947, K949, R951, R955, K965, K968, K1000, K1002, R1003, K1009, K1017, K1022, K1029, K1035, K1054, K1072, K1086, R1094, K1095, K1109, K1118, K1142, K1150, K1158, K1159, R1220, R1226, R1242, and / or R1252 relative to the amino acid position numbering of RuvC domain (BV3L6).
[0187] In certain of the above-described non-naturally occurring Cpf1 enzymes, the enzyme is modified by mutation of one or more residues (within the RAD50 domain), including, but not limited to, positions K324, K335, K337, R331, K369, K370, R386, R392, R393, K400, K404, K406, K408, K414, K429, K436, K438, K459, K460, K464, R670, K675, R681, K686, K689, R699, K705, R725, K729, K739, K748, and / or K752 relative to the amino acid position numbering of AsCpf1 (Acidaminococcus sp. BV3L6).
[0188] In certain of the above-described non-naturally occurring Cpf1 enzymes, the enzyme is modified by mutation of one or more residues, including, but not limited to, positions R912, T923, R947, K949, R951, R955, K965, K968, K1000, R1003, K1009, K1017, K1022, K1029, K1072, K1086, F1103, R1226, and / or R1252, based on the amino acid position numbering of AsCpf1 (Acidaminococcus sp. BV3L6).
[0189] In certain embodiments, the enzyme is modified by mutation of one or more residues, including but not limited to, positions R833, R836, K847, K879, K881, R883, R887, K897, K900, K932, R935, K940, K948, K953, K960, K984, K1003, K1017, R1033, R1138, R1165, and / or R1252 based on the amino acid position numbering of LbCpf1 (Lachnospiraceae bacterium ND2006).
[0190] In certain embodiments, the Cpf1 enzyme is selected from the group consisting of, but not limited to, positions K15, R18, K26, Q34, R43, K48, K51, R56, R84, K85, K87, N93, R103, N104, T118, K123, K134, R176, K177, R192, K200, K226, K273, K280, K290, K300, K310, K320, K330, K340, K350, K360, K370, K380, K390, K410, K420, K430, K440, K450, K460, K470, K480, K510, R56, R84, K85, K87, N93, R103, N104, T118, K123, K134, R176, K177, R192, K200, K226, K273, K280, K29 ...00, K420, K430, K440, K450, K460, K470, K480, K510, R56, R84, K85, K8 75, T291, R301, K307, K369, S404, V409, K414, K436, K438, K468, D482, K516, R51 8, K524, K530, K532, K548, K559, K570, R574, K592, D596, K603, K607, K613, C647, R681, K686, H720, K739, K748, K757, T766, K780, R790, P791, K796, K809, K815, T 816, K860, R862, R863, K868, K897, R909, R912, T923, R947, K949, R951, R955, K9 65, K968, K1000, R1003, K1009, K1017, K1022, K1029, A1053, K1072, K1086, F1103, S1209, R1226, R1252, K1273, K1282, and / or K1288.
[0191] In certain embodiments, the Cpf1 enzyme may be, but is not limited to, FnCpf1 (Francisella novicida) Based on the amino acid numbering of the genus U112), the following positions are used: K15, R18, K26, R34, R43, K48, K51, K56, K87, K88, D90, K96, K106, K107, K120, Q125, K143, R186, K187, R202, K210, K235, K296, K298, K314, K320, K326, K397, K444, K449, E454, A483, E491, K527, K541, K581, R583, K589, K595, K597, K613, K624, K635, K639, K656, K660, K671, K672, K673, K674, K675, K676, K677, K678, K679, K680, K681, K682, K683, K684, K685, K686, K687, K688, K689, K690, K691, K692, K693, K694, K695, K696, K697, K698, K699, K700, K701, K702, K703, K704, K705, K706, K707, K710, K711, K712, K713, K714, K715, K716, K717, K718, K720, K721, K722, K723, K724, K725, K730, K731, and / or K1098.
[0192] In certain embodiments, the Cpf1 enzyme may comprise, but is not limited to, the amino acid sequences at positions K15, R18, K26, K34, R43, K48, K51, R56, K83, K84, R86, K92, R102, K103, K116, K121, R158, R159, R174, R182, K206, K251, R264, R274, R282, R294, R306, R310, R320, R332, R340, R352, R360, R374, R382, R394, R406, R410, R420, R430, R440, R452, R460, R470, R480, R490, R510, R520, R530, R540, R550, R560, R570, R580, R590, R600, R610, R620, R630, R640, R650, R660, R670, R680, R690, R710, R720, R730, R740, R750, R760, R770, R780, R790, R800, R810, R820, R830, R840, R860, R870, R880, R890, R910, R920, R930, R940, R950, R960, R970, R980, R990, R990, R991, R992, R993, R994, R995, R996, R99 , K253, K269, K271, K278, P342, K380, R385, K390, K415, K421, K457, K471, A506 , R508, K514, K520, K522, K538, Y548, K560, K564, K580, K584, K591, K595, K601, K634, K640, R645, K679, K689, K707, T716, K725, R737, R747, R748, K753, K768, K774, K775, K785, K787, R788, Q793, K821, R833, R836, K847, K879, K881, R883, R and / or K1208.
[0193] In certain embodiments, the enzyme may comprise, but is not limited to, the following amino acid positions relative to the amino acid position numbering of MbCpf1 (Moraxella bovoculi): K14, R17, R25, K33, M42, Q47, K50, D55, K85, N86, K88, K94, R104, K105, K118, K123, K131, R174, K175, R190, R198, I221, K267, Q228, I231, I232, I233, I234, I235, I236, I237, I238, I239, I240, I241, I252, I253, I254, I255, I266, I267, I270, I271, I272, I273, I274, I275, I276, I277, I278, I279, I280, I281, I282, I283, I284, I285, I286, I287, I288, I289, I290, I291, I292, I293, I300, I301, I302, I303, I314, I315, I316, I317, I318, I321, I322, I323, I334, I335, I346, I347, I350, I351, I352, I353, I365, I366, I377, I378, I381, I382, I383, I384, I385, I386, I38 69, K285, K291, K297, K357, K403, K409, K414, K448, K460, K501, K515, K550, R552 , K558, K564, K566, K582, K593, K604, K608, K623, K627, K633, K637, E643, K780, Y7 87, K792, K830, Q846, K858, K867, K876, K890, R900, K901, M906, K921, K927, K928 , K937, K939, R940, K945, Q975, R987, R990, K1001, R1034, I1036, R1038, R1042, K1 and / or K1356.
[0194] In any of the non-naturally occurring CRISPR enzymes: There may be a single mismatch between the target and the corresponding sequence at one or more off-target loci; and / or There may be two, three, four or more mismatches between the target and the corresponding sequence at one or more off-target loci; and / or wherein in (ii), said 2, 3 or 4 or more mismatches are consecutive.
[0195] In any of the non-naturally occurring CRISPR enzymes, the enzyme in the CRISPR complex may have a decreased ability to modify one or more off-target loci compared to the unmodified enzyme, and the enzyme in the CRISPR complex has an increased ability to modify the target loci compared to the unmodified enzyme.
[0196] For any non-naturally occurring CRISPR enzyme, when in a CRISPR complex, the relative difference in the modification ability of the enzyme as between the target and at least one off-target locus can be increased compared to the relative difference of the unmodified enzyme.
[0197] In any of the non-naturally occurring CRISPR enzymes, the CRISPR enzyme may include one or more additional mutations, wherein the one or more additional mutations are in one or more catalytically active domains.
[0198] In such non-naturally occurring CRISPR enzymes, the CRISPR enzyme may have reduced or eliminated nuclease activity compared to an enzyme lacking the one or more additional mutations.
[0199] In some such non-naturally occurring CRISPR enzymes, the CRISPR enzyme does not direct cleavage of either DNA strand at that position in the target sequence.
[0200] Where the CRISPR enzyme comprises one or more additional mutations in one or more catalytically active domains, the one or more additional mutations may be in a catalytically active domain of the CRISPR enzyme comprising RuvCI, RuvCII, or RuvCIII.
[0201] Without being bound by theory, in certain embodiments of the present invention, the described methods and mutations enhance the conformational rearrangement of CRISPR enzyme domains (e.g., Cpf1 domains) to positions that result in cleavage at on-target sites and the avoidance of their conformational state at off-target sites. CRISPR enzymes cleave target DNA in a series of coordinated steps. First, the PAM-interacting domain recognizes the PAM sequence 5' of the target DNA. After PAM binding, the first 10-12 nucleotides of the target sequence (the seed sequence) are sampled for gRNA:DNA complementarity, a process that relies on DNA duplex separation. If the seed sequence nucleotides are complementary to the gRNA, the remaining DNA unwinds and the entire length of the gRNA hybridizes to the target DNA strand. The nt groove may stabilize the non-target DNA strand through nonspecific interactions with the positive charges of the DNA phosphate backbone, promoting unwinding. RNA:cDNA and Cas9:ncDNA interactions compete with cDNA:ncDNA rehybridization to drive DNA unwinding. Other CRISPR enzyme domains, such as linkers connecting different domains, can further affect the conformation of the nuclease domain.Therefore, the provided methods and mutations include, without limitation, RuvCI, RuvCIII, RuvCIII and linkers.The conformational changes of, for example, Cpf1 caused by target DNA binding, including seed sequence interaction and interaction with target and non-target DNA strands, determine whether the domain is in a position to initiate nuclease activity.Therefore, the mutations and methods provided herein demonstrate and enable modifications beyond PAM recognition and RNA-DNA base pairing.
[0202] In certain embodiments, the invention provides CRISPR nucleases as defined herein, such as Cpfl, that comprise an improved equilibrium toward a conformation associated with cleavage activity when engaged in on-target interactions and / or an improved equilibrium away from a conformation associated with cleavage activity when engaged in off-target interactions. In one embodiment, the invention provides a Cas (e.g., Cpfl) nuclease with improved proofreading function, i.e., a Cas (e.g., Cpfl) nuclease that adopts a conformation that includes nuclease activity at on-target sites and whose conformation becomes more unfavorable at off-target sites. Sternberg et al., Nature 527(7576):110-3, doi:10.1038 / nature15544, published online 28 October 2015. Epub 2015 Oct 28 used Förster resonance energy transfer (FRET) experiments to detect the relative orientation of Cas (e.g., Cpf1) catalytic domains when associated with on-target and off-target DNA, which can be applied to the CRISPR enzymes (e.g., Cpf1) of the present invention.
[0203] The present invention further provides methods and mutations for modulating nuclease activity and / or specificity using modified guide RNAs. As discussed, on-target nuclease activity can be increased or decreased. Off-target nuclease activity can also be increased or decreased. Furthermore, specificity can be increased or decreased with respect to on-target activity versus off-target activity. Modified guide RNAs include, without limitation, truncated guide RNAs, dead guide RNAs, chemically modified guide RNAs, guide RNAs associated with functional domains, modified guide RNAs containing functional domains, modified guide RNAs containing aptamers, modified guide RNAs containing adaptor proteins, and guide RNAs containing additional or modified loops. In some embodiments, one or more functional domains are associated with a dead gRNA (dRNA). In some embodiments, a dRNA complex with a CRISPR enzyme directs gene regulation by the functional domain at one locus, while the gRNA directs DNA cleavage by the CRISPR enzyme at another locus. In some embodiments, the dRNA is selected to maximize selectivity of regulation for the locus of interest relative to off-target regulation. In some embodiments, dRNAs are selected to maximize target gene regulation and minimize target cleavage.
[0204] For purposes of the following discussion, reference to a functional domain can refer to a functional domain associated with a CRISPR enzyme or a functional domain associated with an adaptor protein.
[0205] In practicing the present invention, the loop of a gRNA may be extended without conflicting with a Cas (e.g., Cpf1) protein by inserting a discrete RNA loop or sequences that can recruit an adaptor protein capable of binding to the discrete RNA loop or sequences. Adaptor proteins may include, but are not limited to, orthogonal RNA-binding protein / aptamer combinations within the diversity of bacteriophage coat proteins. The list of such coat proteins includes, but is not limited to, Qβ, F2, GA, fr, JP501, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s, and PRR1. These adaptor proteins or orthogonal RNA-binding proteins can further recruit effector proteins or fusions containing one or more functional domains. In some embodiments, the functional domain may be selected from the group consisting of a transposase domain, an integrase domain, a recombinase domain, a resolvase domain, an invertase domain, a protease domain, a DNA methyltransferase domain, a DNA hydroxylmethylase domain, a DNA demethylase domain, a histone acetylase domain, a histone deacetylase domain, a nuclease domain, a repressor domain, an activator domain, a transcriptional regulatory protein (or transcription complex recruiting) domain, a domain associated with cellular uptake activity, a nucleic acid binding domain, an antibody presentation domain, a histone modifying enzyme, a recruiter of a histone modifying enzyme; an inhibitor of a histone modifying enzyme, a histone methyltransferase, a histone demethylase, a histone kinase, a histone phosphatase, a histone ribosylase, a histone derivosylase, a histone ubiquitinase, a histone deubiquitinase, a histone biotinase, and a histone tail protease.In some preferred embodiments, the functional domain is a transcriptional activation domain, such as, without limitation, VP64, p65, MyoD1, HSF1, RTA, SET7 / 9, or histone acetyltransferase. In some embodiments, the functional domain is a transcriptional repression domain, preferably KRAB. In some embodiments, the transcriptional repression domain is SID or a concatemer of SIDs (e.g., SID4X). In some embodiments, the functional domain is an epigenetically modified domain, providing an epigenetically modified enzyme. In some embodiments, the functional domain is an activation domain, which may be a P65 activation domain. In some embodiments, the functional domain is a deaminase, such as cytidine deaminase. Cytidine deaminase targets a target nucleic acid where it induces the conversion of cytidine to uridine, resulting in a C to T substitution (G to A on the complementary strand). In such embodiments, nucleotide substitution can be achieved without DNA cleavage.
[0206] In certain aspects, the present invention also provides methods and mutations that modulate Cas (e.g., Cpfl) binding activity and / or binding specificity. In certain embodiments, a Cas (e.g., Cpfl) protein that lacks nuclease activity is used. In certain embodiments, modified guide RNAs that promote Cas (e.g., Cpfl) nuclease binding but not nuclease activity are utilized. In such embodiments, on-target binding can be increased or decreased. In such embodiments, off-target binding can also be increased or decreased. Furthermore, there can be increased or decreased specificity with respect to on-target versus off-target binding.
[0207] In particular embodiments, reduced off-target cleavage is ensured by introducing mutations into the Cpfl enzyme that destabilize strand separation, more particularly that reduce the positive charge in the DNA interacting region (as described herein and further exemplified for Cas9 in Slaymaker et al. 2016 (Science, 1;351(6268):84-8). In further embodiments, reduced off-target cleavage is ensured by introducing mutations into the Cpfl enzyme that affect the interaction between the target strand and the guide RNA sequence, more particularly that disrupt the interaction between Cpfl and the phosphate backbone of the target DNA strand in such a way that the target specific activity is retained but the off-target activity is reduced (as described for Cas9 by Kleinstiver et al. 2016, Nature, 28;529(7587):490-5). In particular embodiments, off-target activity is reduced by modified Cpfl, which has modified interactions with both the target and non-target strands compared to wild-type Cpfl.
[0208] Methods and mutations, which can be used in various combinations to increase or decrease the activity and / or specificity of on-target versus off-target activity, or to increase or decrease the binding and / or specificity of on-target versus off-target binding, can be used to compensate for or enhance mutations or modifications made to promote other effects. Such mutations or modifications made to promote other effects include mutations or modifications to Cas (e.g., Cpf1) and / or mutations or modifications made to the guide RNA. In certain embodiments, the methods and mutations are used with chemically modified guide RNAs. Examples of chemical modifications of guide RNAs include, without limitation, incorporation of 2'-O-methyl (M), 2'-O-methyl 3' phosphorothioate (MS), or 2'-O-methyl 3' thioPACE (MSP) at one or more terminal nucleotides. Such chemically modified guide RNAs can include increased stability and activity compared to unmodified guide RNAs, however, on-target versus off-target specificity is unpredictable (see Hendel, 2015, Nat Biotechnol. 33(9):985-9, doi:10.1038 / nbt.3290, published online 29 June 2015). Chemically modified guide RNAs further include, without limitation, RNAs with phosphorothioate linkages and locked nucleic acid (LNA) nucleotides containing a methylene bridge between the 2' and 4' carbons of the ribose ring. The methods and mutations of the present invention are used to modulate Cas (e.g., Cpf1) nuclease activity and / or binding by chemically modified guide RNAs.
[0209] In certain aspects, the invention provides methods and mutations for altering the binding and / or binding specificity of the Cas (e.g., Cpfl) proteins of the invention as defined herein, including functional domains such as nuclease, transcriptional activator, transcriptional repressor, etc. For example, Cas (e.g., Cpfl) proteins can be made nuclease null or have altered or reduced nuclease activity by introducing mutations, such as the Cpfl mutations described elsewhere herein, e.g., FnCpflp Nuclease-deficient Cas (e.g., Cpfl) proteins are useful for RNA-guided, target-sequence-dependent delivery of functional domains. The present invention provides methods and mutations that modulate binding of Cas (e.g., Cpfl) proteins. In one embodiment, the functional domain comprises VP64, which provides an RNA-guided transcription factor. In another embodiment, the functional domain comprises Fok I, which provides an RNA-guided nuclease activity. U.S. Patent Application Publication Nos. 2014 / 0356959, 2014 / 0342456, 2015 / 0031132, and Mali, P. et al., 2013, Science 339(6121):823-6, doi:10.1126 / science.1232033, published online 3 January 2013, are cited, and throughout the teachings herein, the present invention encompasses the methods and materials of these documents applied in conjunction with the teachings herein. In certain embodiments, on-target binding is increased. In certain embodiments, off-target binding is decreased. In certain embodiments, on-target binding is decreased.In certain embodiments, off-target binding is increased. Thus, the present invention also provides for increasing or decreasing the specificity of on-target versus off-target binding of a functional Cas (e.g., Cpfl) binding protein.
[0210] The use of Cas (e.g., Cpf1) as an RNA-guided binding protein is not limited to nuclease-null Cas (e.g., Cpf1). Cas (e.g., Cpf1) enzymes containing nuclease activity can also function as RNA-guided binding proteins when used with specific guide RNAs. For example, small guide RNAs and guide RNAs containing mismatched nucleotides can promote RNA-guided Cas9 binding to target sequences with little or no target cleavage (see, e.g., Dahlman, 2015, Nat Biotechnol. 33(11):1159-1161, doi:10.1038 / nbt.3390, published online 05 October 2015). In certain aspects, the present invention provides methods and mutations that modulate the binding of Cas (e.g., Cpf1) proteins containing nuclease activity. In certain embodiments, on-target binding is increased. In certain embodiments, off-target binding is decreased. In certain embodiments, on-target binding is decreased. In certain embodiments, off-target binding is increased. In certain embodiments, there is an increase or decrease in the specificity of on-target versus off-target binding. In certain embodiments, the nuclease activity of a guide RNA-Cas (e.g., Cpfl) enzyme is also modulated.
[0211] RNA-DNA heteroduplex formation is important for cleavage activity and specificity throughout the target region, not just the seed region sequence closest to the PAM. Thus, truncated guide RNAs exhibit reduced cleavage activity and specificity. In some embodiments, the present invention provides methods and mutations that increase cleavage activity and specificity using altered guide RNAs.
[0212] The present invention also demonstrates that modifications of Cas (e.g., Cpfl) nuclease specificity can be tailored to target domain modifications. For example, by selecting mutations that alter PAM specificity and combining them with nt groove mutations that increase (or, if desired, decrease) specificity for on-target versus off-target sequences, Cas (e.g., Cpfl) mutants can be designed that have increased target specificity as well as tailored modifications of PAM recognition. In one such embodiment, PI domain residues are mutated to tailor recognition of a desired PAM sequence while one or more nt groove amino acids are mutated to alter target specificity. The Cas (e.g., Cpfl) methods and modifications described herein can be used to counteract loss of specificity caused by altered PAM recognition, enhance gain of specificity caused by altered PAM recognition, counteract gain of specificity caused by altered PAM recognition, or enhance loss of specificity caused by altered PAM recognition.
[0213] The methods and mutations can be used with any Cas (e.g., Cpfl) enzyme with altered PAM recognition. Non-limiting examples of PAMs include those described anywhere herein.
[0214] In a further embodiment, the methods and mutations are used in engineered proteins.
[0215] In any of the non-naturally occurring CRISPR enzymes, the CRISPR enzyme may comprise one or more heterologous functional domains.
[0216] The one or more heterologous functional domains may include one or more nuclear localization signal (NLS) domains. The one or more heterologous functional domains may include at least two or more NLSs.
[0217] The one or more heterologous functional domains include one or more transcription activation domains. The transcription activation domain can include VP64.
[0218] The one or more heterologous functional domains include one or more transcriptional repression domains, which may include a KRAB domain or a SID domain.
[0219] The one or more heterologous functional domains can include one or more nuclease domains. The one or more nuclease domains can include Fok1.
[0220] The one or more heterologous functional domains can have one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription deactivator activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity.
[0221] The at least one heterologous functional domain can be at or near the amino terminus of the enzyme and / or at or near the carboxy terminus of the enzyme.
[0222] The one or more heterologous functional domains may be fused to the CRISPR enzyme, tethered to the CRISPR enzyme, or linked to the CRISPR enzyme by a linker moiety.
[0223] Among any of the non-naturally occurring CRISPR enzymes, the CRISPR enzyme is selected from the group consisting of Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, or Porphyromonas macacae The Cas may comprise a CRISPR enzyme from an organism of the genus CRISPR1, including CRISPR1 (Cpf1) from one of these organisms modified as described herein, and may comprise additional mutations or alterations, or may be a chimeric Cas (e.g., Cpf1).
[0224] In any of the non-naturally occurring CRISPR enzymes, the CRISPR enzyme can comprise a chimeric Cas (e.g., Cpfl) enzyme comprising a first fragment from a first Cas (e.g., Cpfl) ortholog and a second fragment from a second Cas (e.g., Cpfl) ortholog, wherein the first and second Cas (e.g., Cpfl) orthologs are different. At least one of the first and second Cas (e.g., Cpf1) orthologs is / are present in Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus The Cas may include a Cas (e.g., Cpf1) derived from an organism such as Pseudomonas sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis3, Prevotella disiens, or Porphyromonas macacae.
[0225] In any of the non-naturally occurring CRISPR enzymes, the nucleotide sequence encoding the CRISPR enzyme may be codon-optimized for expression in eukaryotes.
[0226] In any of the non-naturally occurring CRISPR enzymes, the cell may be a eukaryotic cell or a prokaryotic cell; wherein a CRISPR complex is operable in the cell such that the enzymes of the CRISPR complex have a reduced ability to modify one or more off-target loci in the cell compared to the unmodified enzymes, and / or such that the enzymes in the CRISPR complex have an increased ability to modify one or more target loci compared to the unmodified enzymes.
[0227] Thus, in one aspect, the present invention provides a eukaryotic cell comprising an engineered CRISPR protein or system as defined herein.
[0228] In certain embodiments, the methods as described herein may include providing a Cas (e.g., Cpf1) transgenic cell in which one or more nucleic acids encoding one or more guide RNAs operably linked within the cell to a regulatory element comprising a promoter of one or more genes of interest are provided or introduced. As used herein, the term "Cas transgenic cell" refers to a cell, such as a eukaryotic cell, in which a Cas gene has been genomically integrated. The nature, type, or origin of the cell is not particularly limited according to the present invention. Additionally, the method of introducing a Cas transgene into a cell may vary and may be any method known in the art. In certain embodiments, the Cas transgenic cell is obtained by introducing a Cas transgene into an isolated cell. In certain other embodiments, the Cas transgenic cell is obtained by isolating cells from a Cas transgenic organism. By way of example and without limitation, the Cas transgenic cell as referred to herein may be derived from a Cas transgenic eukaryotic organism, such as a Cas knock-in eukaryotic organism. Reference is made to International Publication No. WO 2014 / 093622 (PCT / US13 / 74667), which is incorporated herein by reference. The methods of U.S. Patent Application Publication Nos. 20120017290 and 20110265198, assigned to Sangamo BioSciences, Inc., relating to targeting the Rosa locus, can be modified to utilize the CRISPR Cas system of the present invention. The methods of U.S. Patent Application Publication No. 20130236946, assigned to Cellectis, relating to targeting the Rosa locus can also be modified to utilize the CRISPR Cas system of the present invention. As a further example, reference is made to Platt et al. (Cell;159(2):440-455(2014)), which describes Cas9 knock-in mice, which can be applied to the CRISPR enzymes of the present invention as defined herein.The Cas transgene may further comprise a Lox-Stop-PolyA-Lox (LSL) cassette, which allows Cas expression to be inducible by Cre recombinase. Alternatively, Cas transgenic cells may be obtained by introducing the Cas transgene into isolated cells. Transgene delivery systems are well known in the art. For example, the Cas transgene may be delivered using vectors (e.g., AAV, adenovirus, lentivirus) and / or particle and / or nanoparticle delivery, for example, in eukaryotic cells, as also described elsewhere herein.
[0229] Those skilled in the art will understand that cells, such as Cas transgenic cells as referenced herein, may contain additional genomic alterations in addition to having an integrated Cas gene, or may contain mutations, such as one or more oncogenic mutations, that arise due to the sequence-specific action of Cas when complexed with an RNA capable of guiding Cas to a target locus, as described, for example and without limitation, in Platt et al. (2014), Chen et al., (2014), or Kumar et al. (2009).
[0230] The present invention also provides compositions comprising an engineered CRISPR protein as described herein, such as described in this section.
[0231] The present invention also provides non-naturally occurring engineered compositions comprising a CRISPR-Cas complex comprising any of the non-naturally occurring CRISPR enzymes described above.
[0232] In one embodiment, the present invention provides a vector system comprising one or more vectors, wherein the one or more vectors a) a first regulatory element operably linked to a nucleotide sequence encoding an engineered CRISPR protein as defined herein; and optionally, b) a second regulatory element operably linked to one or more nucleotide sequences encoding one or more nucleic acid molecules comprising a guide sequence, a guide RNA comprising a direct repeat sequence, optionally wherein components (a) and (b) are located on the same or different vectors.
[0233] The present invention also provides A delivery system operably configured to deliver CRISPR-Cas complex components or one or more polynucleotide sequences comprising or encoding said components to a cell, wherein said CRISPR-Cas complex is operably present in the cell. Also provided is a naturally occurring engineered composition comprising: The CRISPR-Cas complex component or one or more polynucleotide sequences encoding a CRISPR-Cas complex component for transcription and / or translation in a cell may be (I) a non-naturally occurring CRISPR enzyme described herein (e.g., engineered Cpf1); (II) a CRISPR-Cas guide RNA comprising: Guide sequence, Direct repeat sequences, Contains, where: The enzymes in the CRISPR complex have a reduced ability to modify one or more off-target loci compared to the unmodified enzymes, and / or thereby have an increased ability to modify one or more target loci compared to the unmodified enzymes.
[0234] In some embodiments, the present invention also provides a system comprising an engineered CRISPR protein as described herein, such as described in this section.
[0235] In any such composition, the delivery system may include a yeast system, a lipofection system, a microinjection system, a biolistic system, a virosome, a liposome, an immunoliposome, a polycation, a lipid:nucleic acid conjugate, or an artificial virion, as described anywhere herein.
[0236] In any such composition, the delivery system may comprise a vector system comprising one or more vectors, wherein component (II) comprises a first regulatory element operably linked to a polynucleotide sequence optionally comprising a guide sequence and a direct repeat sequence, and component (I) comprises a second regulatory element operably linked to a polynucleotide sequence encoding a CRISPR enzyme.
[0237] In any such composition, the delivery system may include a vector system comprising one or more vectors, wherein component (II) comprises a first regulatory element operably linked to a guide sequence and a direct repeat sequence, and component (I) comprises a second regulatory element operably linked to a polynucleotide sequence encoding a CRISPR enzyme.
[0238] In any such composition, the composition can include two or more guide RNAs, each guide RNA having a different target, thereby providing multiplexing.
[0239] In any such composition, one or more polynucleotide sequences may be on a single vector.
[0240] The present invention also provides a) a first regulatory element operably linked to a nucleotide sequence encoding the non-naturally occurring CRISPR enzyme of any one of the constructs of the invention herein; and b) a second regulatory element operably linked to one or more nucleotide sequences encoding one or more guide RNAs, wherein the guide RNAs comprise a guide sequence and a direct repeat sequence; Also provided is an engineered non-naturally occurring clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) vector system comprising one or more vectors comprising: Components (a) and (b) are located on the same or different vectors; CRISPR complexes form; The guide RNA targets the target polynucleotide locus, the enzyme alters the polynucleotide locus, and The enzymes in the CRISPR complex have a reduced ability to modify one or more off-target loci compared to the unmodified enzymes, and / or thereby have an increased ability to modify one or more target loci compared to the unmodified enzymes.
[0241] In such a system, component (II) may comprise a first regulatory element operably linked to a polynucleotide sequence comprising a guide sequence and a direct repeat sequence, and component (II) may comprise a second regulatory element operably linked to a polynucleotide sequence encoding a CRISPR enzyme. In such a system, where applicable, the guide RNA may comprise a chimeric RNA.
[0242] In such a system, component (I) can comprise a first regulatory element operably linked to a guide sequence and a direct repeat sequence, and component (II) can comprise a second regulatory element operably linked to a polynucleotide sequence encoding a CRISPR enzyme. Such a system can comprise two or more guide RNAs, each with a different target, thereby providing multiplexing. Components (a) and (b) can be on the same vector.
[0243] In any such system comprising a vector, the one or more vectors may include one or more viral vectors, such as one or more retroviruses, lentiviruses, adenoviruses, adeno-associated viruses, or herpes simplex viruses.
[0244] In any such system that includes regulatory elements, at least one of the regulatory elements can include a tissue-specific promoter that can direct expression in mammalian blood cells, mammalian liver cells, or mammalian eyes.
[0245] In any of the above compositions or systems, the direct repeat sequence can include one or more protein-interacting RNA aptamers. The one or more aptamers can be located in the tetraloop. The one or more aptamers can be capable of binding to MS2 bacteriophage coat protein.
[0246] In any of the above compositions or systems, the cell may be a eukaryotic or prokaryotic cell; wherein a CRISPR complex is operable in the cell such that the enzymes of the CRISPR complex have a reduced ability to modify one or more off-target loci in the cell compared to the unmodified enzymes, and / or such that the enzymes in the CRISPR complex have an increased ability to modify one or more target loci in the cell compared to the unmodified enzymes.
[0247] The present invention also provides a CRISPR complex of any of the compositions described above or from any of the systems described above.
[0248] The present invention also provides a method for modifying a locus of interest in a cell, the method comprising contacting the cell with any of the engineered CRISPR enzymes (e.g., engineered Cpfl), compositions described herein, or any of the systems or vector systems described herein, or wherein the cell comprises any of the CRISPR complexes described herein present within the cell. In such methods, the cell may be a prokaryotic or eukaryotic cell, preferably a eukaryotic cell. In such methods, an organism may comprise the cell. In such methods, the organism need not be a human or other animal.
[0249] Any such method may be ex vivo or in vitro.
[0250] In certain embodiments, the nucleotide sequence encoding at least one of the guide RNA or Cas protein is operably linked in the cell to regulatory elements including a promoter of a gene of interest, such that expression of at least one CRISPR-Cas system component is driven by the promoter of the gene of interest. "Operably linked," as mentioned elsewhere herein, is intended to mean that the nucleotide sequence encoding the guide RNA and / or Cas is linked to one or more regulatory elements in a manner that allows expression of the nucleotide sequence. The term "regulatory element" is also described elsewhere herein. According to the present invention, the regulatory element preferably comprises a promoter of a gene of interest, such as a promoter of an endogenous gene of interest. In certain embodiments, the promoter is in its endogenous genomic location. In such embodiments, the nucleic acid encoding the CRISPR and / or Cas is under the transcriptional control of the promoter of the gene of interest in its natural genomic location. In certain other embodiments, the promoter is provided on a (separate) nucleic acid molecule, such as a vector or plasmid, or on another extrachromosomal nucleic acid, i.e., the promoter is not provided in its natural genomic location. In certain embodiments, the promoter is genomically integrated at a non-native genomic location.
[0251] In any such method, the modification can include modifying gene expression. The modification of gene expression can include activating gene expression and / or repressing gene expression. Thus, in one embodiment, the present invention provides a method of modulating gene expression, the method comprising introducing into a cell an engineered CRISPR protein or system as described herein.
[0252] The present invention also provides methods of treating a disease, disorder, or infection in an individual in need thereof, the methods comprising administering an effective amount of any of the engineered CRISPR enzymes (e.g., engineered Cpfl), compositions, systems, or CRISPR complexes described herein. The disease, disorder, or infection can include a viral infection. The viral infection can be HBV.
[0253] The present invention also provides the use of any of the engineered CRISPR enzymes (e.g., engineered Cpf1), compositions, systems or CRISPR complexes described above for gene or genome editing.
[0254] The present invention also provides a method of altering expression of a genomic locus of interest in a mammalian cell, the method comprising contacting the cell with an engineered CRISPR enzyme (e.g., an engineered Cpfl), composition, system, or CRISPR complex described herein, thereby delivering a CRISPR-Cas (vector) and allowing the CRISPR-Cas complex to form and bind to the target, and determining whether expression of the genomic locus has been altered, such as by increased or decreased expression, or an altered gene product.
[0255] The present invention also provides any of the engineered CRISPR enzymes (e.g., engineered Cpfl), compositions, systems, or CRISPR complexes described above for use as a therapeutic agent, which may be for gene or genome editing, or gene therapy.
[0256] In certain embodiments, the activity of an engineered CRISPR enzyme (e.g., an engineered Cpfl) as described herein includes genomic DNA cleavage, optionally resulting in decreased transcription of the gene.
[0257] In one aspect, the invention provides isolated cells in which expression of a genomic locus has been altered by a method described herein, wherein the altered expression is relative to a cell that has not been subjected to the method to alter expression of the genomic locus. In a related aspect, the invention provides cell lines established from such cells.
[0258] In one aspect, the invention provides a method of modifying an organism or non-human organism by manipulation of a target sequence at a genomic locus of interest, e.g., HSC (hematopoietic stem cell) (e.g., the genomic locus of interest is associated with a mutation associated with aberrant protein expression or a disease condition or state), the method comprising: I. CRISPR-Cas system guide RNA (gRNA) polynucleotide sequences, (a) a guide sequence capable of hybridizing to a target sequence in an HSC; (b) direct repeat sequences; a polynucleotide sequence comprising: II. A CRISPR enzyme, optionally including at least one nuclear localization sequence; delivering to the HSCs, e.g., by contacting the HSCs with particles containing the non-naturally occurring or engineered composition comprising wherein the guide sequence directs sequence-specific binding of the CRISPR complex to the target sequence; and wherein the CRISPR complex comprises: (1) a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence; and The method can also optionally include delivering an HDR template, e.g., via particles that contact the HSCs containing the HDR template or by contacting the HSCs with another particle containing the HDR template, where the HDR template results in expression of a normal or less aberrant form of the protein; "normal" is with respect to wild-type, and "aberrant" can be protein expression that causes a pathology or disease state; and Optionally, the method may include isolating or obtaining HSCs from an organism or non-human organism, optionally expanding the HSC population, contacting the HSCs with one or more particles to obtain a modified HSC population, optionally expanding the population of modified HSCs, and optionally administering the modified HSCs to the organism or non-human organism.
[0259] In one aspect, the invention provides a method of modifying an organism or non-human organism by manipulation of a target sequence at a genomic locus of interest in an HSC (e.g., the genomic locus of interest is associated with a mutation associated with aberrant protein expression or a disease condition or state), the method comprising: delivering to the HSC, e.g., by contacting the HSC with particles containing the same, a non-naturally occurring or engineered composition comprising: I. (a) a guide sequence hybridizable to the target sequence in the HSC, and (b) at least one or more direct repeat sequences, and II. a CRISPR enzyme, optionally having one or more NLSs, wherein the guide sequence directs sequence-specific binding of the CRISPR complex to the target sequence, and wherein the CRISPR complex comprises the CRISPR enzyme complexed with the guide sequence that hybridizes to the target sequence; and The method can also optionally include delivering an HDR template, e.g., via particles that contact the HSCs containing the HDR template or by contacting the HSCs with another particle containing the HDR template, where the HDR template results in expression of a normal or less aberrant form of the protein; "normal" is with respect to wild-type, and "aberrant" can be protein expression that causes a pathology or disease state; and Optionally, the method may include isolating or obtaining HSCs from an organism or non-human organism, optionally expanding the HSC population, contacting the HSCs with one or more particles to obtain a modified HSC population, optionally expanding the population of modified HSCs, and optionally administering the modified HSCs to the organism or non-human organism.
[0260] This delivery can be, for example, delivery of one or more polynucleotides encoding any one or more or all of the CRISPR complexes, advantageously linked to one or more regulatory elements for in vivo expression, via one or more particles containing a vector containing one or more polynucleotides operably linked to one or more regulatory elements. Some or all of the polynucleotide sequences encoding the CRISPR enzyme, guide sequence, and direct repeat sequence can be RNA. When a polynucleotide is referred to as being RNA and "comprising" such a direct repeat sequence feature, it will be understood that the RNA sequence comprises the feature. When a polynucleotide is DNA and is referred to as comprising such a direct repeat sequence feature, the DNA sequence is or can be transcribed into RNA containing the feature in question. When the feature is a protein such as a CRISPR enzyme, the DNA or RNA sequence referred to is or can be translated (first transcribed, in the case of DNA).
[0261] In certain embodiments, the invention provides methods of modifying an organism, e.g., a mammal, including a human, or a non-human mammal or organism, by manipulation of a target sequence in a genomic locus of interest in HSCs, e.g., associated with a mutation associated with aberrant protein expression or a disease condition or state, comprising delivering, e.g., by contacting, a non-naturally occurring or engineered composition with HSCs, wherein the composition comprises one or more particles comprising one or more viruses, plasmids, or nucleic acid molecule vectors (e.g., RNA) operably encoding the composition for expression, the composition comprising: (A) a first regulatory element operably linked to a CRISPR-Cas system RNA polynucleotide sequence, wherein the polynucleotide sequence comprises: (a) a guide sequence capable of hybridizing to a target sequence in a eukaryotic cell; (b) a direct repeat sequence; and II. a second regulatory element operably linked to an enzyme-coding sequence encoding a CRISPR enzyme comprising at least one or more nuclear localization sequences (or optionally at least one or more nuclear localization sequences, as some embodiments may not involve an NLS) [(a), (b), and (c) are arranged in a 5' to 3' direction, components I and II are located on the same or different vectors of the system, when transcribed, and the guide sequence directs sequence-specific binding of a CRISPR complex to a target sequence, and the CRISPR complex comprises the CRISPR enzyme complexed with the guide sequence that hybridizes to the target sequence], or (B) a non-naturally occurring or engineered composition, comprising: I. a first regulatory element operably linked to (a) a guide sequence capable of hybridizing to a target sequence in a eukaryotic cell, and (b) at least one or more direct repeat sequences; II.the composition comprising a vector system comprising one or more vectors comprising a second regulatory element operably linked to an enzyme-coding sequence encoding a CRISPR enzyme, wherein components I and II are located on the same or different vectors of the system, and when transcribed, the guide sequence directs sequence-specific binding of a CRISPR complex to the target sequence, and the CRISPR complex comprises the CRISPR enzyme complexed with the guide sequence that hybridizes to the target sequence; the method optionally comprising, for example, via particles contacting HSCs containing the HDR template or by contacting the HSCs with separate particles containing the HDR template; The method can also include delivering an HDR template, wherein the HDR template results in expression of a normal or less aberrant form of the protein; "normal" refers to wild-type, and "aberrant" can be protein expression that causes a pathology or disease state; and optionally, the method can include isolating or obtaining HSCs from an organism or non-human organism, optionally expanding the HSC population, contacting the HSCs with one or more particles to obtain a modified HSC population, optionally expanding the population of modified HSCs, and optionally administering the modified HSCs to the organism or non-human organism. In some embodiments, components I, II, and III are located on the same vector. In other embodiments, components I and II are located on the same vector, while component III is located on a separate vector. In other embodiments, components I and III are located on the same vector, while component II is located on a separate vector. In other embodiments, components II and III are located on the same vector, while component I is located on a separate vector. In other embodiments, each of components I, II, and III is located on a different vector. The present invention also provides viral or plasmid vector systems as described herein. .
[0262] By manipulation of a target sequence, Applicants also mean epigenetic manipulation of the target sequence. This may be manipulation of the chromatin state of the target sequence, such as by modifying the methylation status of the target sequence (i.e., methylation or methylation pattern or addition or removal of CpG islands), histone modification, increasing or decreasing the accessibility of the target sequence, or promoting three-dimensional folding. When referring to a method of modifying an organism or mammal, including a human, or a non-human mammal or organism by manipulating a target sequence at a genomic locus of interest, it will be understood that this may apply to the organism (or mammal) as a whole, or (if the organism is a multicellular organism) to just a single cell or cell population of that organism. For example, in the case of humans, Applicants specifically contemplate single cells or cell populations, which may be modified ex vivo and then reintroduced. In this case, a biopsy or other tissue or biological fluid sample may be required. Stem cells are also particularly preferred in this regard. However, of course, in vivo embodiments are also contemplated. The present invention is particularly advantageous with respect to HSCs.
[0263] In some embodiments, the present invention provides, for example, I. A first CRISPR-Cas (e.g., Cpf1) system RNA (RNA) polynucleotide sequence, (a) a first guide sequence capable of hybridizing to a first target sequence; (b) a first direct repeat sequence, and a first polynucleotide sequence comprising: II. A second CRISPR-Cas (e.g., Cpf1) system guide RNA polynucleotide sequence, (a) a second guide sequence capable of hybridizing to a second target sequence; (b) a second direct repeat sequence, and a second polynucleotide sequence comprising: III. A polynucleotide sequence encoding a CRISPR enzyme comprising at least one or more nuclear localization sequences and comprising one or more mutations ((a), (b), and (c) arranged in a 5' to 3' direction); or IV. One or more expression products of one or more of I.-III., e.g., first and second direct repeat sequences, a CRISPR enzyme; a method for modifying an organism or non-human organism, the method comprising: delivering, by contacting the HSCs with a particle comprising a non-naturally occurring or engineered composition comprising: when transcribed, first and second guide sequences direct sequence-specific binding of first and second CRISPR complexes to first and second target sequences, respectively; the first CRISPR complex comprises a CRISPR enzyme complexed with the first guide sequence that hybridizes to the first target sequence; and the second CRISPR complex comprises a CRISPR enzyme complexed with the second guide sequence that hybridizes to the second target sequence; the polynucleotide sequence encoding the CRISPR enzyme is DNA or RNA; and the first guide sequence directs cleavage of one strand of a DNA duplex near the first target sequence and the second guide sequence directs cleavage of the other strand near the second target sequence to create a double-stranded break, thereby modifying the organism or non-human organism, the method comprising: delivering, by contacting the HSCs with a particle comprising a non-naturally occurring or engineered composition comprising: when transcribed, first and second guide sequences direct sequence-specific binding of first and second CRISPR complexes to the first and second target sequences, respectively; the first CRISPR complex comprises a CRISPR enzyme complexed with the first guide sequence that hybridizes to the first target sequence; and the second CRISPR complex comprises a CRISPR enzyme complexed with the second guide sequence that hybridizes to the second target sequence, the polynucleotide sequence encoding the CRISPR enzyme is DNA or RNA; and the first guide sequence directs cleavage of one strand of a DNA duplex near the first target sequence and the second guide sequence directs cleavage of the other strand near the second target sequence to create a double-stranded break, thereby modifying the organism or non-human organism, the method comprising: and the method optionally can also include delivering an HDR template, for example via particles that contact the HSCs containing the HDR template or by contacting the HSCs with other particles that contain the HDR template, wherein the HDR template results in expression of a normal or less aberrant form of the protein; "normal" is with respect to wild-type, and "aberrant" can be protein expression that causes a pathology or disease state; and optionally the method includes isolating or obtaining HSCs from the organism or non-human organism, optionally expanding the HSC population, contacting the HSCs with one or more particles to obtain a modified HSC population, optionally expanding the population of modified HSCs, and optionally administering the modified HSCs to the organism or non-human organism.In some methods of the present invention, some or all of the polynucleotide sequence encoding the CRISPR enzyme, the first and second guide sequences, and the first and second direct repeat sequences are RNA. In further embodiments of the present invention, the polynucleotides encoding the CRISPR enzyme, the first and second guide sequences, and the first and second direct repeat sequences are RNA and are delivered by liposomes, nanoparticles, exosomes, microvesicles, or gene guns; however, delivery by particles is advantageous. In certain embodiments of the present invention, the first and second direct repeats share 100% identity. In some embodiments, the polynucleotides may be contained within a vector system comprising one or more vectors. In a preferred embodiment, the first CRISPR enzyme has one or more mutations that make it a complementary strand nicking enzyme, and the second CRISPR enzyme has one or more mutations that make it a non-complementary strand nicking enzyme. Alternatively, the first enzyme may be a non-complementary strand nicking enzyme, and the second enzyme may be a complementary strand nicking enzyme. In a preferred method of the present invention, a first guide sequence induces cleavage of one strand of a DNA duplex near a first target sequence, and a second guide sequence induces cleavage of the other strand near a second target sequence, resulting in a 5' overhang. In an embodiment of the present invention, the 5' overhang is at most 200 base pairs, preferably at most 100 base pairs, or more preferably at most 50 base pairs. In an embodiment of the present invention, the 5' overhang is at least 26 base pairs, preferably at least 30 base pairs, or more preferably 34-50 base pairs.
[0264] In some embodiments, the present invention provides, for example, I. A first regulatory element, (a) a first guide sequence capable of hybridizing to a first target sequence, and (b) at least one direct repeat sequence a first regulatory element operably linked to II. A second regulatory element, (a) a second guide sequence capable of hybridizing to a second target sequence, and (b) at least one direct repeat sequence a second regulatory element operably linked to III. A third regulatory element operably linked to an enzyme coding sequence encoding a CRISPR enzyme (e.g., Cpf1), and One or more expression products of one or more of VI to IV., e.g., first and second direct repeat sequences, CRISPR enzyme; [Once transcribed, components I, II, III, and IV are located on the same or different vectors of the system, and the first and second guide sequences direct sequence-specific binding of first and second CRISPR complexes to first and second target sequences, respectively, wherein the first CRISPR complex comprises (1) a CRISPR enzyme complexed with the first guide sequence that hybridizes to the first target sequence, and the second CRISPR complex comprises (1) a CRISPR enzyme complexed with the second guide sequence that hybridizes to the second target sequence. wherein the polynucleotide sequence encoding the CRISPR enzyme is DNA or RNA, and the first guide sequence directs cleavage of one strand of the DNA duplex near a first target sequence and the second guide sequence directs cleavage of the other strand near a second target sequence to create a double-stranded break, thereby modifying the organism or non-human organism, e.g., a genomic locus of interest in the HSCs. and optionally, the method can also include the step of delivering an HDR template, for example, via particles contacting the HSCs containing the HDR template or by contacting the HSCs with another particle containing the HDR template, wherein the HDR template results in expression of a normal or less aberrant form of the protein; "normal" is with respect to wild-type, and "aberrant" can be protein expression that causes a pathology or disease state; and optionally, the method can include the steps of isolating or obtaining HSCs from the organism or non-human organism, optionally expanding the HSC population, contacting the HSCs with one or more particles to obtain a modified HSC population, optionally expanding the population of modified HSCs, and optionally administering the modified HSCs to the organism or non-human organism.
[0265] The present invention also provides a vector system as described herein. The system may include one, two, three, or four different vectors. Thus, components I, II, III, and IV may be located on one, two, three, or four different vectors, and all possible combinations of component locations are contemplated herein, for example: components I, II, III, and IV may be located on the same vector in all possible combinations of locations; components I, II, III, and IV may each be located on a different vector; components I, II, III, and IV may be located on a total of two or three different vectors, etc. In some methods of the present invention, some or all of the polynucleotide sequence encoding the CRISPR enzyme, the first and second guide sequences, and the first and second direct repeat sequences are RNA. In a further embodiment of the present invention, the first and second direct repeat sequences share 100% identity. In a preferred embodiment, the first CRISPR enzyme has one or more mutations that make it a complementary strand nicking enzyme, and the second CRISPR enzyme has one or more mutations that make it a non-complementary strand nicking enzyme. Alternatively, the first enzyme can be a non-complementary strand nicking enzyme, and the second enzyme can be a complementary strand nicking enzyme. In a further embodiment of the present invention, one or more viral vectors are delivered by liposomes, nanoparticles, exosomes, microvesicles, or gene guns; however, particle delivery is preferred.
[0266] In a preferred method of the present invention, a first guide sequence induces cleavage of one strand of a DNA duplex near a first target sequence, and a second guide sequence induces cleavage of the other strand near a second target sequence, resulting in a 5' overhang. In an embodiment of the present invention, the 5' overhang is at most 200 base pairs, preferably at most 100 base pairs, or more preferably at most 50 base pairs. In an embodiment of the present invention, the 5' overhang is at least 26 base pairs, preferably at least 30 base pairs, or more preferably 34-50 base pairs.
[0267] The invention, in some embodiments, encompasses a method of modifying a genomic locus of interest in, e.g., an HSC, e.g., associated with a mutation associated with aberrant protein expression or a disease condition or state, by introducing into the HSC, e.g., by contacting the HSC with particles comprising a Cas protein having one or more mutations and two guide RNAs that target a first strand and a second strand, respectively, of a DNA molecule in the HSC, whereby the guide RNAs target the DNA molecule and the Cas protein nicks each of the first strand and the second strand of the DNA molecule, thereby altering the target in the HSC (and where the Cas protein and the two guide RNAs do not naturally occur together), and this method optionally includes, e.g., The method can also include delivering the HDR template, for example, via particles that contact the HSCs containing the HDR template or by contacting the HSCs with other particles containing the HDR template, where the HDR template results in normal or mildly aberrant protein expression; "normal" refers to wild-type, and "aberrant" can be protein expression that causes a pathology or disease state; and optionally, the method can include isolating or obtaining HSCs from an organism or non-human organism, optionally expanding the HSC population, contacting the HSCs with one or more particles to obtain a modified HSC population, optionally expanding the population of modified HSCs, and optionally administering the modified HSCs to the organism or non-human organism. In a preferred method of the present invention, a Cas protein nicks each of the first and second strands of the DNA molecule, resulting in a 5' overhang. In an embodiment of the present invention, the 5' overhang is at most 200 base pairs, preferably at most 100 base pairs, or more preferably at most 50 base pairs. In embodiments of the invention, the 5' overhang is at least 26 base pairs, preferably at least 30 base pairs, or more preferably 34-50 base pairs. In aspects of the invention, the Cas protein is codon-optimized for expression in eukaryotic cells, preferably mammalian or human cells.Aspects of the invention relate to reducing expression of a gene product, or introducing a template polynucleotide into a DNA molecule encoding the gene product, or precisely excising an intervening sequence by allowing reannealing and ligation of two 5' overhangs, or altering the activity or function of a gene product, or increasing expression of a gene product. In certain embodiments of the invention, the gene product is a protein.
[0268] In some embodiments, the present invention provides, for example, a) a first regulatory element operably linked to each of two CRISPR-Cas system guide RNAs that target a first strand and a second strand, respectively, of a double-stranded DNA molecule of an HSC; and b) a second regulatory element operably linked to a Cas (e.g., Cpf1) protein; or c) one or more expression products of a) or b); wherein components (a) and (b) are located on the same or different vectors of the system, by contacting the HSCs with one or more particles comprising the HDR template, whereby the guide RNA targets the DNA molecule of the HSCs and the Cas protein nicks each of the first and second strands of the DNA molecule of the HSCs (and wherein the Cas protein and the two guide RNAs do not naturally occur together), thereby modifying a genomic locus of interest in the HSCs, e.g., a genomic locus of interest associated with a mutation associated with aberrant protein expression or a disease condition or state; and the method optionally includes, for example, via particles contacting the HSCs containing the HDR template, or The method may also include delivering the HDR template by contacting the HSCs with another particle containing the HDR template, wherein the HDR template results in normal or less aberrant protein expression; "normal" refers to wild-type, and "aberrant" can be protein expression that causes a pathology or disease state; and optionally, the method may include isolating or obtaining HSCs from an organism or non-human organism, optionally expanding the HSC population, contacting the HSCs with one or more particles to obtain a modified HSC population, optionally expanding the population of modified HSCs, and optionally administering the modified HSCs to the organism or non-human organism. In an embodiment of the present invention, the guide RNA may comprise a guide sequence and a tracr sequence fused to a direct repeat sequence. Aspects of the invention relate to reducing expression of a gene product, or introducing a template polynucleotide into a DNA molecule encoding the gene product, or precisely excising an intervening sequence by allowing reannealing and ligation of two 5' overhangs, or altering the activity or function of a gene product, or increasing expression of a gene product. In some embodiments of the invention, the gene product is a protein. In preferred embodiments of the invention, the vector of the system is a viral vector. In further embodiments, the vector of the system is delivered by liposomes, nanoparticles, exosomes, microvesicles, or gene guns; and particles are preferred.In one aspect, the present invention provides a method for modifying a target polynucleotide in HSCs. In some embodiments, the method includes a step of allowing a CRISPR complex to bind to a target polynucleotide, causing cleavage of the target polynucleotide and thereby modifying the target polynucleotide, wherein the CRISPR complex includes a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence within the target polynucleotide, the guide sequence being linked to a direct repeat sequence. In some embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene containing the target sequence. In some embodiments, the method further includes a step of delivering one or more vectors or one or more expression products thereof, e.g., by one or more particles, to the HSCs, e.g., wherein the one or more vectors drive expression of one or more of the CRISPR enzyme, the guide sequence linked to the direct repeat sequence, and / or the target polynucleotide. In some embodiments, the vectors are delivered to, e.g., the subject's HSCs. In some embodiments, the modification occurs in the HSCs in cell culture. In some embodiments, the method further includes a step of isolating the HSCs from the subject prior to the modification. In some embodiments, the method further comprises returning the HSCs and / or cells derived therefrom to the subject.
[0269] In one aspect, the present invention provides methods for generating HSCs containing, for example, a mutant disease gene. In some embodiments, the disease gene is any gene associated with an increased risk of developing or having a disease. In some embodiments, the method includes (a) introducing one or more vectors or one or more expression products thereof into HSCs, e.g., via one or more particles, where the one or more vectors drive expression of one or more of a CRISPR enzyme, a guide sequence linked to a direct repeat sequence, and (b) allowing a CRISPR complex to bind to a target polynucleotide, resulting in cleavage of the target polynucleotide within the disease gene, where the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence within the target polynucleotide, and optionally, where applicable, thereby generating HSCs containing a mutant disease gene. In some embodiments, the cleavage comprises cleavage of one or both strands at the location of the target sequence by the CRISPR enzyme. In some embodiments, the cleavage results in decreased transcription of the target gene. In some embodiments, the method further comprises repairing the cleaved target polynucleotide by homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides in the target polynucleotide. In some embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence. In some embodiments, the modified HSCs are administered to an animal, thereby generating an animal model.
[0270] In one aspect, the present invention provides a method for modifying a target polynucleotide, for example, in HSCs. In some embodiments, the method comprises a step of allowing a CRISPR complex to bind to a target polynucleotide, causing cleavage of the target polynucleotide, thereby modifying the target polynucleotide, wherein the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence in the target polynucleotide, wherein the guide sequence is linked to a direct repeat sequence. In another embodiment, the present invention provides a method for modifying the expression of a polynucleotide, for example, in a eukaryotic cell derived from HSCs. The method comprises a step of increasing or decreasing the expression of the target polynucleotide using a CRISPR complex that binds to a polynucleotide in the HSC; advantageously, the CRISPR complex is delivered via one or more particles.
[0271] In some methods, the target polynucleotide can be inactivated to result in altered expression, for example, in HSCs. For example, when a CRISPR complex binds to a target sequence in a cell, the target polynucleotide is inactivated so that the sequence is not transcribed, the encoded protein is not produced, or the sequence does not function like the wild-type sequence.
[0272] In some embodiments, the RNA of the CRISPR-Cas system, such as the guide or gRNA, can be modified; for example, it can include an aptamer or a functional domain. Aptamers are synthetic oligonucleotides that bind to specific target molecules; nucleic acid molecules that have been engineered by repeated rounds of in vitro selection or SELEX (enrichment by evolution of sequences) to bind to various molecular targets, such as small molecules, proteins, nucleic acids, and even cells, tissues, and organisms. Aptamers are useful in that they offer molecular recognition properties that compete with antibodies. In addition to their discriminatory recognition, aptamers offer advantages over antibodies, including the fact that they induce little or no immunogenicity in therapeutic applications. Thus, in the practice of the present invention, either or both of the enzyme or RNA can include a functional domain.
[0273] In some embodiments, the functional domain is a transcription activation domain, preferably VP64. In some embodiments, the functional domain is a transcription repression domain, preferably KRAB. In some embodiments, the transcription repression domain is SID or a concatemer of SIDs (e.g., SID4X). In some embodiments, the functional domain is an epigenetic modification domain, thus providing an epigenetic modification enzyme. In some embodiments, the functional domain is an activation domain, which may be a P65 activation domain. In some embodiments, the functional domain comprises nuclease activity. In one such embodiment, the functional domain comprises Fok1.
[0274] The present invention also provides an in vitro or ex vivo cell comprising any of the modified CRISPR enzymes, compositions, systems, or complexes described above, or by any of the methods described above. The cell may be a eukaryotic or prokaryotic cell. The present invention also provides progeny of such cells. The present invention also provides a product of any such cell or any such progeny, wherein the product is a product of the one or more target loci as modified by the modified CRISPR enzyme of a CRISPR complex. The product may be a peptide, polypeptide, or protein. Some such products may be modified by the modified CRISPR enzyme of a CRISPR complex. In some such modified products, the product of the target locus is physically different from the product of the target locus that is not modified by the modified CRISPR enzyme.
[0275] The present invention also provides a polynucleotide molecule comprising a polynucleotide sequence encoding any of the non-naturally occurring CRISPR enzymes described above.
[0276] Any such polynucleotide may further comprise one or more regulatory elements operably linked to the polynucleotide sequence encoding the non-naturally occurring CRISPR enzyme.
[0277] In any such polynucleotide comprising one or more regulatory elements, the one or more regulatory elements can be configured to be operable for the expression of a non-naturally occurring CRISPR enzyme in a eukaryotic cell.The eukaryotic cell can be a human cell.The eukaryotic cell can be a rodent cell, optionally a mouse cell.The eukaryotic cell can be a yeast cell.The eukaryotic cell can be a Chinese hamster ovary (CHO) cell.The eukaryotic cell can be an insect cell.
[0278] In any such polynucleotide that includes one or more regulatory elements, the one or more regulatory elements can be operably configured for expression of a non-naturally occurring CRISPR enzyme in a prokaryotic cell.
[0279] In any such polynucleotide comprising one or more regulatory elements, the one or more regulatory elements can be operably configured for expression of a non-naturally occurring CRISPR enzyme in an in vitro system.
[0280] The present invention also provides expression vectors comprising any of the polynucleotide molecules described above. The present invention also provides one or more such polynucleotide molecules, e.g., such polynucleotide molecules operably configured to express one or more protein and / or nucleic acid components, as well as one or more such vectors.
[0281] The invention further provides a method of making a mutation in a Cas (e.g., Cpfl) or a mutated or modified Cas (e.g., Cpfl) that is orthologous of a CRISPR enzyme according to the invention described herein, comprising the steps of identifying one or more amino acids in the ortholog, which may be adjacent to or in contact with a nucleic acid molecule, e.g., DNA, RNA, gRNA, etc., for modification and / or mutation, and / or one or more amino acids similar or corresponding to one or more amino acids identified herein in a CRISPR enzyme according to the invention described herein, and synthesizing, or preparing, or expressing the ortholog comprising, consisting of, or consisting essentially of the one or more modifications and / or one or more mutations, or mutating, e.g., modifying, e.g., changing or mutating, as discussed herein, a neutral amino acid to a charged, e.g., positively charged amino acid, e.g., alanine to, e.g., lysine. Such modified orthologs can be used in CRISPR-Cas systems; and one or more nucleic acid molecules that express them can be used in vectors or other delivery systems that deliver molecules or encode CRISPR-Cas system components as discussed herein.
[0282] In some embodiments, the present invention provides efficient on-target activity and minimizes off-target activity. In some embodiments, the present invention provides efficient on-target cleavage by CRISPR proteins and minimizes off-target cleavage by CRISPR proteins. In some embodiments, the present invention provides guide-specific binding of CRISPR proteins at a locus without DNA cleavage. In some embodiments, the present invention provides efficient on-target binding, guided by a CRISPR protein at a locus, and minimizes off-target binding of the CRISPR protein. Thus, in some embodiments, the present invention provides target-specific gene regulation. In some embodiments, the present invention provides guide-specific binding of CRISPR enzymes at a locus without DNA cleavage. Thus, in some embodiments, the present invention provides cleavage at one locus and gene regulation at another locus using a single CRISPR enzyme. In some embodiments, the present invention provides orthogonal activation and / or inhibition and / or cleavage of multiple targets using one or more CRISPR proteins and / or enzymes.
[0283] In another aspect, the present invention provides a method for functional screening of genes in a genome in a pool of cells ex vivo or in vivo, the method comprising administering or expressing a library comprising a plurality of CRISPR-Cas system guide RNAs (gRNAs), wherein the screening further comprises the use of a CRISPR enzyme, and wherein the CRISPR complex is modified to comprise a heterologous functional domain. In some aspects, the present invention provides a method for screening a genome comprising administering the library to a host or expressing it in vivo in a host. In some aspects, the present invention provides a method as discussed herein, further comprising an activator administered to or expressed in the host. In some aspects, the present invention provides a method as discussed herein, wherein the activator is added to a CRISPR protein. In some aspects, the present invention provides a method as discussed herein, wherein the activator is added to the N-terminus or C-terminus of the CRISPR protein. In some aspects, the present invention provides a method as discussed herein, wherein the activator is added to the gRNA loop. In some embodiments, the invention provides a method as discussed herein, further comprising a repressor administered to or expressed in the host, hi some embodiments, the invention provides a method as discussed herein, wherein screening comprises affecting and detecting gene activation, gene inhibition, or truncation of the gene locus.
[0284] In some embodiments, the present invention provides a method as discussed herein, wherein the host is a eukaryotic cell. In some embodiments, the present invention provides a method as discussed herein, wherein the host is a mammalian cell. In some embodiments, the present invention provides a method as discussed herein, wherein the host is a non-human eukaryotic cell. In some embodiments, the present invention provides a method as discussed herein, wherein the non-human eukaryotic cell is a non-human mammalian cell. In some embodiments, the present invention provides a method as discussed herein, wherein the non-human mammalian cell may include, but is not limited to, a primate, bovine, ovine, porcine, canine, rodent, lagomorph, such as a monkey, cow, sheep, pig, dog, rabbit, rat, or mouse cell. In some embodiments, the present invention provides a method as discussed herein, wherein the cell may be a non-mammalian eukaryotic cell, such as a poultry avian (e.g., chicken), vertebrate fish (e.g., salmon), or crustacean (e.g., oyster, clam, lobster, shrimp) cell. In an aspect, the invention provides a method as discussed herein, wherein the non-human eukaryotic cell is a plant cell. The plant cell may be a monocotyledonous or dicotyledonous plant or a crop or cereal plant, such as cassava, maize, sorghum, soybean, wheat, oat, or rice. The plant cell may also be an algae, a tree or a productive plant, fruit or vegetable (e.g., a citrus tree, e.g., an orange, grapefruit or lemon tree; a peach or nectarine tree; an apple or pear tree; a nut tree such as an almond, walnut or pistachio tree; a Solanaceae plant; a Brassica plant; a Lactuca plant; a Spinacia plant; a Capsicum plant; a cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.).
[0285] In some embodiments, the present invention provides a method as discussed herein, comprising delivery of a CRISPR-Cas complex, or one or more components thereof, or one or more nucleic acid molecules encoding same, wherein said one or more nucleic acid molecules are operably linked to one or more regulatory sequences and expressed in vivo. In some embodiments, the present invention provides a method as discussed herein, wherein in vivo expression is by lentivirus, adenovirus, or AAV. In some embodiments, the present invention provides a method as discussed herein, wherein delivery is by particle, nanoparticle, lipid, or cell-penetrating peptide (CPP).
[0286] In specific embodiments, it may be beneficial to target the CRISPR-Cas complex to the chloroplast. This targeting can often be achieved by the presence of an N-terminal extension called a chloroplast transit peptide (CTP) or plastid transit peptide. If an expressed polypeptide is to be compartmentalized in a plant plastid (e.g., a chloroplast), a chromosomal transgene from a bacterial source must have a sequence encoding the CTP sequence fused to the sequence encoding the expressed polypeptide. Therefore, localization of an exogenous polypeptide to the chloroplast is often achieved by operably linking a polynucleotide sequence encoding the CTP sequence to the 5' region of a polynucleotide encoding the exogenous polypeptide. The CTP is removed in a processing step during translocation to the plastid. However, processing efficiency can be affected by the amino acid sequence of the CTP and the nearby sequence at the NH2-terminus of the peptide. Other options described for targeting to chloroplasts are the maize cab-m7 signal sequence (U.S. Pat. No. 7,022,896; WO 97 / 41228), the pea glutathione reductase signal sequence (WO 97 / 41228), and the CTP described in U.S. Patent Application Publication No. 2009029861.
[0287] In some embodiments, the present invention provides a pair of CRISPR-Cas complexes, each comprising a guide RNA (gRNA) that includes a guide sequence capable of hybridizing to a target sequence in a genomic locus of interest in a cell, wherein at least one loop of each gRNA is modified by the insertion of one or more distinct RNA sequences that bind to one or more adaptor proteins, and the adaptor proteins are associated with one or more functional domains, and each sgRNA of each CRISPR-Cas comprises a functional domain with DNA cleavage activity. In some embodiments, the present invention provides a pair of CRISPR-Cas complexes as discussed herein, wherein the DNA cleavage activity is due to Fok1 nuclease.
[0288] In some embodiments, the present invention provides a method for cleaving a target sequence at a genomic locus of interest, comprising delivering ...
Claims
1. 1. An engineered, non-naturally occurring clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system comprising: a) one or more V-type CRISPR-Cas polynucleotide sequences comprising a guide RNA comprising a guide sequence linked to a direct repeat sequence, the guide sequence being hybridizable to a target sequence, or one or more nucleotide sequences encoding the one or more V-type CRISPR-Cas polynucleotide sequences; and b) a Cpf1 effector protein or one or more nucleotide sequences encoding said Cpf1 effector protein; wherein the one or more guide sequences hybridize to the target sequence, the target sequence is 3' to a protospacer adjacent motif (PAM), and the guide RNA forms a complex with the Cpf1 effector protein.
2. 1. An engineered, non-naturally occurring clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) vector system comprising: a) a first regulatory element operably linked to one or more nucleotide sequences encoding one or more V-type CRISPR-Cas polynucleotide sequences comprising a guide RNA comprising a guide sequence linked to a direct repeat sequence, the guide sequence being hybridizable to a target sequence; b) a second regulatory element operably linked to the nucleotide sequence encoding a Cpf1 effector protein; and one or more vectors comprising components (a) and (b) are located on the same or different vectors of the system, Upon transcription, the one or more guide sequences hybridize to the target sequence, the target sequence being 3' to a protospacer adjacent motif (PAM), and the guide RNA forms a complex with the Cpf1 effector protein.
3. The system of claim 1 or 2, wherein the target sequence is intracellular.
4. The system of claim 3 , wherein the cells comprise eukaryotic cells.
5. 3. The system of claim 1 or 2, wherein, upon transcription, the one or more guide sequences hybridize to the target sequence and the guide RNA forms a complex with the Cpf1 effector protein, which causes cleavage distal to the target sequence.
6. The system of claim 5 , wherein the cleavage generates a sticky-end type double-stranded break with a 4 or 5 nt 5′ overhang.
7. The system of claim 1 or 2, wherein the PAM comprises a 5' T-rich motif.
8. 3. The system of claim 1 or 2, wherein the effector protein is a Cpf1 effector protein derived from a bacterial species listed in Figure 64.
9. The Cpf1 effector protein is a virulent inhibitor of Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parkobacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadae 9. The system of claim 8, wherein the bacterial species is selected from the group consisting of Porphyromonas inada, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens and Porphyromonas macacae.
10. The system of claim 9, wherein the PAM sequence is TTN, where N is A / C / G or T, and the effector protein is FnCpf1, or the PAM sequence is TTTV, where V is A / C or G, and the effector protein is PaCpf1p, LbCpf1, or AsCpf1.
11. The system of claim 1 or 2, wherein the Cpf1 effector protein comprises one or more nuclear localization signals.
12. 3. The system of claim 1 or 2, wherein the nucleic acid sequence encoding the Cpf1 effector protein is codon-optimized for expression in eukaryotic cells.
13. 3. The system according to claim 1 or 2, wherein components (a) and (b) or the nucleotide sequence are on one vector.
14. 10. A method for modifying a target locus of interest, comprising the step of delivering the system described in claim 1 or 2 to said locus or to a cell containing said locus.
15. 1. A method for modifying a target locus of interest, the method comprising the step of delivering to the locus a non-naturally occurring or engineered composition comprising a Cpf1 effector protein and one or more nucleic acid components, wherein the Cpf1 effector protein forms a complex with the one or more nucleic acid components, and upon binding of the complex to the target locus of interest 3' to a protospacer adjacent motif (PAM), the effector protein induces modification of the target locus of interest.
16. 16. The method of claim 15, wherein the target locus of interest is intracellular.
17. 17. The method of claim 16, wherein the cell is a eukaryotic cell.
18. 17. The method of claim 16, wherein the cell is an animal or human cell.
19. The method of claim 16 , wherein the cell is a plant cell.
20. 16. The method of claim 15, wherein the target locus of interest is contained in a DNA molecule in vitro.
21. 16. The method of claim 15, wherein the non-naturally occurring or engineered composition comprising a Cpf1 effector protein and one or more nucleic acid components is delivered to the cell as one or more polynucleotide molecules.
22. 16. The method of claim 15, wherein the target locus of interest comprises DNA.
23. 23. The method of claim 22, wherein the DNA is relaxed or supercoiled.
24. The method of claim 15 , wherein the composition comprises a single nucleic acid component.
25. 25. The method of Claim 24, wherein the single nucleic acid component comprises a guide sequence linked to a direct repeat sequence.
26. 16. The method of claim 15, wherein the modification of the target locus of interest is a strand break.
27. 27. The method of claim 26, wherein the strand break comprises a sticky-end type DNA double-strand break with a 4 or 5 nt 5' overhang.
28. 27. The method of claim 26, wherein the target locus of interest is modified by integration of a DNA insert into the sticky-end DNA double-strand break.
29. 16. The method of claim 15, wherein the Cpf1 effector protein comprises one or more nuclear localization signals (NLS).
30. 22. The method of claim 21, wherein the one or more polynucleotide molecules are contained within one or more vectors.
31. 22. The method of claim 21 , wherein the one or more polynucleotide molecules comprise one or more regulatory elements operably configured to express the Cpf1 effector protein and / or one or more of the nucleic acid components, and optionally the one or more regulatory elements comprise an inducible promoter.
32. 22. The method of claim 21, wherein the one or more polynucleotide molecules or the one or more vectors are comprised in a delivery system.
33. 22. The method of claim 21, wherein the system or the one or more polynucleotide molecules are delivered by a particle, a vesicle, or one or more viral vectors.
34. 34. The method of claim 33, wherein the particle comprises a lipid, a sugar, a metal, or a protein.
35. 34. The method of claim 33, wherein the vesicle comprises an exosome or a liposome.
36. 34. The method of claim 33, wherein the one or more viral vectors comprise one or more of an adenovirus, one or more lentivirus, or one or more adeno-associated virus.
37. 16. The method of claim 15, wherein the method is a method of modifying a cell, cell line or organism by manipulation of one or more target sequences in a genomic locus of interest.
38. 38. The cell from the method of claim 37, or a progeny thereof, which comprises a modification that is not present in a cell not subjected to said method.
39. 39. The cell of claim 38, or a progeny thereof, wherein the cell not subjected to the method contains a disorder and the cell from the method has the disorder addressed or corrected.
40. 39. A cell product from the cell of claim 38 or a progeny thereof that is altered in quality or quantity compared to a cell product from a cell not subjected to said method.
41. 41. The cellular product of claim 40, wherein the cells not subjected to the method contain a disorder, and the cellular product reflects the disorder that has been addressed or corrected by the method.
42. 3. An in vitro, ex vivo or in vivo host cell or cell line or progeny thereof comprising the system of claim 1 or 2.
43. 43. The host cell or cell line or progeny thereof of claim 42, wherein the cell is a eukaryotic cell.
44. 44. The host cell or cell line or progeny thereof of claim 43, wherein the cell is an animal cell.
45. 34. The host cell or cell line or progeny thereof of claim 33, wherein the cell is a human cell.
46. 32. The host cell, cell line or progeny thereof of claim 31, comprising a stem cell or stem cell line.
47. 31. The host cell or cell line or progeny thereof of claim 30, wherein the cell is a plant cell.
48. 16. A method for producing a plant having a modified trait of interest encoded by a gene of interest, the method comprising contacting a plant cell with the system of claim 1 or 2 or subjecting the plant cell to the method of claim 15, thereby modifying or introducing the gene of interest, and regenerating a plant from the plant cell.
49. 16. A method for identifying a trait of interest in a plant, said trait of interest being encoded by a gene of interest, said method comprising the step of contacting a plant cell with the system of claim 1 or 2 or subjecting said plant cell to the method of claim 15, whereby said gene of interest is identified.
50. 50. The method of claim 49, comprising introducing the identified gene of interest into a plant cell or plant cell line or plant germplasm and producing a plant therefrom, whereby the plant contains the gene of interest.
51. 51. The method of claim 50, wherein the plant exhibits the trait of interest.
52. A particle comprising the system according to claim 1 or 2.
53. 53. The particle of claim 52, comprising the Cpf1 effector protein complexed with the guide RNA.
54. 16. The system or method of claim 1, 2 or 15, wherein the complex, guide RNA or protein is conjugated to at least one sugar moiety, optionally N-acetylgalactosamine (GalNAc), in particular triantennary GalNAc.
55. Mg 2+ 16. The system or method of claim 1, 2 or 15, wherein the concentration of is from about 1 mM to about 15 mM.