Reengineered CASCADE components and CASCADE complexes
By modifying the type I CRISPR-Cas system, fusing Cas8 and FokI proteins and optimizing guide polynucleotides, the problems of the Cascade complex being difficult to express in eukaryotic cells and low DNA targeting efficiency were solved, achieving efficient genome editing and precise cutting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-12
- Publication Date
- 2026-03-31
AI Technical Summary
The application of type I CRISPR-Cas systems in eukaryotic genome modification is limited, mainly due to the difficulty of heterologous expression of the Cascade complex and the limitations of DNA target cleavage methods.
By modifying type I CRISPR-Cas effector complexes to include fusion proteins such as Cas8 and FokI, and by binding specific guide polynucleotides, the composition and function of the Cascade complex were optimized, enhancing its DNA targeting capability.
It enables efficient genome editing in eukaryotic cells, improves DNA targeting efficiency and cutting precision, expands PAM selectivity, and enhances the flexibility of genome editing.
Smart Images

Figure CN119490976B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese Patent Application No. 201980038867.7, which was filed on June 12, 2019, the entire contents of which are incorporated herein by reference.
[0002] Cross-references to related applications
[0003] This application is a partial continuation of U.S. Patent Application Series 16 / 420,061, filed May 22, 2019, which is a continuation of U.S. Patent Application Series 16 / 420,061, filed January 30, 2019, which is a continuation of U.S. Patent Application Series 16 / 262,773, filed August 17, 2018, which is a continuation of U.S. Patent Application Series 16 / 104,875, filed August 17, 2018, which is a continuation of U.S. Patent Application No. 10,227,576, issued March 12, 2019. This application also claims the benefit of U.S. Provisional Patent Application Series 62 / 684,735, filed June 13, 2018, which is a continuation of U.S. Provisional Patent Application Series 62 / 807,717, filed February 19, 2019, the contents of which are incorporated herein by reference in their entirety.
[0004] Statement regarding federally sponsored research or development
[0005] not applicable.
[0006] sequence list
[0007] This application contains a sequence list electronically submitted in ASCII format, which is incorporated herein by reference in its entirety. An ASCII copy generated on June 12, 2019, is named CBI032-30_ST25.txt and is 3.1 MB in size. Technical Field
[0009] This disclosure generally relates to a modified type I CRISPR-Cas (Cascade) system comprising a multi-protein effector complex, a nucleoprotein complex containing a type I CRISPR-Cas subunit protein and a nucleic acid guide, a polynucleotide encoding a type I CRISPR-Cas subunit protein, and a guide polynucleotide. This disclosure also relates to compositions and methods for preparing and using the modified type I CRISPR-Cas system of the present invention. Background Technology
[0011] Regularly spaced clustered short palindromic repeats (CRISPR) and CRISPR-associated proteins (Cas) constitute the CRISPR-Cas system. The CRISPR-Cas system provides adaptive immunity against foreign polynucleotides in bacteria and archaea (see, for example, Barrangou, R., et al., Science 315:1709-1712 (2007); Makarova, KS, et al., Nature Reviews Microbiology 9:467-477 (2011); Garneau, JE, et al., Nature 468:67-71 (2010); Sapranauskas, R., et al., Nucleic Acids Res. 39:9275-9282 (2011); Koonin, EV, et al., Curr. Opin. Microbiol. 37:67-78 (2017)). Various CRISPR-Cas systems are able to target DNA (Type I; Type II and Type V), RNA (Type II VI), and both DNA and RNA (Type I III) in their natural hosts (see, for example, Makarova, KS, et al., Nat. Rev. Microbiol. 13:722-736 (2015); Shmakov, S., et al., Nat. Rev. Microbiol. 15:169-182 (2017); Abudayyeh, OO, et al., Science 353:1-17 (2016)).
[0012] The classification of CRISPR-Cas systems has reached numerous sub-levels. Koonin, EV, et al. (Curr. Opin. Microbiol. 37:67-78 (2017)) proposed a classification system that considers tag cas genes specific to various types and subtypes of CRISPR-Cas systems. This classification also considers sequence similarity among multiple common Cas proteins, phylogeny of optimally conserved Cas proteins, gene organization, and the structure of CRISPR arrays. This approach provides a classification scheme that divides CRISPR-Cas systems into two distinct types: the first category includes multi-protein effector complexes (Type I (CRISPR-associated complex (“Cascade”) effector complex for antiviral defense), Type III (Cmr / Csm effector complex), and Type IV); and the second category includes single effector proteins (Type II (Cas9), Type V (Cas12a, formerly known as Cpf1), and Type VI (Cas13a, formerly known as C2c2)). In the first system, type I is the most common and diverse, type III is more common in archaea than in bacteria, and type IV is the least common.
[0013] The type I system includes the tag Cas3 protein. The Cas3 protein possesses both helicase and DNase domains responsible for cleaving DNA target sequences. To date, seven subtypes of the type I system (i.e., IA, IB, IC, ID, IE, IF (and variants of IF (e.g., I-Fv1, I-Fv2)) and IU) have been identified, each with a variable number of Cas genes. Type I Cas genes include, but are not limited to, the following genes: cas7, cas5, cas8, cse2, csa5, cas3, cas2, cas4, cas1, and cas6. Examples of organisms with type I systems are as follows: IA, Archaeoglobus fulgidus; IB, Clostridium kluyveri; IC, Bacillus halodurans; IU, Geobacter sulfurreducens; ID, Cyanothece sp. 8802; IE, Escherichia coli K12; IF, Yersinia pseudo-tuberculosis; IF variant, Shewanella putrefaciens CN-32 (Koonin, EV, et al., Curr. Opin. Microbiol. 37:67-78 (2017)).The characteristics of Cas3 protein-mediated DNA cleavage and progressive degradation have been described (see, for example, Plagens, A., et al., Nucleic Acids Res. 42:5125-5138 (2014); Maier, L., et al., RNA Biol. 10:865–874 (2013); Hochstrasser, M., et al., Proc. Natl. Acad. Sci. USA 111:6618–6623 (2014); Sinkunas, T., et al., EMBO J. 30:1335–1342 (2011); Westra, E., et al., Mol. Cell 46:595–605 (2012); Mulepati, S., et al.). al., J. Biol. Chem. 288:22184–22192 (2013); Sinkunas, T., et al., EMBO J. 32: 385–394 (2013); Mulepati, S., et al., J. Biol. Chem. 288: 22184–22192 (2013); Redding, S., et al. al., Cell 163:854–865 (2015); Sinkunas, T., et al., EMBO J. 32: 385–394 (2013); Westra, E., et al., Mol. Cell 46: 595–605 (2012)).
[0014] Type I systems typically encode proteins that combine with CRISPR RNA (crRNA or "guide RNA") to form Cascade complexes. These complexes contain multiple proteins and crRNAs, both transcribed from the CRISPR site. In Type I systems, the initial processing of the crRNA precursor is catalyzed by Cas6. This typically results in the production of crRNA with an 8-nucleotide 5' stalk, a spacer region, and a 3' stalk; both the 5' and 3' stalks are derived from repetitive sequences. In some systems, the 3' stalk forms a stem-loop structure; in others, secondary processing of the 3' end of the crRNA is catalyzed by a ribonuclease (see, for example, van der Oost, J., et al., Nature Reviews Microbiology 12:479-492 (2014)).
[0015] The Cascade effector complex of a type I CRISPR-Cas system contains a backbone of unknown proteins with paralogous repetitive sequences (RAMPs; e.g., Cas7 and Cas5 proteins) containing RNA recognition motif (RRM) folds and additional "large" and "small" subunits (see, e.g., Koonin, EV, et al., Curr. Opin. Microbiol. 37:67-78, (2017), Figure 2). These Cascade effector complexes typically contain a Cas5 subunit and several Cas7 subunits. Such Cascade effector complexes also contain a guide RNA. The Cascade effector complex contains different subunits arranged asymmetrically along the length of the guide RNA. The Cas5 subunit and the large subunit (Cas8 protein) are located at one end of the complex, wrapping around the 5' end of the guide RNA. Several copies of the small subunit contact the guide RNA backbone, which is bound to multiple copies of the Cas7 subunit. The Cas6 subunit (another RAMP protein) binds to the Cascade effector complex primarily by binding to the 3′ stalk (repetition region) of crRNA. The Cas6 subunit is typically used as a repeat-sequence-specific RNase participating in the processing of crRNA precursors; however, in the type I system, Cas5 acts as the repeat-sequence-specific RNase, and Cas6 is absent.
[0016] The major sequences of the CRISPR-Cas type I Cascade subunit proteins have very little sequence identity; however, the presence of homologous RAMP modules and the overall structural similarity of multi-protein effector complexes support a common origin for these effector complexes (see, for example, Koonin, EV, et al., Curr. Opin. Microbiol. 37:67-78 (2017)).
[0017] The adaptive immune mechanism in the type I CRISPR-Cas system mainly involves three phases: adaptation, expression, and intervention. In the adaptation phase, foreign DNA or RNA infects the host, and proteins encoded by various Cas genes bind to the infected DNA or RNA region. Such a region is called the preseptal region. The preseptal region sequence adjacent motif (PAM) is a short nucleotide sequence (e.g., a 2-6 base pair DNA sequence) adjacent to the preseptal region. PAM sequences are typically recognized by the Cas1 subunit / Cas2 subunit protein complex, where the active PAM sensing site is associated with the Cas1 subunit protein (see, e.g., Jackson, SA, et al., Science 356:356(6333)(2017)).
[0018] During the expression phase, a CRISPR array containing multiple spacer repeat elements is transcribed into a single transcript. Each spacer repeat element is processed into a single crRNA by an endonuclease (e.g., type I, Cas6 protein; and type IC, Cas5 protein). Cas subunit proteins are expressed and bind to the crRNA to form the Cascade effector complex.
[0019] The Cascad effector complex scans for foreign polynucleotides in the infected host to identify DNA complementary to the spacer region. In type I systems, interference occurs when the effector complex recognizes a sequence complementary to the spacer region of a neighboring PAM; and the Cas3 protein is recruited to the DNA-binding Cascade effector complex to cleave and progressively digest the foreign polynucleotide.
[0020] Makarova, KS, et al. (Cell 168:946 (2017)) provided summary information on the genes, homologues, Cascade complex, and mechanism of action of the type I CRISPR-Cas system.
[0021] Therefore, the use of the type I CRISPR-Cas system in eukaryotic genome editing is currently limited, partly due to the difficulty of heterologous expression of the Cascade complex and the way the type I CRISPR-Cas system cuts DNA targets. Invention Overview
[0023] This invention generally relates to compositions comprising a modified type I CRISPR-Cas effector complex and its components including protein components, modified or differentially altered guide polynucleotides, and combinations thereof.
[0024] One embodiment of the present invention is a composition comprising:
[0025] The first modified type I CRISPR-Cas effector complex comprises:
[0026] The first Cse2 subunit protein, the first Cas5 subunit protein, the first Cas6 subunit protein, and the first Cas7 subunit protein.
[0027] A first fusion protein comprising a first Cas8 subunit protein and a first FokI, wherein the N-terminus or C-terminus of the first Cas8 subunit protein is covalently linked to the C-terminus or N-terminus of the first FokI via a first linker polypeptide, and wherein the first linker polypeptide has a length of 10 to 40 amino acids.
[0028] Contains a first guide polynucleotide capable of binding to a first spacer region of a first nucleic acid target sequence; and
[0029] The second modified type I CRISPR-Cas effector complex comprises:
[0030] Second Cse2 subunit protein, second Cas5 subunit protein, second Cas6 subunit protein, and second Cas7 subunit protein.
[0031] A second fusion protein comprising a second Cas8 subunit protein and a second FokI, wherein the N-terminus of the second Cas8 subunit protein or the C-terminus of the second Cas8 protein is covalently linked to the C-terminus or N-terminus of the second FokI via a second linker polypeptide, and wherein the second linker polypeptide has a length of 10 to 40 amino acids.
[0032] The second guide polynucleotide contains a second spacer region capable of binding to a second nucleic acid target sequence, wherein the pre-interstitial adjacent motif (PAM) of the second nucleic acid target sequence and the PAM of the first nucleic acid target sequence have a spacer interval of 20 to 42 base pairs.
[0033] In some embodiments, the length of the first linker polypeptide and / or the second linker polypeptide is 15 to 30 amino acids or 17 to 20 amino acids. In one embodiment, the lengths of the first linker polypeptide and the second linker polypeptide are the same.
[0034] The spacing between the second nucleic acid target sequence and the first nucleic acid target sequence includes, but is not limited to, 22 to 40 base pairs, 26 to 36 base pairs, 29 to 35 base pairs, or 30 to 34 base pairs.
[0035] The first FokI and the second FokI can be monomeric subunits that can combine to form homodimers, or different subunits that can combine to form heterodimers.
[0036] In some embodiments, the N-terminus of the first Cas8 subunit protein is covalently linked to the C-terminus of the first FokI via a first linker polypeptide, the C-terminus of the first Cas8 subunit protein is covalently linked to the N-terminus of the first FokI via a first linker polypeptide, the N-terminus of the second Cas8 subunit protein is covalently linked to the C-terminus of the second FokI via a second linker polypeptide, and the C-terminus of the second Cas8 subunit protein is covalently linked to the N-terminus of the second FokI via a second linker polypeptide, as well as combinations thereof. Each of the first and second Cas8 subunit proteins may contain Cas8 subunit proteins with different sequences, or both the first and second Cas8 subunit proteins may contain the same amino acid sequence.
[0037] Similarly, each of the first Cse2 subunit protein and the second Cse2 subunit protein may contain different or the same Cse2 subunit protein amino acid sequences, each of the first Cas5 subunit protein and the second Cas5 subunit protein may contain different or the same Cas5 subunit protein amino acid sequences, each of the first Cas6 subunit protein and the second Cas6 subunit protein may contain different or the same Cas6 subunit protein amino acid sequences, each of the first Cas7 subunit protein and the second Cas7 subunit protein may contain different or the same Cas7 subunit protein amino acid sequences, and combinations thereof.
[0038] In a preferred embodiment, the guide polynucleotide comprises RNA.
[0039] In another embodiment, the invention includes a modified type I CRISPR Cas3 mutant protein (“mCas3 protein”) capable of reducing the movement along DNA relative to the wild-type type I CRISPR Cas3 protein (“wtCas3 protein”).
[0040] The present invention also includes intracellular genome editing using the above-described composition, and a method for preparing the above-described composition.
[0041] Other embodiments of the invention will readily become apparent to those skilled in the art in light of the disclosure herein.
[0042] Brief description of the attached figures
[0043] The diagrams are not drawn to scale, nor are they drawn to a fixed scale. The positions of the indicators are approximate.
[0044] Figure 1A A generalized view of the type I CRISPR-Cas effector complex is presented. Figure 1B A generalized view of type I CRISPR-Cas crRNA is presented.
[0045] Figure 2A , Figure 2B and Figure 2C Illustrative examples of two modified type I CRISPR-Cas effector complexes with fusion domains that bind to adjacent spacer sequences are presented.
[0046] Figure 3A and Figure 3B Examples of proteins arranged in a circular pattern are presented.
[0047] Figure 4A , Figure 4B , Figure 5A , Figure 5B , Figure 6A , Figure 6B , Figure 6C , Figure 7A , Figure 7B , Figure 8 , Figure 9 , Figure 10A and Figure 10B Various examples of the modified type I CRISPR-Cas effector complex of the present invention are shown.
[0048] Figure 11A and Figure 11B An example of a substrate channel is shown.
[0049] Figure 12A , Figure 12B and Figure 12C This presents a generalized view of the site-directed recruitment of functional protein domains fused to the Cascade subunit by the dCas9:NATNA complex.
[0050] Figure 13A , Figure 13B , Figure 14A , Figure 14B and Figure 14C An example of the modified type I CRISPR-Cas effector complex of the present invention is shown.
[0051] Figure 15A , Figure 15B , Figure 15C , Figure 16A , Figure 16B , Figure 16C , Figure 17A , Figure 17B , Figure 17C , Figure 18A , Figure 18B , Figure 18C , Figure 18D , Figure 19A , Figure 19B , Figure 20A and Figure 20B Examples of the modified type I CRISPR-Cas effector complex of the present invention and its usage are presented.
[0052] Figure 21A , Figure 21B , Figure 21C , Figure 21D , Figure 22A , Figure 22B , Figure 22C and Figure 22D An embodiment of the present invention using a Cas3 protein containing active endonuclease activity is shown.
[0053] Figure 23A , Figure 23B , Figure 23C , Figure 23D , Figure 23E , Figure 24 , Figure 25 , Figure 26 and Figure 27 Schematic diagrams of various Cascade component expression systems are presented.
[0054] Figure 28 , Figure 29 , Figure 30 , Figure 31A , Figure 31B , Figure 32 , Figure 33A , Figure 33B and Figure 34 Data related to genome editing using the modified Cascade system of this invention are presented.
[0055] Figure 35 An example of a minimal CRISPR array containing paired guide RNAs (gRNAs) is shown.
[0056] Figure 36A , Figure 36B , Figure 36C and Figure 36D Data related to genome editing in human cells via RNP and plasmid delivery of a modified type I CRISPR-Cas complex are presented.
[0057] Figure 37A , Figure 37B , Figure 37C , Figure 37D , Figure 37E , Figure 37F and Figure 37G The data related to the repair results is presented.
[0058] Figure 38A , Figure 38B and Figure 38C Data related to how mismatches between gRNAs and target DNA inhibit genome editing of modified type I CRISPR-Cas complexes are presented.
[0059] Figure 39A , Figure 39B , Figure 39C and Figure 39D Data related to expanded screening of PAM selectivity for three Cascade homologue variants are presented.
[0060] Figure 40A , Figure 40B , Figure 40C , Figure 40D , Figure 40E and Figure 40FData related to exemplary variations in the editing efficiency of the modified type I CRISPR-Cas complex are presented.
[0061] Figure 41A , Figure 41B and Figure 41C Data related to expanded screening of FokI-Cas8 connector length and spacer spacing for three Cascade homologue variants are presented.
[0062] Figure 42A and Figure 42B An example of PCR amplification using oligomers as templates is shown.
[0063] Figure 43 Data on the percentage of genome editing are presented, shown as a function of FokI-Cascade homologue variants and spacer intervals.
[0064] Figure 44 Linear plots of the functional domains of the EcoCas3 protein and the relative positions of mutants created within the sequence are shown.
[0065] Figure 45A , Figure 45B , Figure 45C and Figure 45D Data related to genome editing using the EcoCascade RNP complex, which contains wild-type or mutant EcoCas3 proteins, are shown.
[0066] Figure 46A , Figure 46B , Figure 46C , Figure 47A and Figure 47B Data related to the dCas9-VP64 / sgRNA RNP complex roadblocks and their impact on EcoCascade RNP complex cleavage targets are presented.
[0067] Figure 48 The example edit data for Cas3[D452A] / -EcoCascade or mCas3[D452A]-EcoCascade is shown.
[0068] Figure 49 Data on genome editing at eight TRAC target sites using the PseCascade RNP complex are presented.
[0069] By citation and inclusion in this article
[0070] All patents, publications and patent applications cited in this specification are incorporated herein by reference as if each individual patent, publication or patent application were expressly and individually indicated as being incorporated herein by reference in its entirety for all purposes. Invention Details
[0072] It should be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in this specification and claims, the singular forms “a,” “an,” and “the” include plural indicators unless the context clearly indicates otherwise. Thus, for example, reference to “a polynucleotide” includes one or more polynucleotides, and reference to “a vector” includes one or more vectors.
[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Although other methods and materials similar to or equivalent to those described herein may be used in this invention, preferred materials and methods are described herein.
[0074] Based on the teachings of this specification and the examples, those skilled in the art can apply conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant polynucleotides, such as those taught in the following standard texts: Cellular and Molecular Immunology, 9th Edition, AKA Abbas., et al., Elsevier (2017), ISBN 978-0323479783; Cancer Immunotherapy Principles and Practice, 1st Edition, LHB Butterfield, et al., Demos Medical (2017), ISBN 978-1620700976; Janeway's Immunobiology, 9th Edition, Kenneth Murphy, Garland Science (2016), ISBN 978-0815345053; Clinical Immunology and Serology: A Laboratory Perspective, 4th Edition, C. Dorresteyn Stevens, et al., FADAvis Company(2016),ISBN978-0803644663;Antibodies:A Laboratory Manual,Second Edition,EAGreenfield,ColdSpring Harbor Laboratory Press(2014),ISBN 978-1-936113-81-1;Culture of AnimalCells:A Manual of Basic Technique and Specialized Applications, seventh edition, RIFreshney, Wiley-Blackwell (2016), ISBN 978-1118873656; Transgenic Animal Technology, third edition: ALaboratory Handbook, CAPinkert, Elsevier (2014), ISBN 978-0124104907; The Laboratory Mouse, second edition, H.Hedrich, Academic Press(2012), ISBN978-0123820082; Manipulating the Mouse Embryo: A Laboratory Manual, Fourth Edition, R.Behringer, et al., Cold Spring Harbor Laboratory Press (2013), ISBN 978 - 1936113019; PCR 2: A Practical Approach, M.J. McPherson, et al., IRL Press (1995), ISBN 978 - 0199634248; Methods in Molecular Biology (Series), J.M. Walker, ISSN1064 - 3745, Humana Press; RNA: A Laboratory Manual, D.C. Rio, et al., Cold Spring Harbor Laboratory Press (2010), ISBN 978 - 0879698911; Methods in Enzymology (Series), Academic Press; Molecular Cloning: A Laboratory Manual (Fourth Edition), M.R. Green, et al., Cold Spring Harbor Laboratory Press (2012), ISBN 978 - 1605500560; Bioconjugate Techniques, Third Edition, G.T. Hermanson, Academic Press (2013), ISBN 978 - 0123822390; Methods in Plant Biochemistry and Molecular Biology, W.V. Dashek, CRC Press (1997), ISBN 978 - 0849394805; Plant Cell Culture Protocols (Methods in Molecular Biology), V.M. Loyola - Vargas, et al., Humana Press (2012), ISBN 978 - 1617798177; Plant Transformation Technologies, C.N. Stewart, et al., Wiley - Blackwell (2011), ISBN 978 - 0813821955; Recombinant Proteins from Plants (Methods in Biotechnology), C. Cunningham, et al.,Humana Press(2010),ISBN 978-1617370212;Plant Genomics:Methods and Protocols(Methods in MolecularBiology),W.Busch,Humana Press(2017),ISBN 978-1493970018;Plant Biotechnology:Methods in Tissue Culture and Gene Transfer,R.Keshavachandran,et al., Orient Blackswan (2008), ISBN 978-8173716164. .
[0075] Regularly spaced clustered short palindromic repeats (CRISPR) and associated CRISPR-related proteins (Cas proteins) constitute the CRISPR-Cas system (see, for example, Barrangou, R., et al., Science 315:1709-1712 (2007)).
[0076] As used herein, “Cas protein,” “CRISPR-Cas protein,” “CRISPR-Cas subunit protein,” and “Cas subunit protein,” unless otherwise indicated, refer to type I CRISPR-Cas protein. Typically, for use in aspects of the present invention, the Cas subunit protein is capable of interacting with one or more homologous polynucleotides (most typically, crRNA) to form a type I effector complex (most typically, RNP complex).
[0077] Over time, various conventions have been used to name the genes encoding Cascade in the IE-type CRISPR-Cas system, which can be confusing when comparing recent and older literature. Generally, this specification uses the nomenclature shown in Koonin, E., et al. (Curr. Opin. Microbiol. 37:67-78 (2017)), where the gene sequence in the reference *E. coli* K12 operon is: cas3, cas8, cas11, cas7, cas5, cas6, cas1, and cas2. For simplicity, the "e" qualifier in cas8e is sometimes used to distinguish the cas8 gene between different subtypes in the type I system. The stoichiometry of wild-type *E. coli* IE-type CRISPR-Cas is Cas51-Cas61-Cas76-Cas81-Cas112-gRNA1.
[0078] However, for cross-reference purposes: cas8 was formerly known as cse1 and casA, and also as the "large subunit"; cas11 was formerly known as cse2 and casB, and also as the "small subunit"; cas7 was formerly known as cse4 and casC; cas5 was formerly known as casD, and sometimes given the qualifier cas5e; and cas6 was formerly known as cse3 and casE, and usually given the qualifier cas6e. Table 1 lists the genes encoding Cas subunit proteins.
[0079]
[0080]
[0081] *As defined by Makarova, KS, et al., Nat. Rev. Microbiol. 13:722-736 (2015); Koonin, EV, et al., Curr Opin Microbiol. 37:67-78 (2017).
[0082] PAM sequences are typically recognized by the Cas1 / Cas2 subunit complex, with the active PAM sensing site associated with the Cas1 subunit (see, e.g., Jackson, SA, et al., Science 356:356(6333)(2017)). Cas1 and Cas2 proteins are present in the vast majority of known CRISPR-Cas systems and are sufficient to insert spacer regions into CRISPR boxes (see, e.g., Yosef, I, et al., Nucleic Acids Res.40:5569–5576(2012)). These two proteins form a complex used in adaptation processes. The Cas1 protein possesses endonuclease activity required for spacer region integration, while the Cas2 protein appears to perform a non-enzymatic function (see, e.g., Nunz, J., et al., Nat Struct Mol Biol. 21:528–534 (2014); Richter, C., et al., PLoS One. 2012; 7:e49549). The Cas1-Cas2 protein complex represents a highly conserved information processing module of the CRISPR-Cas system, which appears to be quasi-autonomous relative to the rest of the system (see, e.g., Makarova, K., et al., Methods Mol. Biol. 1311:47–75 (2015)). The endonuclease Cas1 protein is essential to ensure the unique ability of the CRISPR system to retain memory of previous infectious agent encounters.
[0083] The terms “Type I CRISPR-Cas effector complex,” “Type I CRISPR-Cas nucleoprotein (NP) complex,” “Cascade nucleoprotein (NP) complex,” and “Type I nucleoprotein (NP) complex” are used interchangeably herein and generally refer to the Cascade protein that forms a complex containing a guide polynucleotide. When referring to the protein component of the Cascade NP complex, the terms “Cascade complex” and “Type I complex” are generally used. The terms “Cascade RNP complex,” “Type I CRISPR-Cas RNP complex,” and “Type I RNP complex” refer to the Cascade complex containing crRNA relative to the more general guide polynucleotide (i.e., as in the Cascade NP complex). An example of a wild-type Type I CRISPR-Cas effector complex is shown in… Figure 1A middle. Figure 1A This is an adaptation of Makarova, KS, et al. (Cell 168:946(2017); Makarova, K., et al., Nature Reviews Microbiology 13:722-736(2015)). Figure 1A The diagram shows six Cas7 proteins, Cas5 proteins, Cas8 proteins, two Cse2 proteins, Cas6 proteins, and crRNA that combine to form the Cascade complex. Figure 1A Cas7, Cas5, Cas8, Cse2, and Cas6; the dashed box around Cas6 indicates its interaction with the crRNA hairpin; cRNA is shown as the black line including the hairpin. The complex is capable of binding to nucleic acid target sequences. In the wtCas3 protein ( Figure 1A The Cas3 subunit (enclosed in the dashed box) binds to the Cascade complex, which then cleaves the target nucleic acid sequence. As shown in Table 1, the total number of some Cas subunit proteins may vary within the Cascade complex.
[0084] In this article, “Cas3” and “Cas3 protein” are used interchangeably to refer to type I CRISPR-Cas3 protein, its modifications, and variants. The type I CRISPR-Cas effector complex binds to foreign DNA complementary to the crRNA guide and recruits Cas3, a trans-acting nuclease-helicase required for target degradation. The Cas3 protein possesses a motif characteristic of helicases from superfamily 2 and includes a DEAD / DEAH box region and a conserved C-terminal domain. Cas3 proteins and their variants are known in the art (see, for example, Westra, ER, et al., Mol. Cell. 46:595–605 (2012); Sinkunas, T., et al., EMBO J. 30:1335-1342 (2011); Beloglazova, N., et al., EMBO J. 30:4616-4627 (2011); Mulepati, S., et al., J. Biol. Chem. 286:31896-31903 (2011)). As used herein, the term “mCas3 protein” refers to a Cas3 protein containing one or more mutations relative to its corresponding wtCas3 protein. mCas3 proteins include, but are not limited to: mCas3 proteins (e.g., Examples 23A, 23B, and 23C), dblmCas3 proteins (e.g., Examples 26A, 26B, and 26C), and dCas3* (mutant Cas3 proteins that do not have any nuclease activity and / or helicase activity).
[0085] As used herein, the term "nuclease" refers to an enzyme capable of cleaving phosphodiester bonds, such as those linking two nucleotides, found in double-stranded (ds) nucleic acids (e.g., dsDNA, genomic DNA (gDNA), dsRNA), single-stranded (ss) nucleic acids (e.g., ssDNA, RNA), or hybrid dsRNA / DNA. "Endonuclease" typically affects ss- (crack) or ds- gaps in its target molecules. An example of a DNA endonuclease is the FokI enzyme. "FokI endonuclease" and "FokI" are used interchangeably herein and refer to the FokI enzyme, FokI homologues, the enzymatic active domain of the FokI enzyme, and variants of the FokI enzyme. FokI dimerization is typically required for DNA cleavage. FokI dimers can comprise two monomeric subunits that combine to form a homodimer or two different monomeric subunits that combine to form a heterodimer (see, for example, Bitinaite, J., et al., Proc. Natl. Acad. Sci. USA 95:10570-10575 (1998); Ramalingam, S., et al., J. Mol. Biol. 405:630-641 (2011)). An example of a FokI variant is the Sharkey variant described by Guo, et al. (Guo, J., et al., J. Mol. Biol. 400:96-107 (2010)). Other DNA and RNA nucleases are known in the art.
[0086] As used herein, “CRISPR RNA,” “crRNA,” and “guide RNA” refer to one or more RNAs that can interact with Cas subunit proteins to form a guide complex that preferentially binds to a type I effector complex of a polynucleotide (as opposed to a polynucleotide that does not contain a target nucleic acid sequence). As used herein, “guide” and “guide polynucleotide” refer to a polynucleotide component of a type I effector complex comprising a ribonucleotide base (e.g., RNA) and ribose, as well as different components and combinations thereof, including but not limited to: deoxyribonucleotide bases, nucleotide analogs, modified nucleotides, different nitrogenous bases, fundamentally different nucleotide bases, chemically different molecules, mixtures of bases (e.g., RNA bases, DNA bases, and / or modified bases), and combinations thereof, as well as synthetic backbones, naturally occurring backbones, non-naturally occurring backbones, fundamentally different backbone residues, chemically different residues or bonds, modified backbones, mixtures (e.g., ribose and deoxyribose components of a backbone), and combinations thereof. Examples of guide polynucleotides are described herein. An example of type I CRISPR-Cas crRNA binding to a nucleic acid target sequence via a crRNA spacer region is shown. Figure 1B middle. Figure 1B Adapted from Hochstrasser, ML, et al., Mol. Cell 63:840-851 (2016). Figure 1B In China, PAM ( Figure 1B (104) binds to the nucleic acid target sequence, and shows the 5' and 3' strands of the double-stranded nucleic acid ( Figure 1B Vertical lines represent hydrogen bonds. Guide polynucleotides ( Figure 1B ,106) typically includes a 5' handle region ( Figure 1B ,101), the interval region containing the seed region ( Figure 1B ,103), and a 3' hairpin containing two hydrogen bond repeating regions ( Figure 1B , 102); the horizontal line represents hydrogen bonds. This paper discusses PAM sequences associated with numerous type I Cascade homologues. PAM sequences are adjacent pre-spacer sequences ( Figure 1B ,105). Figure 1B The Cascade complex spacer region that binds to the nucleic acid target sequence is shown. Figure 1B (The vertical lines represent hydrogen bonds). Figure 1B The anterior septum region is also shown. Figure 1B The spacer region can contain a crRNA region of approximately 6 to approximately 56 nucleotides, where the spacer region is complementary to the nucleic acid target sequence in the polynucleotide. In IE-type CRISPR-Cas systems, the spacer region length can be altered to finely tune Cascade activity. The Cascade complex can be incorporated with additional Cas7 subunits, with 6 nucleotides added to the crRNA spacer region, and additional Cse2 subunits, with 12 nucleotides added to the spacer region (see, e.g., Luo, ML, et al., Nucleic Acids Res. 44(15):7385-7394(2016)). The spacer region typically contains a region of approximately 32 to approximately 36 nucleotides.
[0087] The terms “spacer region,” “spacer region sequence,” and “nucleic acid target binding sequence” are used interchangeably in this article.
[0088] The terms “target,” “target sequence,” “nucleic acid target sequence,” and “intermediate target sequence” are used interchangeably in this document to refer to a nucleic acid sequence that is fully or partially complementary to the target-binding sequence (e.g., the spacer region of crRNA) guided by the Cascade nucleoprotein complex (e.g., the Cascade RNP complex). Typically, the target-binding sequence is chosen to be 100% complementary to the target sequence to which the Cascade nucleoprotein complex targets; however, a lower percentage of complementarity can be used to reduce binding to the target sequence. “Off-target” sequence binding refers to the binding of the Cascade nucleoprotein complex to a nucleic acid sequence with less than 100% complementarity to the target-binding sequence (spacer region). Double-stranded DNA sequences typically contain the target sequence on one strand. Figure 1B (The hydrogen that binds to the guide RNA). The "target region" contains the nucleic acid target sequence.
[0089] As used herein, a “stem element” or “stem structure” refers to two nucleic acid strands known or predicted to form a double-stranded region (“stem element”). A “stem-loop element” or “stem-loop structure” refers to a stem structure in which the 3’ end sequence of one strand is covalently linked to the 5’ end sequence of the second strand by a nucleotide sequence that is typically a single-stranded nucleotide (“stem-loop element nucleotide sequence”). In some embodiments, the loop element comprises a loop element nucleotide sequence of about 3 to about 20 nucleotides in length, preferably about 4 to about 10 nucleotides in length. In a preferred embodiment, the loop element nucleotide sequence is a single-stranded nucleotide sequence of unpaired nucleic acid bases that generate the stem element within the loop element nucleotide sequence without hydrogen bonding interactions. The term “hairpin element” is also used herein to refer to a stem-loop structure. Such structures are well known in the art. Base pairing may be precise; however, as is known in the art, stem elements do not require precise base pairing. Therefore, stem elements may include one or more mismatched or unpaired bases. Examples of stem-loop structures in guide polynucleotides include Figure 1B As shown.
[0090] The terms “linker element nucleotide sequence,” “linker nucleotide sequence,” and “linker polynucleotide” are used interchangeably herein and refer to a single-stranded or double-stranded nucleic acid sequence of one or more nucleotides covalently attached to a first nucleic acid sequence (e.g., 5'-linker nucleotide sequence-first nucleic acid sequence-3'). In some embodiments, the linker nucleotide sequence links two different nucleic acid sequences to form a single polynucleotide (e.g., 5'-first nucleic acid sequence-linker nucleotide sequence-second nucleic acid sequence-3'). Other examples of linker nucleotide sequences include, but are not limited to, 5'-first nucleic acid sequence-linker nucleotide sequence-3' and 5'-linker nucleotide sequence-first nucleic acid sequence-linker nucleotide sequence-3'. In some embodiments, the linker element nucleotide sequence may be a single-stranded nucleotide sequence of unpaired nucleic acid bases that do not interact with each other through hydrogen bonding to create secondary structures (e.g., stem-loop structures) within the linker element nucleotide sequence. In some embodiments, the two linker element nucleotide sequences may interact with each other through hydrogen bonds between the two linker element nucleotide sequences. In some embodiments, the linker polynucleotide encodes a “linker polypeptide.” Such linker polynucleotides typically link the 3' end of a first polynucleotide encoding a first polypeptide to the 5' end of a second polynucleotide encoding a second polypeptide, thereby forming a single polynucleotide encoding a fusion protein comprising N-first polypeptide-linker polypeptide-second polypeptide-C. In some embodiments of the invention, more than two polypeptide sequences can be tandemly linked by a linker polypeptide (e.g., N-first polypeptide-first linker polypeptide-second polypeptide-second linker polypeptide-third polypeptide-C). The terms "linker polypeptide," "linker polypeptide sequence," "amino acid linker sequence," and "linker sequence" are used interchangeably herein.
[0091] As used in this article, "linking nucleotide sequence" refers to a single-stranded nucleic acid sequence linker sequence that covalently links a first nucleic acid sequence to a second nucleic acid sequence.
[0092] As used herein, the terms “interspacer,” “interspacerregion,” and “interspacer-in” are used interchangeably and refer to the generally PAM-in distance between the PAM of a first nucleic acid target sequence (e.g., a first DNA target sequence) and the PAM of a second nucleic acid target sequence (e.g., a second DNA target sequence), wherein a first type I CRISPR-Cas effector complex contains a first spacer region capable of binding the first nucleic acid target sequence, and a second type I CRISPR-Cas effector complex contains a second spacer region capable of binding the second nucleic acid target sequence. Figure 2A , Figure 2B and Figure 2CIllustrative examples of two type I CRISPR-Cas effector complexes containing fusion proteins are presented. Figure 2A “Cascade1”, a solid-lined box encompassing “crRNA1”; and “Cascade2”, a dashed-lined box encompassing “crRNA2”. Figure 2A “FP1” and “FP2” are represented as circular portions; for example, FP1 and FP could be FokI), which are linked to each Cascade complex via linker polynucleotides. Figure 2A "Linker 1" and "Linker 2"), in which the CRISPR-Cas effector complex binds to a nearby nucleic acid target sequence on the double-stranded DNA. Figure 2A "dsDNA" is represented by paired horizontal dashed lines. This indicates the PAM sequence associated with each nucleic acid target sequence. Figure 2A “PAM1” is a hollow frame, and “PAM2” is a hollow frame. Figure 2A The spacer interval between two target sites in a PAM-containing (PAM-containing / PAM-containing) configuration is shown (displayed as...). Figure 2A (The horizontal double-headed line at the top). Figure 2B The spacer interval between two target sites in the PAM-containing / PAM-free configuration is shown (displayed as...). Figure 2B (The horizontal double-headed line at the top). Figure 2C The spacer interval between two target sites is shown in the PAM-free (PAM-free / PAM-free) configuration. Figure 2C (The horizontal double-headed line at the top). Figure 2A , Figure 2B and Figure 2C The separation of the two strands of dsDNA is also shown. The Cascade complex recognizes the dsDNA target sequence adjacent to the PAM. The PAM sequence is recognized by Cse1. Base pairing between the crRNA and the complementary target DNA strand results in the production of an R loop with a substituted non-complementary target DNA strand (see, for example, Beloglazova, N., et al., Nucleic Acids Res. 43:530–543 (2015)).
[0093] As used herein, the term "homologous" refers to interacting biomolecules such as cell surface receptors (e.g., chemokine receptors) and their ligands (e.g., chemokines expressed on tumor cells or in the tumor microenvironment); site-directed peptides and their guides; site-directed peptide / guide complexes (i.e., nucleoprotein complexes) capable of site-directed binding to nucleic acid target sequences complementary to the guide-binding sequence; and the like. Additionally, the term "homologous" refers to a group of Cas subunit proteins (e.g., Cse2, Cas5, Cas6, Cas7, and Cas8), and one or more guide polynucleotides (e.g., type I CRISPR-Cas RNA) capable of forming nucleoprotein complexes that are site-directed to bind to nucleic acid target sequences complementary to the spacer region present in one or more guide polynucleotides.
[0094] The terms “wild-type,” “naturally occurring,” and “unmodified” are used herein to refer to the typical (or most common) form, appearance, phenotype, or strain that is naturally occurring; for example, the typical form in which it occurs in cells, organisms, polynucleotides, proteins, macromolecular complexes, genes, RNA, DNA, or genomes, and which can be isolated from natural sources. The wild-type form, appearance, phenotype, or strain serves as the original parent prior to anticipated modifications, alterations, mutations, and / or significantly different structural changes. Therefore, mutants, variants, engineered, recombinant, and modified forms are non-wild-type forms.
[0095] The terms “modified,” “genetically modified,” “genetically altered,” “recombinant,” “modified,” “non-natural,” and “unnatural” refer to intentional human or mechanical manipulation of the genome of an organism or cell. The terms encompass methods of genome modification, including genome editing as defined herein, as well as techniques for altering gene expression or inactivation, enzyme engineering, directed evolution, knowledge-based design, random mutagenesis methods, gene shuffling, codon optimization, etc. Methods used for genetic modification are known in the art.
[0096] The terms "covalent bond," "covalently attached," "covalently bound," "covalently linked," "covalently connected," and "molecular bond" are used interchangeably in this document and refer to shared chemical bonds involving electron pairs between atoms. Examples of covalent bonds include, but are not limited to, phosphodiester bonds, thiophosphate bonds, disulfide bonds, and peptide bonds (-CO-NH-).
[0097] The terms “non-covalent bond,” “non-covalent attachment,” “non-covalent bond,” “non-covalent linkage,” “non-covalent interaction,” and “non-covalent connection” are used interchangeably in this document and refer to any relatively weak chemical bond that does not involve the sharing of electron pairs. Multiple non-covalent bonds often stabilize the structure of macromolecules and mediate specific interactions between molecules. Examples of non-covalent bonds include, but are not limited to, hydrogen bonds, ionic interactions (e.g., Na+), and ionic interactions (e.g., Na+). + Cl - ), van der Waals interactions and hydrophobic bonds.
[0098] As used herein, the terms “hydrogen-bonded,” “hydrogen-base pair,” and “hydrogen-bonded” are used interchangeably and refer to both typical and atypical hydrogen-bonded combinations, including, but not limited to: “Watson-Crick-Hydrogen-Binded Base Pairs” (WC-Hydrogen-Binded Base Pairs or WC Hydrogen Bonding); “Hoogsteen-Hydrogen-Binded Base Pairs” (Hoogsteen Hydrogen Bonding); and “Wobble-Binded Base Pairs” (Wobble-Binded Hydrogen Bonding). WC hydrogen bonding, including reverse WC hydrogen bonding, refers to purine-pyrimidine base pairing, such as adenine:thymine, guanine:cytosine, and uracil:adenine. Hoogsteen hydrogen bonding, including reverse Hoogsteen hydrogen bonding, refers to a change in base pairing in nucleic acids, where two nucleic acid bases of one type on each strand are held together by hydrogen bonds in the major groove. This non-WC hydrogen bonding allows a third strand to wrap around the double helix and form a triple-stranded helix. Wobble hydrogen bonds, including reverse wobble hydrogen bonds, refer to pairings between two nucleotides in an RNA molecule that do not follow the Watson-Crick pairing rule. There are four main wobble base pairs: guanine:cytosine, inosine (hypoxanthine):cytosine, inosine-adenine, and inosine-cytosine.The rules governing typical and atypical hydrogen bonding are known to those skilled in the art (see, for example, *The RNAWorld, 3rd Edition (Cold Spring Harbor Monograph Series)*, RF. Gesteland, Cold Spring Harbor Laboratory Press (2005), ISBN 978-0879697396; *The RNAWorld, 2nd Edition (Cold Spring Harbor Monograph Series)*, RF. Gesteland, et al., Cold Spring Harbor Laboratory Press (1999), ISBN 978-0879695613; *The RNAWorld (Cold Spring Harbor Monograph Series)*, RF. Gesteland, et al., Cold Spring Harbor Laboratory Press (1993), ISBN 978-0879694562; see, for example, *Appendix 1: Structures of Base Pairs Involving at Least Two Hydrogen Bonds*, I. Tinoco; *Principles of Nucleic Acid Structure*, W. Saenger, Springer International Publishing). AG (1988), ISBN 978-0-387-90761-1; Principles of Nucleic Acid Structure, first edition, S. Neidle, Academic Press (2007), ISBN 978-01236950791).
[0099] The terms “connect,” “connected,” and “connecting” are used interchangeably in this article and refer to covalent or non-covalent bonds between two macromolecules (e.g., polynucleotides, proteins, etc.).
[0100] As used herein, the terms “nucleic acid sequence,” “nucleotide sequence,” and “oligonucleotide” are used interchangeably and refer to a polymeric form of nucleotides. As used herein, the term “polynucleotide” refers to a polymeric form of nucleotides having a 5’ end and a 3’ end and may contain one or more nucleic acid sequences. “Circular polynucleotide” refers to a polynucleotide having a covalent bond between its 5’ end and its 3’ end, thus forming a circular polynucleotide. Nucleotides can be deoxyribonucleotides (DNA), ribonucleotides (RNA), analogs of the above, or combinations thereof (e.g., as described in the context of the above-described guide polynucleotides) and can have any length. Polynucleotides can perform any function and can have a variety of secondary and tertiary structures. The term includes native nucleotides and known analogs of nucleotides modified in the base, sugar, and / or phosphate ester moieties. Analogs of a particular nucleotide have the same base pairing specificity (e.g., analogs of A base pairs versus T). Polynucleotides may contain one or more modified nucleotides. Examples of modified nucleotides include, but are not limited to, fluorinated nucleotides, methylated nucleotides, and nucleotide analogs. Nucleotide structures can be modified before or after polymer assembly. After polymerization, polynucleotides can be further modified, for example, by conjugation to a labeled component or a target-binding component. The nucleotide sequence may contain non-nucleotide components. It also includes modified backbone residues or linkages, i.e., synthetic, naturally occurring, and / or non-naturally occurring nucleic acids that have similar binding properties to a reference polynucleotide (e.g., DNA or RNA). Examples of such analogues include, but are not limited to, phosphate thioesters, aminophosphate esters, methylphosphonates, chiral methylphosphonates, 2-O-methylribonucleotides, peptide-nucleic acids (PNAs), and locked nucleic acids (LNAs). TM (Exiqon, Inc., Woburn, MA) nucleoside, ethylene glycol nucleic acid, bridged nucleic acid, and morpholino structure.
[0101] Peptide-nucleic acids (PNAs) are synthetic homologues of nucleic acids in which the polynucleotide phosphate sugar backbone is replaced by a flexible pseudopeptide polymer, and the nucleobases are linked to the polymer. PNAs have the ability to hybridize with complementary sequences of RNA and DNA with high affinity and specificity.
[0102] In phosphate-thioester nucleic acids, the phosphate-thioester (PS) bond replaces the sulfur atom with a non-bridging oxygen atom in the polynucleotide phosphate backbone. This modification makes the internucleotide linkages resistant to nuclease degradation. In some embodiments, a phosphate-thioester bond is introduced between the last 3-5 nucleotides at the 5' or 3' end of the polynucleotide sequence to inhibit exonuclease degradation. Placing a phosphate-thioester bond throughout the oligonucleotide also helps reduce endonuclease degradation.
[0103] Threonine nucleic acid (TNA) is an artificial genetic polymer. The backbone of TNA consists of repeating threose groups linked by phosphodiester bonds. TNA polymers are resistant to nuclease degradation. TNA can self-assemble into a double-stranded structure via base-pair hydrogen bonding.
[0104] The linker can be reversed into the polynucleotide using "reverse phosphoramiditization" (see, for example, www.ucalgary.ca / dnalab / synthesis / -modifications / linkages). The 3'-3' linker at the polynucleotide end stabilizes the polynucleotide against exonuclease degradation by creating an oligonucleotide with two 5'-OH ends but lacking a 3'-OH end. Typically, such a polynucleotide has a phosphoramidit group at the 5'-OH position and a dimethoxytriphenylmethyl (DMT) protecting group at the 3'-OH position. Typically, the DMT protecting group is at the 5'-OH, and the phosphoramidititium is at the 3'-OH.
[0105] Polynucleotide sequences are shown in the conventional 5'–3' orientation in this paper unless otherwise indicated.
[0106] As used herein, “sequence identity” generally refers to the percentage of nucleotide base or amino acid identity obtained by comparing a first polynucleotide or polypeptide with a second polynucleotide or polypeptide using algorithms with different weighting parameters. Sequence identity between two polynucleotides or two polypeptides can be determined using sequence alignment with various methods and computer parameters (e.g., BLAST, CS-BLAST, PSI-BLAST, FASTA, HMMER, L-ALIGN, etc.) available on the World Wide Web, including but not limited to GENBANK (www.ncbi.nlm.nih.gov / genbank / ) and EMBL-EBI (www.ebi.ac.uk). Sequence identity between two polynucleotides or two polypeptide sequences is typically calculated using standard default parameters of various methods or computer programs. A high degree of sequence identity between two polynucleotides or two polypeptides, as used herein, is typically from about 90% to 100% identity, for example, about 90% identity or higher, preferably about 95% identity or higher, and more preferably about 98% identity or higher. As used herein, moderate sequence identity between two polynucleotides or two polypeptides is typically about 80% to about 85% identity, for example, about 80% identity or higher, preferably about 85% identity. Low sequence identity between two polynucleotides or two polypeptides is typically about 50% to 75% identity, for example, about 50% identity, preferably about 60% identity, more preferably about 75% identity. For example, Cas proteins containing amino acid substitutions (e.g., IE type Cse2, Cas5, Cas6, Cas7 and / or Cas8) may have low, moderate, or high sequence identity in length with reference Cas proteins (e.g., wild type IE type Cse2, Cas5, Cas6, Cas7 and / or Cas8, respectively). As another example, guide polynucleotides may have low, moderate, or high sequence identity in length compared to reference wild-type guide polynucleotides that complex with reference Cas proteins (e.g., guide polynucleotides that form complexes with IE-type Cse2, Cas5, Cas6, Cas7, and / or Cas8).
[0107] As used herein, “hybridization,” “hybridize,” or “hybridizing” is the process of combining two complementary single-stranded DNA or RNA molecules through hydrogen base pairing to form single- or double-stranded molecules (DNA / DNA, DNA / RNA, RNA / RNA). Hybridization strictness is typically determined by the hybridization temperature and the salt concentration of the hybridization buffer; for example, high temperature and low salt provide highly strict hybridization conditions. Examples of salt concentration and temperature ranges for different hybridization conditions are as follows: high strictness, approximately 0.01 M to approximately 0.05 M salt, hybridization temperature 5°C–10°C lower than Tm; moderate strictness, approximately 0.16 M to approximately 0.33 M salt, hybridization temperature 20°C–29°C lower than Tm; and low strictness, approximately 0.33 M to approximately 0.82 M salt, hybridization temperature 40°C–48°C lower than Tm. Tm of double-stranded nucleic acid sequences can be calculated using standard methods well-known in the art (see, for example, Maniatis, T., et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press: New York (1982); Casey, J., et al., Nucleic Acids Res. 4: 1539-1552 (1977); Bodkin, DK, et al., J. Virological Methods 10: 45-52 (1985); Wallace, RB, et al., Nucleic Acids Res. 9: 879-894 (1981)). Algorithmic prediction tools for Tm estimation are also widely available. High hybridization stringency generally refers to conditions where polynucleotides complementary to the target sequence hybridize primarily with the target sequence and substantially with little hybridization with non-target sequences. Typically, hybridization conditions are of moderate stringency, with high stringency being preferred.
[0108] As used herein, “complementarity” refers to the ability of a nucleic acid sequence to form hydrogen bonds with another nucleic acid sequence (e.g., via canonical Watson-Crick base pairing). Percentage complementarity indicates the percentage of residues in a nucleic acid sequence that can form hydrogen bonds with the second nucleic acid sequence. If two nucleic acid sequences have 100% complementarity, then the two sequences are perfectly complementary, meaning that all consecutive residues in the first polynucleotide are hydrogen-bonded to the same number of consecutive residues in the second polynucleotide.
[0109] As used herein, “binding” refers to a non-covalent interaction between macromolecules (e.g., between a protein and a polynucleotide, between polynucleotides, between proteins, etc.). This non-covalent interaction is also referred to as “binding” or “interaction” (e.g., if a first macromolecule interacts with a second macromolecule, the first macromolecule binds to the second macromolecule in a non-covalent manner). Some parts of a binding interaction can be sequence-specific (the terms “sequence-specific binding,” “sequence-specific binding,” “site-specific binding,” and “site-specific binding” are used interchangeably herein). As used herein, sequence-specific binding generally refers to the ability to form a complex with a type I CRISPR-Cas subunit protein (e.g., Cse2, Cas5, Cas6, Cas7, and Cas8), preferably with respect to a second nucleic acid sequence (e.g., a second DNA sequence) that does not have a nucleic acid target binding sequence (e.g., a DNA target binding sequence), resulting in the protein binding to one or more guide polynucleotides containing a nucleic acid target sequence (e.g., a DNA target sequence). Not all components of a binding interaction need to be sequence-specific, such as protein contact with phosphate residues in the DNA backbone. The characteristics of binding interactions can be found in the dissociation constant (Kd). "Binding affinity" refers to the strength of the binding interaction. Increased binding affinity is associated with a lower Kd.
[0110] As used in this article, an effector complex is described as “targeting” a polynucleotide if such a complex binds to or cleaves a polynucleotide within a nucleic acid target sequence.
[0111] As used in this article, a "double-strand break" (DSB) refers to the cleavage of both strands of a double-stranded DNA fragment. In some cases, if such a break occurs, one strand may have a "sticky end," where nucleotides are exposed and do not bind to hydrogen bonds with nucleotides on the other strand. In other cases, a "blunt end" may occur, where both strands remain perfectly base-paired with each other.
[0112] The terms “donor polynucleotide,” “donor oligonucleotide,” and “donor template” are used interchangeably herein and can be double-stranded polynucleotides (e.g., DNA), single-stranded polynucleotides (e.g., DNA or RNA), or combinations thereof. The donor polynucleotide may be contained in homologous arms flanking the inserted sequence (e.g., DSBs in DNA). The length of the homologous arms on each side can vary (e.g., 1–50 bases, 50–100 bases, 100–200 bases, 200–300 bases, 300–500 bases, 500–1000 bases). The homologous arms can be symmetrical or asymmetrical in length. The parameters used for designing or constructing donor polynucleotides are well known in the art (see, for example, Ran, F., et al., Nature Protocols 8:2281-2308 (2013); Smithies, O., et al., Nature 317:230-234 (1985); Thomas, K., et al., Cell 44:419-428 (1986); Wu, S., et al., Nature Protocols 3:1056-1076 (2008); Singer, B., et al., Cell 31:25-33 (1982); Shen, P., et al., Genetics 112:441-457 (1986); Watt, V., et al., Proc. Natl. Acad. Sci. USA). 82:4768-4772 (1985); Sugawara, N., et al., J. Mol. Bio. 12:563-575 (1992); Rubnitz, J., et al., J. Mol. Bio. 4:2253-2258 (1984); Ayares, D., et al., Proc. Natl. Acad. Sci. USA 83:5199-5203 (1986); Liskay, R., et al., Genetics 115:161-167 (1987)). In some embodiments, the donor polynucleotide comprises a chimeric antigen receptor (e.g., CAR).
[0113] The terms “chimeric antigen receptor” and “CAR” are used interchangeably in this document and refer to a polypeptide molecule typically containing at least two components, created in the laboratory: an extracellular antigen recognition domain (also known as a target-binding domain or extracellular ligand-binding domain) and an intracellular activation domain (e.g., containing one or more intracellular signaling domains and typically containing one or more costimulatory signaling domains). CARs may also contain a hinge domain and a transmembrane domain. A typical CAR polypeptide structure is as follows: N-terminus-extracellular-[antigen recognition domain-hinge domain]-transmembrane-[transmembrane domain]-intracellular-[intracellular activation domain]-C-terminus; or N-terminus-intracellular-[intracellular activation domain]-transmembrane-[transmembrane domain]-extracellular-[antigen recognition domain-hinge domain]-C-terminus.
[0114] Examples of extracellular antigen recognition domains include portions for binding to antigens, and include, but are not limited to, single-chain immunoglobulin variable fragments (scFv), antigen-binding fragments (Fab; typically the antibody region that binds to the antigen and consists of a constant domain and a variable domain for each heavy and light chain), nanobodies, single-chain antibodies from the camel family or sharks, modified protein-binding scaffolds (e.g., DARPins and Centyrins), or natural ligands that bind to their homologous receptors.
[0115] Examples of hinge domains include, but are not limited to, peptide hinges of variable length (e.g., one or more amino acids), hinge regions of CD8α, hinge regions of CD28, hinge regions of IgG4, and combinations thereof.
[0116] Examples of transmembrane domains include, but are not limited to, transmembrane regions derived from transmembrane proteins such as CD8α, CD28, DAP10, DAP12, NKG2D, and combinations thereof.
[0117] Examples of intracellular activation domains include, but are not limited to, intracellular signaling domains of CD28, 4-1BB, CD3ζ, OX40, 2B4, DAP10, DAP12, truncated and mutated signaling domains (e.g., mutations and truncations in the three ITAM domains of CD3ζ), or other intracellular signaling domains, and combinations thereof.
[0118] When the extracellular ligand-binding domain binds to a homologous ligand, the extracellular signaling domain of the CAR activates lymphocytes (for descriptions of CAR-T cells, see, for example, Brudno, J., et al., Nature Rev. Clin. Oncol. 15:31-46 (2018); Maude, S., et al., N. Engl. J. Med. 371:1507-1517 (2014); Sadelain, M., et al., Cancer Disc. 3:388-398 (2013); U.S. Patent No. 7,446,190; U.S. Patent No. 8,399,645) (for descriptions of CAR-NK cells, see, for example, Rezvani, K., et al., Mol. Ther., 25:1769-1781 (2017); Siegler, E., et al., Cell Stem Cell.23:160-161(2018); Li,Y.,et al.,Cell StemCell.23:181-192(2018); Lin,C.,et al.,Biochim.Biophys.Acta.Rev.Cancer.1869:200-215(2018); Hu,Y.,et al., Acta. Pharmacol. Sin. 39: 167-176 (2018); Fang, F., et al., Semin. Immunol. 31: 37-54 (2017); Glienke, W., et al., Front Pharmacol. 6: 21 (2015)).
[0119] Table 2 presents exemplary cellular targets and scFvs / binding proteins that bind to these targets. Such scFvs / binding proteins, or portions thereof, can be incorporated into CAR constructs.
[0120]
[0121]
[0122]
[0123] As used herein, “homology-directed repair” (HDR) refers to DNA repair occurring within a cell, such as during DSB repair in gDNA. HDR requires nucleotide sequence homology and uses a donor or template polynucleotide to repair the sequence, where a DSB occurs (e.g., in a DNA target sequence). The donor polynucleotide typically has the required sequence homology with the sequences flanking the DSB, allowing it to be used as a suitable template for repair. HDR results in the transfer of genetic information from, for example, the donor polynucleotide to the DNA target sequence. If the donor polynucleotide sequence differs from the DNA target sequence, and part or all of the donor polynucleotide is incorporated into the DNA target sequence, HDR can cause alterations to the DNA target sequence (e.g., insertions, deletions, or mutations). In some embodiments, the entire donor polynucleotide, a portion of the donor polynucleotide, or a copy of the donor polynucleotide is integrated at a site in the DNA target sequence. For example, the donor polynucleotide can be used for the repair of breaks in the DNA target sequence, where the repair results in the transfer of genetic information from or immediately adjacent to the break site in the DNA from the donor polynucleotide. Therefore, new genetic information can be inserted or copied at the DNA target sequence.
[0124] A "genomic region" is a chromosomal segment of the host cell's genome located on either side of a nucleic acid target sequence site, or optionally including a portion of the nucleic acid target sequence site. The homologous arms of the donor polynucleotide possess sufficient homology to undergo homologous recombination with the corresponding genomic region. In some implementations, the homologous arms of the donor polynucleotide exhibit significant sequence homology with genomic regions immediately flanking the nucleic acid target sequence site; it is generally accepted that homologous arms can be designed to have sufficient homology with genomic regions distant from the nucleic acid target sequence site.
[0125] As used in this article, "non-homologous end joining" (NHEJ) refers to repairing DSBs in DNA by directly joining one broken end to the other, without the need for a donor polynucleotide. NHEJ is a DNA repair pathway that can be used to repair DNA in cells without the use of a repair template. In the absence of a donor polynucleotide, NHEJ typically results in random insertions or deletions of nucleotides at the DSB site.
[0126] Microhomology-mediated end joining (MMEJ) is a pathway for repairing DSBs in gDNA. MMEJ involves deletions flanking the DSB and alignment of microhomological sequences within the break site prior to joining. MMEJ is gene-defined and requires activities such as CtIP, poly(ADP-ribose) polymerase 1 (PARP1), DNA polymerase θ (Polθ), DNA ligase 1 (Lig 1), or DNA ligase 3 (Lig 3). Additional genetic components are known in the art (see, for example, Sfeir, A., et al., Trends in Biochemical Sciences 40:701-714 (2015)).
[0127] As used herein, “DNA repair” includes any process by which cellular mechanisms repair damage to DNA molecules contained within a cell. The repaired damage can include single-strand breaks or DSBs. At least three mechanisms exist for repairing DSBs: HDR, NHEJ, and MMEJ. “DNA repair” is also used herein to refer to DNA repair produced by human or machine manipulation, where the target site is modified, for example, by inserting, deleting, or substituting nucleotides; all of these represent forms of genome editing.
[0128] As used in this article, "recombination" refers to the process of exchanging genetic information between two polynucleotides.
[0129] As used herein, the terms “regulatory sequence,” “regulatory element,” and “control element” are interchangeable and refer to the polynucleotide sequence to be expressed upstream (5' non-coding sequence), internal, or downstream (3' non-translated sequence) of a polynucleotide target. Regulatory sequences influence, for example, the timing of transcription; the amount or level of transcription; RNA processing or stability; and / or the translation of related structural nucleotide sequences. Regulatory sequences may include activator-binding sequences, enhancers, introns, polyadenylation recognition sequences, promoters, transcription start sites, repressor-binding sequences, stem-loop structures, translation initiation sequences, internal ribosome entry sites (IRES), translation leader sequences, transcription termination sequences (e.g., polyadenylation signals and poly-U sequences), translation termination sequences, primer binding sites, etc.
[0130] Regulatory elements include those that direct the constitutive, inducible, and repressible expression of nucleotide sequences in many types of host cells, as well as those that direct the expression of nucleotide sequences only in certain host cells (e.g., tissue-specific regulatory sequences). In some embodiments, the vector comprises one or more pol III promoters, one or more pol II promoters, one or more pol I promoters, or combinations thereof. Examples of pol III promoters include, but are not limited to, the U6 and H1 promoters. Examples of pol II promoters include, but are not limited to, the retroviral Ruis sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer; see, for example, Boshart, M., et al., Cell 41:521-530 (1985)), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the glycerol phosphokinase (PGK) promoter, and the EF1α promoter, as well as modified artificial promoters (e.g., the MND promoter and the CAG promoter). Those skilled in the art will understand that the design of expression vectors can depend on factors such as the selection of host cells to be transformed and the desired expression level. Vectors can be introduced into host cells to produce RNA transcripts, proteins, or peptides encoded by nucleic acid sequences, including fusion proteins or peptides, as described herein.
[0131] As used in this article, "gene" refers to a multinucleotide sequence containing exons and associated regulatory sequences. Genes may also contain introns and / or untranslated regions (UTRs).
[0132] As used herein, the term "operably linked" refers to polynucleotide or amino acid sequences that are functionally related to each other. For example, if a regulatory sequence regulates or contributes to the transcriptional regulation of a polynucleotide, then the regulatory sequence (e.g., a promoter or enhancer) is "operably linked" to the polynucleotide encoding the gene product. Operatively linked regulatory elements are typically adjacent to the coding sequence. However, enhancers can function if they are no more than a few thousand bases or more away from the promoter. Additionally, polycistronic constructs can include multiple coding sequences, using only one promoter by including 2A self-cleaving peptides, IRES elements, etc. Thus, some regulatory elements can be operably linked to polynucleotide sequences but are not adjacent to them. Similarly, translational regulatory elements contribute to the regulation of protein expression from polynucleotides.
[0133] As used herein, “expression” refers to the transcription of polynucleotides from a DNA template to produce, for example, messenger RNA (mRNA) or other RNA transcripts (e.g., non-coding, such as structural or scaffold RNAs). The term also refers to the process by which transcribed mRNA is translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as “gene products.” If the polynucleotide is derived from gDNA, expression in eukaryotic cells may involve splicing mRNA.
[0134] A "coding sequence," or sequence that "encodes" a selected polypeptide, is a nucleic acid molecule that, when placed under the control of appropriate regulatory sequences, is transcribed (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vitro or in vivo. The boundaries of the coding sequence are determined by the start codon at the 5' end and the translation stop codon at the 3' end.
[0135] As used herein, “artificial transcription activator (ATA)” or “artificial transcription factor (ATF)” refers to a complex capable of recruiting the RNA polymerase II holoenzyme to its associated gene, thereby causing ectopic expression of the target gene. Such an activator comprises at least two components: (1) a catalytically inactivating polynucleotide-binding domain that directly recognizes and binds to homologous nucleotide sequences, or a polynucleotide-binding domain that directs such sequences for binding (e.g., a nucleoprotein complex comprising a nucleic acid-binding domain and a guide as described herein); and (2) an activation domain (also called an “effect domain”) that interacts with the various proteins that constitute the transcriptional mechanisms that upregulate transcription.
[0136] A "catalytically inactivated polynucleotide binding domain" refers to a molecule that binds to but does not cleave the nucleic acid target site bound by the binding domain. Representative examples of this type of domain are described in detail in this article.
[0137] As used herein, the term "regulation" refers to a change in the number, degree, or amount of function. For example, the type I CRISPR nucleoprotein complex disclosed herein can regulate the activity of a promoter sequence by binding to a nucleic acid target sequence at or near the promoter or transcription start site or regulatory site. Depending on the effect that occurs after binding, the type I CRISPR nucleoprotein complex can induce, enhance, suppress, or inhibit the transcription of genes operatively linked to the promoter sequence. Therefore, "regulation" of gene expression includes both gene activation and gene repression.
[0138] Regulation can be analyzed by identifying any characteristics that directly or indirectly affect the expression of target genes. These characteristics include, for example, changes in RNA or protein levels, protein activity, product levels, gene expression, or reporter gene activity. Therefore, the terms “regulating gene expression,” “repressing gene expression,” and “activating gene expression” can refer to the ability of type I CRISPR nucleoprotein complexes to alter, activate, or repress gene transcription.
[0139] Functions (e.g., enzymatic functions) can be upregulated (e.g., increased, enhanced, amplified, or strengthened) or downregulated (e.g., decreased, weakened, reduced, or diminished). In one embodiment, the binding of mCas3 protein to single-stranded DNA (ssDNA) or via ATP binding / hydrolysis of mCas3 protein can be upregulated or downregulated relative to the corresponding wtCas3 protein.
[0140] As used herein, “vector” and “plasmid” refer to polynucleotide vectors that introduce genetic material into cells. Vectors can be linear or circular. Vectors may contain a replication sequence that enables the vector to replicate in a suitable host cell (e.g., an origin of replication). After transformation into a suitable host, the vector may replicate and function independently of the host genome or integrate into the host genome. Among other things, vector design depends on the intended use of the vector and the host cell, and the design of the vectors of the present invention for a specific use and host cell is within the scope of the art. The four main types of vectors are plasmids, viral vectors, granules, and artificial chromosomes. Typically, vectors contain an origin of replication, a multiple cloning site, and / or optional markers. Expression vectors typically contain an expression cassette. By “recombinant virus”, it means a virus that has been genetically altered, for example, by adding or inserting a heterologous nucleic acid construct into or into the viral genome or a portion thereof.
[0141] As used herein, an "expression cassette" refers to a polynucleotide construct produced using recombinant methods or synthetic means, containing a regulatory sequence operatively linked to a selected polynucleotide to promote its expression in a host cell. For example, the regulatory sequence may promote the transcription of the selected polynucleotide in a host cell, or the transcription and translation of the selected polynucleotide in a host cell. Expression cassettes may be integrated into the genome of a host cell or exist within a vector to form an expression vector.
[0142] As used herein, a “targeting vector” is a recombinant DNA construct that typically contains specially designed DNA arms homologous to gDNA, located flanking a target gene or nucleic acid target sequence (e.g., a DSB). The targeting vector contains a donor polynucleotide. Elements of the target gene can be modified in various ways, including deletion and / or insertion. A defective target gene can be replaced with a functional target gene, or optionally, a functional gene can be knocked out. The donor polynucleotide of the targeting vector contains a selection cassette containing selectable markers introduced into the target gene. Target regions (containing nucleic acid target sequences) adjacent to or within the target gene can be used to regulate gene expression.
[0143] As used herein, the term “between” includes endpoint values within a given range (e.g., a length between 1 and 50 nucleotides includes 1 nucleotide and 50 nucleotides; a length between 5 and 50 amino acids includes 5 amino acids and 50 amino acids).
[0144] As used herein, the term “amino acid” (aa) refers to natural and synthetic (non-natural) amino acids, including amino acid analogs, modified amino acids, peptide mimics, glycine, and D or L optical isomers.
[0145] As used herein, the terms “peptide,” “polypeptide,” “protein,” and “subunit protein” are interchangeable and refer to polymers of amino acids. Polypeptides can have any length. They can be branched or linear, can contain non-amino acids, and can include modified amino acids. The term also refers to amino acid polymers that have been modified by, for example, acetylation, disulfide bond formation, glycosylation, lipoylation, phosphorylation, PEGylation, biotinylation, crosslinking, and / or conjugation (e.g., using labeled components or ligands). Unless otherwise stated, polypeptide sequences in this document are shown in the conventional N-terminal to C-terminal orientation.
[0146] Peptides and polynucleotides can be prepared using conventional techniques in the field of molecular biology (see, for example, the standard text listed above). Furthermore, virtually any peptide or polynucleotide is available from commercial sources.
[0147] As used herein, the terms "fusion protein" and "chimeric protein" refer to a single protein generated by linking two or more proteins, protein domains, protein fragments, or circularly arranged polypeptides that are not naturally present together in a single protein. In some embodiments, linker polynucleotides may be used to link a first protein, protein domain, or protein fragment or circularly arranged polypeptide to a second protein, protein domain, protein fragment, or circularly arranged polypeptide. For example, a fusion protein may comprise a type I CRISPR-Cas protein (e.g., Cas8, Cas3) and a domain from another protein (e.g., FokI; see, for example, U.S. Patent No. 9,885,026). Modifying a fusion protein to include such a domain can impart additional activity to the modified type I CRISPR-Cas protein. Such activities may include nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, superoxide dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, and / or myristylation activity or demyristylation activity for modifying polypeptides associated with nucleic acid target sequences (e.g., histones).
[0148] In some implementations, the fusion protein may include an epitope tag (e.g., a histidine tag, an HA tag, etc.). (Sigma Aldrich, St. Louis, MO) tags, Myc tags, nuclear localization signal (NLS) tags, SunTag, reporter protein sequences (e.g., glutathione S-transferase, β-galactosidase, luciferase, green fluorescent protein, cyan fluorescent protein, yellow fluorescent protein), and / or nucleic acid sequence binding domains (e.g., DNA binding domain or RNA binding domain).
[0149] Fusion proteins may also contain an activator domain (e.g., the activator of the heat shock transcription factor NFKB) or an inhibitor domain (e.g., the KRAB domain). As described in Lupo, A., et al., Current Genomics 14:268-278 (2013), the KRAB domain is a potent transcriptional repression module and is located in the N-terminal sequence of most C2H2 zinc finger proteins (see, for example, Margolin, J., et al., Proc. Natl. Acad. Sci. USA 91:4509-4513 (1994); Witzgall, R., et al., Proc. Natl. Acad. Sci. USA 91:4514-4518 (1994)). The KRAB domain typically binds to co-repressor proteins and / or transcription factors via protein-protein interactions, causing transcriptional repression of genes that bind to KRAB zinc finger proteins (KRAB-ZFPs) (see, for example, Friedman, JR, et al., Genes & Development 10:2067-2678 (1996)). In some embodiments, the linker nucleic acid sequence is used to link two or more proteins, protein domains, or protein fragments.
[0150] As used herein, “CASCADEa” (Cascade activation) is a CRISPR method or system that activates the expression of a gene associated with a site on a target nucleic acid sequence of the Cascade RNP complex. In some embodiments, one or more proteins of the Cascade complex are fused to an effector domain (e.g., VP16 or VP64), and the Cascade RNP complex containing the fusion and guide polynucleotide is used to recruit endogenous transcription factors. In some embodiments, the guide polynucleotide may be fused at 5' or 3' to a nucleotide effector domain, such as MS2-binding RNA that also recruits transcription factors.
[0151] As used herein, “CASCADEi” (Cascade repression) is a CRISPR method or system in which the expression of a gene associated with a site on the target nucleic acid sequence of the Cascade RNP complex is downregulated (i.e., the Cascade RNP complex is used to downregulate gene expression). For the recruitment of endogenous repressor factors, one or more proteins in the Cascade complex are typically fused to an effector domain (e.g., KRAB). In some embodiments, the guide polynucleotide may be fused at 5' or 3' to a nucleotide effector domain that also recruits endogenous transcriptional repressor effector proteins.
[0152] As used herein, “part” refers to a portion of a molecule. A part can be a functional group or can describe a portion of a molecule having multiple functional groups (e.g., sharing common structural aspects). The terms “part” and “functional group” are generally used interchangeably herein; however, “functional group” can more specifically refer to a portion of a molecule that includes some common chemical behaviors. “Part” is generally used as a structural description. In some embodiments, the 5' end, 3' end, or 5' and 3' ends (e.g., a non-natural 5' end and / or a non-natural 3' end in a first stem element) may comprise one or more parts.
[0153] As used herein, “adoptive cells” refers to cells that can be genetically modified for use in cell therapy, such as for the treatment of cancer and / or the prevention of graft-versus-host disease (GvHD), as well as other adverse side effects of cell therapy, such as, but not limited to, cytokine storms, oncogenic transformation of the administered genetically modified material, and neurological disorders. Adoptive cells include, but are not limited to, stem cells, induced pluripotent stem cells (iPSCs), umbilical cord blood stem cells, lymphocytes, macrophages, erythrocytes, fibroblasts, endothelial cells, epithelial cells, and pancreatic progenitor cells.
[0154] As used in this article, “cell therapy” refers to the treatment of diseases or conditions using genetically modified cells. Genetic modifications can be introduced using the methods described herein, including methods such as viral vectors, nuclear transfection, gene gun delivery, sonication, cell extrusion, lipid transfection, or the use of other chemicals, cell-penetrating peptides, etc.
[0155] As used in this article, “adoptive cell therapy (ACT)” refers to a therapy that treats a patient using genetically modified adoptive cells derived from a specific patient (autologous cell therapy) or a third-party donor (allogeneic cell therapy). ACTs include, but are not limited to, bone marrow transplantation, stem cell transplantation, T-cell therapy, CAR-T cell therapy, and natural killer (NK) cell therapy.
[0156] As used herein, “lymphocyte” refers to white blood cells (leukocytes) that are part of the vertebrate immune system. The term “lymphocyte” also includes hematopoietic stem cells or induced pluripotent stem cells (iPSCs) that produce lymphoid cells. Lymphocytes include T cells used for cell-mediated cytotoxic adaptive immunity, such as CD4+ and / or CD8+ cytotoxic T cells; α / β T cells and γ / δ T cells; regulatory T cells, such as Treg cells; NK cells that play a role in cell-mediated cytotoxic innate immunity; B cells used for humoral, antibody-driven adaptive immunity; NK / T cells; cytokine-induced killer cells (CIK cells); and antigen-presenting cells (APCs), such as dendritic cells. Lymphocytes can be mammalian cells, such as human (Homo sapiens; H. sapiens) cells. The term “lymphocyte” also includes genetically modified T cells and NK cells that are modified to produce chimeric antigen receptors (CARs) on the surface of T or NK cells (CAR-T cells and CAR-NK cells). These CAR-T cells recognize specific soluble antigens or target cell surfaces, such as the surface of tumor cells or antigens on cells in the tumor microenvironment.
[0157] As used herein, the term "lymphocyte" also includes T-cell receptor-modified T cells (TCRs), which are genetically modified to express one or more specific naturally occurring or modified T-cell receptors that can recognize protein or (glyco)lipid antigens of target cells presented by the major histocompatibility complex (MHC). Small fragments of these antigens, such as peptides or fatty acids, are shuttled to the surface of target cells and presented to T-cell receptors, which are part of the MHC. The binding of T-cell receptors to antigen-carrying MHCs activates lymphocytes.
[0158] Lymphocyte activation occurs when lymphocytes are triggered by antigen-specific receptors on their cell surface. This leads to cell proliferation and differentiation into specific effector lymphocytes. These "activated" lymphocytes are typically characterized by a set of receptors on their cell surface. Surface markers of activated T cells include CD3, CD4, CD8, PD1, and IL2R. Activated cytotoxic lymphocytes can kill target cells after binding to their homologous receptors on the surface of target cells.
[0159] As used herein, the term "lymphocyte" also includes tumor-infiltrating lymphocytes (TILs). TILs are immune cells that have penetrated the tumor and its surrounding environment ("tumor microenvironment"). TILs are typically isolated from tumor cells and the tumor microenvironment and selected in vitro for high reactivity against tumor antigens. TILs are grown in vitro under conditions that overcome the effects of tolerance present in vivo and then introduced into the patient for treatment.
[0160] T cells typically present in many subtypes, such as "naive T cells" (Tn), "stem cell memory T cells" (Tscm), "central memory T cells" (Tcm), "effector memory T cells" (Tem), "effector T cells" (Teff), and "regulatory T cells" (Treg). Each T cell subgroup is characterized by a set of cell surface markers.
[0161] As used herein, the term "affinity tag" generally refers to one or more portions that increase the binding affinity of one macromolecule to another, for example, to facilitate the formation of a modified type I CRISPR-Cas nucleoprotein complex. In some embodiments, an affinity tag can be used to increase the binding affinity of one Cas subunit protein to another Cas subunit protein (e.g., a first Cas7 protein to a second Cas7 protein). In some embodiments, an affinity tag can be used to increase the binding affinity of one or more Cas subunit proteins to homologous guide polynucleotides. Some embodiments of the invention introduce one or more affinity tags at the N-terminus of a Cas subunit protein sequence, at the C-terminus of a Cas subunit protein sequence, at a position between the N-terminus and C-terminus of a Cas subunit protein sequence, or a combination thereof. In some embodiments of the invention, one or more guide polynucleotides contain an affinity tag that increases the binding affinity of the guide polynucleotide to one or more Cas subunit proteins. A wide variety of affinity tags are disclosed in U.S. Patent Application Publication No. 2014-0315985, published October 23, 2014. The ligand and the ligand binding site are paired affinity tags.
[0162] As used herein, "crosslinking" is a bond that links one polymer chain (e.g., a polynucleotide or polypeptide) to another. Such a bond can be a covalent bond or an ionic bond. In some embodiments, a polynucleotide can be linked to another polynucleotide by crosslinking the polynucleotide. In other embodiments, a polynucleotide can be crosslinked to a polypeptide. In still other embodiments, a polypeptide can be crosslinked to another polypeptide.
[0163] As used herein, the term "crosslinked portion" generally refers to a portion suitable for providing crosslinking between two macromolecules. Crosslinked portions are another example of affinity tags.
[0164] As used in this article, "host cell" generally refers to a biological cell. A cell is the basic structural, functional, and / or biological unit of a living organism. Cells can originate from any organism that has one or more cells. Examples of host cells include, but are not limited to, prokaryotic cells, eukaryotic cells, bacterial cells, archaea cells, cells of unicellular eukaryotes, cells of eukaryotes, protozoan cells, cells from plants, algal cells (e.g., *Botryococcus braunii*, *Chlamydomonas reinhardtii*, *Nanochloropsis gaditana*, *Chlorella pyrenoidosa*, *Sargassum patens* C. agardh), seaweed (e.g., giant kelp), fungal cells (e.g., yeast cells or cells from mushrooms), animal cells, cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), and cells from vertebrates, including mammals (e.g., pigs, cattle, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). Furthermore, host cells can be stem cells or progenitor cells, as well as immune cells, such as any immune cells described herein. The host cell can be a human cell. In some embodiments, the human cell is located outside the human body. In some embodiments, somatic cells of a living organism (e.g., the human body) are manipulated in vitro (i.e., outside the living body). Ex vivo generally refers to a medical procedure in which organs, cells, or tissues are taken from a living organism (e.g., the human body) for treatment or surgery and then returned to the living organism.
[0165] As used in this article, "stem cell" refers to a cell with the ability to self-renew, that is, the ability to undergo many cycles of cell division while remaining undifferentiated. Stem cells can be totipotent, pluripotent, oligopotent, or unipotent. Stem cells can be embryonic, fetal, amniotic, adult, or induced pluripotent stem cells.
[0166] As used herein, “induced pluripotent stem cells” refers to a class of pluripotent stem cells artificially derived from non-pluripotent cells, typically somatic cells. In some embodiments, the somatic cells are human somatic cells. Examples of somatic cells include, but are not limited to, dermal fibroblasts, bone marrow-derived mesenchymal cells, cardiomyocytes, keratinocytes, hepatocytes, gastric cells, neural stem cells, lung cells, kidney cells, spleen cells, and pancreatic cells. Other examples of somatic cells include cells of the immune system, including but not limited to B cells, dendritic cells, granulocytes, innate lymphoid cells, megakaryocytes, monocytes / macrophages, myeloid-derived suppressor cells, natural killer (NK) cells, T cells, thymocytes, and hematopoietic stem cells.
[0167] As used in this article, "hematopoietic stem cells" refers to undifferentiated cells that have the ability to differentiate into hematopoietic cells such as lymphocytes.
[0168] As used herein, “plant” means the whole plant, plant organ, plant tissue, germplasm, seed, plant cell, and its offspring. Plant cells include, but are not limited to, cells derived from seeds, suspension cultures, embryos, meristematic regions, callus, leaves, roots, branches, gametophytes, sporophytes, pollen, and microspores. Plant parts include differentiated and undifferentiated tissues, including but not limited to roots, stems, branches, leaves, pollen, seeds, tumor tissue, and various forms of cells and cultures (e.g., single cells, protoplasts, embryos, and callus). Plant tissues can be in the plant or in plant organs, tissues, or cell cultures. “Plant organ” means a plant tissue or group of tissues that constitutes a morphologically and functionally distinct part of the plant.
[0169] The terms “object,” “individual,” or “patient” are used interchangeably herein and refer to any member of the phylum Chordata, including but not limited to humans and other primates, including non-human primates such as rhesus monkeys, chimpanzees, and other monkey and ape species; farm animals such as cattle, sheep, pigs, goats, and horses; domesticated mammals such as dogs and cats; laboratory animals including rabbits, mice, rats, and guinea pigs; birds, including domesticated, wild, and playful birds such as chickens, turkeys, and other chicken-like birds, ducks, and geese; and so on. The term does not indicate a specific age or sex. Therefore, the term includes adult, young, and newborn individuals, as well as males and females. In some embodiments, the host cell is derived from the object (e.g., lymphocytes, stem cells, progenitor cells, or tissue-specific cells). In some embodiments, the object is a non-human object. In some embodiments, the object is a human (Homo sapiens) object.
[0170] The term "effective amount" or "therapeutic effective amount" for compositions or reagents, such as the genetically modified adoptive cells described herein, refers to an amount sufficient to provide a desired response, such as prevention or elimination of one or more adverse side effects associated with allogeneic adoptive cell therapy. Such a response will depend on the specific target disease. For example, in patients treated with adoptive cell therapy for cancer, the desired response includes, but is not limited to, treatment or prevention of GvHD, host resistance to graft rejection, cytokine release syndrome (CRS), cytokine storm, and reduction of the carcinogenic transformation effects of the administered genetically modified cells. The exact amount required will vary from subject to subject, depending on the type of subject, age and general condition, severity of the condition being treated, the specific modified lymphocytes used, the route of administration, etc. In any individual case, a person skilled in the art can determine the appropriate "effective" amount using routine experiments.
[0171] "Treatment" or "treating" a specific disease, such as a cancerous condition or GvHD, includes: (1) preventing the disease, for example, preventing the development of the disease or causing the disease to occur at a lower intensity in subjects who may be susceptible to the disease but have not yet experienced or shown symptoms of the disease; (2) suppressing the disease, for example, slowing the rate of development, stopping the development, or reversing the disease state; and / or (3) alleviating the symptoms of the disease, for example, reducing the number of symptoms experienced by the subject.
[0172] As used herein, “gene editing” or “genome editing” refers to a class of genetic engineering processes that result in modifications to genes, such as the insertion, deletion, or substitution of nucleotide sequences at specific sites in the cell’s genome, or even a single base. This term includes, but is not limited to, heterologous gene expression, gene or promoter insertion or deletion, nucleic acid mutation, and disruptive genetic modifications as defined herein.
[0173] An epitope is a specific site on a molecule where B cells and T cells respond. In the unique spatial configuration of an epitope, it can contain three or more amino acids. Typically, an epitope consists of at least five such amino acids, and more commonly, at least eight to ten. Methods for determining the spatial configuration of amino acids are known in the art and include, for example, X-ray crystallography, electron microscopy, and two-dimensional nuclear magnetic resonance. Furthermore, the identification of epitopes in a given protein can be readily accomplished using techniques well-known in the art, such as through hydrophobicity studies and site-directed serology.
[0174] "Mimetic epitopes" are large molecules, such as peptides, that mimic the structure of an epitope. Due to this property, they induce antibody responses similar to those triggered by the epitope itself. An antibody against a given epitope antigen will recognize the mimetic epitope that mimics that epitope. Mimetic epitopes are typically obtained from phage display libraries through biopanning.
[0175] "Antibody" refers to a "recognition," that is, a molecule that specifically binds to a target epitope presented in a polypeptide, such as a ligand-binding domain. Through "specific binding," antibodies interact with epitopes in a "lock and key" manner to form a complex between the antigen and the antibody. As used herein, the term "antibody" includes antibodies derived from monoclonal formulations, as well as the following: hybrid (chimeric) antibody molecules; F(ab')2 and F(ab) fragments; Fv molecules (non-covalent heterodimers); single-chain Fv molecules (scFv); dimer and trimer antibody fragment constructs; small antibodies; humanized antibody molecules; single-chain antibodies; nanobody (Ablynx NV, Zwijnaarde, Belgium) antibodies; and any functional fragments derived from these molecules, wherein such fragments retain the immunobinding properties of the parent antibody molecule. Antibodies can be derived from various species, such as humans, mice, rats, rabbits, camels, chickens, etc. Antibodies and antibody fractions can then be further obtained using in vitro techniques, such as phage display and yeast display. Fully humanized antibodies with modified humanized B cell libraries can be obtained from human plasma, human B cell clones, mice, rats, rabbits, chickens, etc. Antibodies can then be further modified using affinity maturation and other methods, such as fucosylation or IgG Fc engineering.
[0176] As used herein, the term "monoclonal antibody" refers to an antibody composition having a homogeneous antibody population. This term is not limited to the type or source of the antibody, nor is it intended to be limited by its preparation method. The term includes complete immunoglobulins as well as fragments such as Fab, F(ab')2, Fv, and other fragments, and chimeric and humanized homogeneous antibody populations, all of which exhibit the immunobinding properties of the parental monoclonal antibody molecules of the present invention.
[0177] Antibody-dependent cell-mediated cytotoxicity (ADCC), also known as antibody-dependent cell cytotoxicity, refers to the mechanism by which effector cells of the immune system actively lyse target cells, such as adoptive cells, when a ligand-binding domain on the cell membrane is bound by a specific antibody. Effector cells are typically natural killer (NK) cells. However, macrophages, neutrophils, and eosinophils can also mediate ADCC. ADCC is independent of complement-dependent cytotoxicity (CDC), which can also lyse targets by disrupting the membrane without the involvement of antibodies or immune system cells.
[0178] As used herein, “transformation” refers to the insertion of a foreign polynucleotide into a host cell, regardless of the method used for insertion. Transformation can occur, for example, through direct uptake, transfection, infection, etc. The foreign polynucleotide can remain as a non-integrating vector, such as an episome, or optionally, can be integrated into the host genome. As used herein, “transgenic organism” refers to an organism containing genetic material in which DNA from an unrelated organism has been artificially introduced. This term includes the offspring of the transgenic organism (any generation), provided that the offspring have genetic modifications. In some embodiments, the transgenic organism is a non-human transgenic organism.
[0179] As used herein, “isolated” can refer to a molecule (e.g., a polynucleotide or polypeptide) that exists outside its natural environment due to human intervention and is therefore not a natural product. When referring to polypeptides, “isolated” means that the molecule is separate from and discontinuous with the complete organism found in nature with the molecule, or appears in the absence of substantially any other biomacromolecule of the same type. The term “isolated” for polynucleotides refers to nucleic acid molecules that are wholly or partially lack the sequence normally associated with them in nature; or sequences that, when present in nature, have a heterologous sequence associated with them; or molecules isolated from chromosomes.
[0180] As used herein, the term "purified" preferably means having at least 75% by weight, more preferably at least 85% by weight, still more preferably at least 95% by weight, and most preferably at least 98% by weight of the same molecules.
[0181] As used herein, a “substrate channel” refers to the direct transfer of a reactant from one enzymatic reaction to another without first diffusing into the overall environment (see, for example, Wheeldon, I., et al., Nat. Chem. 8:299-309 (2016)). These intermediates of enzymatic steps are out of equilibrium with the overall solution, which allows for increased efficiency and yield in the enzymatic process. Typically, in naturally occurring metabolic processes, enzymes have developed ways to co-locate and assemble into controlled aggregates.
[0182] As used herein, a "substrate channel element" refers to a component of a metabolic pathway. In some embodiments, a substrate channel element is an enzyme that catalyzes a chemical reaction.
[0183] As used in this article, a “substrate channel complex” refers to multiple substrate channel elements that are co-located together in some manner.
[0184] As used in this article, "RNA scaffold" refers to an RNA molecule in which peptides can be used as binding substrates.
[0185] The data presented in this paper demonstrate that the fusion between Cascade components and nuclease domains (e.g., dimerization-dependent nonspecific FokI nuclease domains; see, for example, Urnov, FD, et al., Nature Reviews Genetics 11:636–646 (2010); Joung, JK, et al., Nat. Rev. Mol. Cell Biol. 14:49–55 (2013); Guilinger, JP, et al., Nat. Biotechnol. 32:577–582 (2014); Tsai, SQ, et al., Nat. Biotechnol. 32:569–576 (2014)) mediates efficient programmable RNA-guided gene editing in human cells using type I systems. Data demonstrate that modified type I CRISPR-Cas systems (e.g., including FokI-Cascade component fusions) can be directly transfected as complete ribonucleoprotein (RNP) complexes or assembled in cells via delivery of components encoded by a single plasmid. As described herein, all CRISPR-related (Cas) genes were assembled onto a single polycistronic vector, resulting in a simplified two-component Cas protein-guided RNA expression system. Furthermore, the length / composition design of the nuclease (e.g., FokI) / Cascade component linker sequence, appropriate DNA geometry settings, and selective Cascade homolog selection provided modified type I CRISPR-Cas complexes with editing efficiencies of up to approximately 50%. Key characteristics of modified type I CRISPR-Cas systems (e.g., including FokI–Cascade component fusion proteins) involving PAM requirements and mismatch susceptibility during DNA targeting were identified.
[0186] In a first aspect, the present invention relates to modified polynucleotides encoding Cascade components, including but not limited to Cascade subunit proteins and Cascade guide polynucleotides.
[0187] In one embodiment, the present invention relates to modified polynucleotides encoding Cascade components derived from the Cascade IE type system. Example 1 presents an exemplary polynucleotide construct comprising Cascade proteins and Cascade crRNAs. Example 1, Table 15, and SEQ ID NO:1 to SEQ ID NO:20 present polynucleotide DNA sequences encoding genes encoding five subunit proteins of IE type Cascade specifically derived from *E. coli* strain K-12MG1655, and amino acid sequences of the resulting protein components. The polynucleotide sequences are derived from *E. coli* gDNA and are specifically codon-optimized for expression in *E. coli*, and / or specifically for expression in eukaryotic cells (e.g., human cells). When this polynucleotide is transcribed into precursor crRNA and treated with a Cascade RNA endonuclease, mature crRNA is generated to act as a guide RNA to target complementary DNA sequences in the genome. The minimal CRISPR array comprises two repeat sequences (underlined in the CRISPR array sequences presented in Example 1) flanking an exemplary spacer region sequence, representing the guide portion of the crRNA. RNA processing via the Cascade endonuclease produces crRNA with repetitive sequences flanking the guide sequence at the 5' and 3' ends. Those skilled in the art, referring to the teachings of this specification and the examples, can select appropriate spacer sequences to target the binding of the Cascade complex to a chosen target sequence (e.g., in gDNA).
[0188] Following the instructions in this manual, and using bioinformatics tools such as BLAST and PSI-BLAST, polynucleotide sequences encoding Cascade components from other bacterial and archaeal species can be identified and designed to locate homologues of the Cascade subunit gene, for example, from *E. coli* strain K-12MG1655. Then, examining the genomic proximity flanking the Cascade gene allows for the location and identification of genes encoding the remaining Cascade subunit proteins (see, for example, Examples 14A, 14B, 15A, and 15B). Because Cascade genes coexist as conserved operons, they are typically arranged in a consistent order within the same type I subtype, facilitating their identification and selection for subsequent analysis and experiments. For example, promising bacterial species can be identified by locating Cas8 homologues for homologous Cascade testing, and other type IE systems can be identified by obtaining or designing polynucleotide sequences encoding other protein components of Cas8 and Cascade from those homologous CRISPR-Cas systems.
[0189] The polynucleotide DNA sequences encoding the subunit proteins of Cascade from many species (some having Cascade complexes homologous to those from Escherichia coli strain K-12MG1655) (listed in Tables 3 and 4), the amino acid sequences of the resulting protein components, and exemplary minimal CRISPR arrays are presented as SEQ ID NO:22 to SEQ ID NO:213 (Table 3).
[0190] Table 3 shows the polynucleotide sequences of genes encoding Cascade proteins from 12 species.
[0191]
[0192]
[0193]
[0194]
[0195]
[0196]
[0197]
[0198]
[0199]
[0200]
[0201] The polynucleotide sequences of the proteins are derived from the gDNA of the host bacteria and are codon-optimized specifically for expression in *E. coli*, and / or for expression in eukaryotic cells (e.g., human cells). The polynucleotide DNA sequences encoding the corresponding minimal CRISPR arrays are based on repetitive sequences from 12 species and can be used to generate mature crRNAs used as guide RNAs. In Table 4, the minimal CRISPR arrays contain two repetitive sequences (lowercase, with underscores) flanking the exemplary “spacer” sequence, representing the guide portion of the crRNA. Processing the RNA with the Cascade subunit of the endonuclease produces crRNAs with repetitive sequences at both the 5' and 3' ends flanking the guide sequence.
[0202]
[0203]
[0204] In another embodiment, the present invention relates to modified polynucleotide sequences encoding cascade components from other bacterial or archaeal species having other type I phenotypes; including but not limited to variants of IB, IC, IF, and IF, which can be identified and designed according to the guidance of this specification and by using bioinformatics tools such as BLAST and PSI-BLAST to locate homologues of cascade genes from marker systems representing each subtype (see, for example, Makarova, KS, et al., Nat. Rev. Microbiol. 13:722-736 (2015); Koonin, EV, et al., Curr. Opin. Microbiol. 37:67-78 (2017)). After identifying the desired homologues, the genomic neighborhoods flanking the cascade gene can be examined to locate and identify genes for the remaining cascade subunit proteins disclosed herein. For example, other IF-type variant 2 systems can be identified by locating Cas8 homologues (and by locating Cas5 homologues) and identifying promising bacterial species for homologous Cascade testing, and then obtaining or designing polynucleotide sequences encoding Cas8, Cas5 and other protein components of Cascade from those homologous CRISPR-Cas systems.
[0205] SEQ ID NO:214 to SEQ ID NO:351 show polynucleotide DNA sequences encoding three, four, or five subunit proteins of Cascade from 12 other homologous Cascade complexes of type II (IB, IC, IF, and IF variants), amino acid sequences of the resulting protein components, and exemplary minimal CRISPR arrays (Table 3). The polynucleotide sequences of the subunit proteins are derived from the gDNA of the host bacteria and are codon-optimized specifically for expression in *E. coli* and / or for expression in eukaryotic cells (e.g., human cells). The polynucleotide DNA sequences encoding the corresponding minimal CRISPR arrays are based on repetitive sequences derived from the 12 species and can be used to generate mature crRNAs used as guide RNAs. In Table 5, the minimal CRISPR arrays contain two repetitive sequences (lowercase, underlined) flanking an exemplary “spacer” sequence, representing the guide portion of the crRNA. Processing the RNA with the Cascade subunit endonuclease produces crRNAs with repetitive sequences at both the 5' and 3' ends flanking the guide sequence.
[0206]
[0207]
[0208] Examples 19A to 19I and Examples 22A to 22C describe the design and testing of various Cascade complex homologues, each containing a Cas subunit protein-FokI fusion protein, to evaluate the genome editing efficiency of each Cascade complex. The highest editing was observed using a variant from *Pseudomonas* S-6-2, while other homologues (i.e., *Salmonella enterica*, *Geothermalobacterium* EPR-M, *Methanococcus insica* MRE50, and *Streptococcus thermophilus* (strain ND07)) showed editing approximately equal to that of *Escherichia coli*. Editing was also observed using modified *Vibrio cholerae* strain L15 (IF type) FokI-Cascade complex and *Vibrio cholerae* strain HE48 (I-Fv2 type) FokI-Cascade complex. In one embodiment, the different PAM requirements of these different homologues are capable of increasing the target density in target polynucleotides (e.g., gDNA in cells). Therefore, the collection of Cascade complex homologues provides greater flexibility in the selection of nucleic acid target sequences in target polynucleotides (e.g., gDNA in cells).
[0209] In a second aspect, the present invention relates to modified Cascade subunit proteins. Suitable Cascade subunit proteins for modification include, but are not limited to, Cascade subunit proteins of the species described herein.
[0210] In one embodiment, the present invention relates to modified cyclically arranged Cascade subunit proteins. Such cyclically arranged Cascade subunit proteins result in protein structures having original linear sequences of amino acids with varying connectivity but generally similar three-dimensional shapes (see, for example, Bliven, S., et al., PLoS Comput. Biol. 8:e1002445 (2012)). Cyclicly arranged Cascade subunit proteins can have many advantages. For example, cyclically arranged Cas7 subunit proteins can generate novel N-termini and novel C-termini designed to be positioned for linking with additional polypeptide sequences to form fusion proteins or linker regions without disrupting the Cas7 protein folding or Cascade complex assembly. Figure 3A and Figure 3B Three instances of cyclically arranged Cas7 (circular Cas7, cpCas7) are shown. Figure 3A and Figure 3B The image shows three parts of a protein: the N-terminal portion of a natural protein (…). Figure 3AVertical stripes, for example, in Cas7 protein), and the central portion of natural proteins ( Figure 3A (gray shading) and the C-terminal portion of natural proteins ( Figure 3A (No shadow). Figure 3A This demonstrates repositioning the N-terminal portion of a native protein to its C-terminal position to produce a cyclically arranged protein. Figure 3A cpCas7), in which the N-terminal portion of the native protein is now located at the N-terminus of cpCas7 and is linked to the central portion of the native protein via a linker polypeptide. Figure 3A (connector). Figure 3B This shows the C-terminal portion of a natural protein ( Figure 3B Cas7) was repositioned to the N-terminus of the native protein. Figure 3B cpCas7), wherein the C-terminal portion of the native protein is now located at the N-terminus of cpCas7 and is linked to the central portion of the native protein via a linker polypeptide. Figure 3B (connector).
[0211] The data presented in Examples 10A, 10B, and 10 demonstrate that purification of the Cascade complex containing a cyclically arranged Cas7 subunit protein variant proves that the cyclically arranged IE type CRISPR-Cas subunit protein can be successfully used to form a Cascade complex having a composition (based on molecular weight) substantially the same as that of the Cascade complex containing the wild-type protein.
[0212] In another embodiment, the present invention relates to a Cascade subunit protein fused to an additional polypeptide sequence to produce a fusion protein, and to a polynucleotide encoding such a fusion protein. The additional polypeptide sequence may include, but is not limited to, proteins, protein domains, protein fragments, and functional domains. Examples of such additional polypeptide sequences include, but are not limited to, sequences derived from transcription activator or repressor domains and nucleotide deaminases (e.g., cytidine deaminases or adenine deaminases, as described in Komor, et al., Nature 553:420-424 (2016); Koblan, et al., Nat. Biotechnol. doi:10.1038 / nbt.4172 (May 29, 2018)). Additional functional domains of the fusion protein are illustrated herein.
[0213] Additional polypeptide sequences can be fused to any Cascade subunit protein, wherein the additional polypeptide sequence is encoded by an additional polynucleotide sequence typically attached to the 5' or 3' end of a polynucleotide encoding a sequence containing a Cascade subunit protein. In some embodiments, the additional polynucleotide sequence encoding an amino acid linker links the Cascade subunit protein to the additional target polypeptide sequence. In some embodiments, the polynucleotide sequences of the fusion protein chaperone and linker sequences can be derived from naturally occurring gDNA sequences or can be codon-optimized for bacterial expression in *E. coli* or eukaryotic expression in mammalian cells (e.g., human cells). Example 1 illustrates the inclusion of affinity tags (e.g., His6, Strep-). II (IBA GMBH LLC, Examples of fusion proteins of maltose-binding protein and FokI (Germany) sequence, nuclear localization signal (NLS), and FokI are also disclosed in Example 1. An exemplary amino acid linker sequence is also disclosed.
[0214] Example 11A describes the fusion of the Cascade subunit protein with FokI, and the fusion of the Cascade subunit protein with domains of cytidine deaminase, endonuclease, restriction enzyme, nuclease / helicase, or more. Example 11B describes the fusion of the Cascade subunit protein with other Cascade subunit proteins, and the fusion of the Cascade subunit protein with other Cascade subunit fusion proteins and enzymatic protein domains (Example 11D). In some embodiments, the ability of type I CRISPR subunit proteins to generate protein fusions at N-terminal, C-terminal, or N-terminal and C-terminal locations can be evaluated in a computer. In some embodiments, one or more peptide linkers can be used to link type I CRISPR subunit proteins to one or more fusion domains at N-terminal, C-terminal, or N-terminal and C-terminal locations. In some embodiments, the Cascade subunit protein can be fused to a single-stranded FokI (e.g., single-stranded FokI fused to the EcoCascade RNP complex; nucleotide sequence, SEQ ID NO: 1926; protein sequence, SEQ ID NO: 1927). Exemplary peptide linkers are shown in Examples 1, 11, 18 and 19.
[0215] Figure 4A and Figure 4B The diagram shows a Cas8 subunit protein that includes a fusion with another protein sequence (e.g., FokI). Figure 4A , Figure 4BCas7, Cas5, Cas8, Cse2, Cas6, with the dashed box around Cas6 indicating its interaction with the hairpin of crRNA; cRNA is shown as a black line including the hairpin; and Cas8, indicating the Cascade complex ("C" for the C-terminus, "N" for the N-terminus). Figure 4A This demonstrates the use of a linker polypeptide to link the C-terminus of the Cas8 subunit protein. Figure 4A (The black curve) Other protein sequences ( Figure 4A ,FP) instances. Figure 4B This demonstrates the use of a linker polypeptide to the N-terminus of the Cas8 subunit protein. Figure 4B (The black curve) Other protein sequences ( Figure 4B Examples of FP). Example 11A describes the computer-aided design, cloning, expression, and purification of IE-type Cas8 fused to the N-terminus of the FokI nuclease domain.
[0216] Figure 5A and Figure 5B Further examples of Cascade complexes, comprising Cascade subunit proteins fused to additional protein sequences, are shown. Figure 5A and Figure 5B In the diagram, cRNA is shown as a black line including a hairpin, and the relative positions of the Cas proteins in the Cascade complex are shown. Figure 5A , Figure 5B Cas7, Cas5, Cas8, Cse2, Cas6; the dashed box around Cas6 indicates its interaction with the crRNA hairpin. Figure 5A Each linker polypeptide is shown. Figure 5A (black curve) fused to the detectable portion of each of the six Cas7 subunit proteins (e.g., green fluorescent protein); Figure 5A Examples of GFP (GFP) are provided. Such Cascade complexes can be used to detect the binding of the complex to a nucleic acid target sequence by providing significant signal amplification due to the presence of multiple detectable moieties associated with the Cascade complex. Figure 5B This demonstrates the use of linker peptides to connect with Cas6 subunit proteins. Figure 5B (Black curve) Other protein sequences ( Figure 5B ,FP) instances.
[0217] Examples of fusion proteins containing E. coli type IE Cascade subunits include, but are not limited to, the following: identical subunits (e.g., Cse2 linker_Cse2), circularly arranged subunits (e.g., cpCas7 linker_cpCas7 linker_cpCas7 linker_cpCas7 linker_cpCas7 linker_cpCas7 linker_cpCas7 linker_cpCas7), type IE Cascade proteins fused to nucleases (e.g., FokI linker_Cas8, Cas3 linker_Cas8, Cas6 linker_FokI, S1 nuclease linker_Cse2 linker_Cse2), and type IE Cascade proteins fused to cytidine deaminases (e.g., Cas8 linker_AID, Cse2 linker_Cse2_...). The linker (APOBEC3G) and one or more other IE-type Cascade proteins fused with IE-type Cascade proteins (e.g., Cas6_linker_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas7_cpCas7_linker_cpCas7_cpCas7_linker_cpCas7_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas7_linker_cpCas5_linker_cpCas7 ...
[0218] Figure 6A , Figure 6B and Figure 6C A diagram of a modified type I CRISPR-Cas effector complex containing cpCas7 is provided. Figure 6A , Figure 6B and Figure 6C In the text, "cpCas7" refers to the Cas7 protein with a circular arrangement. Figure 6A , Figure 6B , Figure 6C : cpCas7, Cas5, Cas8, Cse2, and Cas6; the dashed box around Cas6 indicates its interaction with the crRNA hairpin; cRNA is shown as a black line including the hairpin; for cpCas7, the shading corresponds to Figure 3A The image shows the ring-shaped arrangement of proteins and the relative positions of the Cas proteins in the Cascade complex. Figure 6A The Cascade complex, comprising six individual cpCas7 subunits, is shown. Figure 6A ,cpCas7). Figure 6BThe Cascade complex, comprising six fused cpCas7 subunits, is shown, wherein the C-terminus of the cpCas7 subunits ( Figure 6B cpCas7 utilizes a linker polypeptide to connect to the N-terminus of a neighboring cpCas7 subunit protein. Figure 6B The linker polypeptide is shown as a dark black line connecting the cpCas7 subunit protein. Figure 6C An embodiment is shown, wherein the Cascade complex comprises six fused cpCas7 subunit proteins (“backbone”), wherein the C-terminus of the first cpCas7 subunit protein is linked to the N-terminus of the second cpCas7 subunit protein via a linker polypeptide. Figure 6C The linker polypeptide is shown as a dark black line connecting the cpCas7 subunit protein. The C-terminus of the second cpCas7 subunit protein utilizes the linker polypeptide ( Figure 6C (The black straight line connecting cpCas7 and FP) and different protein sequences ( Figure 6C The N-terminus of cpCas7 is linked to that of FP (e.g., cytidine deaminase), and the C-terminus of this protein-coding sequence is linked to the N-terminus of a third cpCas7 via a linker polypeptide. One advantage of such a fusion backbone as the cpCas7 subunit protein is that additional protein sequences can be introduced at specific locations along the backbone to provide additional protein sequences close to different positions along the length of the guide-guided binding nucleic acid target sequence with the Cascade complex.
[0219] Figure 7A and Figure 7B Other embodiments of a modified type I CRISPR-Cas effector complex incorporating a fusion protein are shown. Figure 7A and Figure 7B The image shows the relative positions of the Cas proteins in the Cascade complex. Figure 7A , Figure 7B Cas7, Cas5, Cas8, Cse2, and Cas6; the dashed box around Cas6 indicates its interaction with the crRNA hairpin; cRNA is shown as a black line including the hairpin. Figure 7A The Cascade complex, which includes the Cse2-Cse2 fusion protein, is shown. Figure 7A (The two Cse2 proteins are connected by a curve with a black line). The computer-aided design, cloning, expression, purification, and electrophoretic mobility change determination of the Cascade complex containing the Cse2-Cse2 fusion protein are described in Examples 11B and 11C. Figure 7B It shows the presence of peptides via linkers ( Figure 7B The black curve connecting the Cse2 protein and FP) and other protein sequences ( Figure 7BCascade complex of Cse2-Cse2 fusion protein fused to cytidine deaminase. Example 11D describes the computer-aided design, cloning, expression, and purification of Cse2-Cse2 protein fused to cytidine deaminase.
[0220] In some implementations, one or more nuclear localization signals may be added to the modified N-terminus or C-terminus of the Cascade protein subunit (e.g., Cas8-FokI fusion protein, cpCas7 protein, or Cse2-Cse2 fusion protein).
[0221] In some embodiments of the fusion peptide, the linker peptide links two or more protein-coding sequences. Exemplary linker peptide lengths are described in the examples. Typically, linker lengths include, but are not limited to, about 10 to about 40 amino acids, about 15 to about 30 amino acids, and about 17 to about 20 amino acids. The amino acid composition of the linker peptide typically contains polar, small, and / or charged amino acids (e.g., Gly, Ala, Leu, Val, Gln, Ser, Thr, Pro, Glu, Asp, Lys, Arg, His, Asn, Cys, Tyr). In other embodiments, the linker peptides are designed to be methionine-free and designed for fusion to avoid hidden translation initiation sites. Following the guidance of this specification, linker peptides were designed to provide appropriate spacing and location of functional domains and cascade proteins within the fusion protein (see, for example, Chichili, C., et al., Protein Science 22:153-167 (2013); Chen, X., et al., 65:1357-1369 (2013); George, R., et al., Protein Engineering, Design and Selection 15:871–879 (2002)). Further examples of linker peptides useful in the practice of this invention are linker peptides that link the coding sequences of cascade proteins to each other, as identified in organisms comprising cascade systems (e.g., the linker peptide linking Cas8 to Cas3 in Streptomyces griseus, as described by Westra, ER, et al., Mol, Cell. 46:595–605 (2012)).
[0222] Codon optimization can be performed on the DNA sequence encoding the fusion protein for expression in selected organisms such as bacteria, archaea, plants, fungi, or mammalian cells. Codon optimization programs are widely available, such as on the Integrated DNA Technology website (www.idtdna.com / CodonOpt), or through... (Genscript, Piscataway, NJ) service. To facilitate cloning into receptor expression vectors, additional sequences overlapping with vectors compatible with SLIC clones can be appended at the 5' and 3' ends of the DNA sequence (see, for example, Li, M., et al., Methods Mol. Biol. 852:51-59 (2012)).
[0223] In other embodiments, the Cascade subunit protein may be fused to a transcriptional activation and / or repression domain. In some embodiments, the fusion protein may comprise an activator domain (e.g., heat shock transcription factor, NFKB activator, VP16, and VP64 (see, for example, Eguchi, A. et al., Proc. Natl. Acad. Sci. USA 113: E8257-E8266 (2016); Perez-Pinera, P. et al., Nature Methods 10: 973-6 (2013); Gilbert, LA, et al. Cell 159: 647-61 (2014)) or a repressor domain (e.g., the KRAB domain). In some embodiments, the linker nucleic acid sequence is used to link two or more coding sequences of the protein, protein domain, or protein fragment.
[0224] The Cascade complex, which contains a type I CRISPR-Cas subunit protein fused to a transcription activator, can be used to activate gene expression. The target site may include a transcription initiation site (TSS), which typically has one or more binding sites for cellular transcriptional activation mechanisms (factors). Figure 8 A Cascade complex comprising six fusion proteins is shown, the fusion proteins comprising linker polypeptides ( Figure 8 The black curve connecting cpCas7 and VP64 is linked to the transcriptional activator VP64 via cpCas7 (and...). Figure 3A (Compared to). Figure 8 In the diagram, crRNA is shown as a dark black line including a hairpin, and the relative positions of the Cas proteins in the Cascade complex are shown. Figure 8: cpCas7, Cas5, Cas8, Cse2, and Cas6; the dashed box around Cas6 indicates its interaction with the crRNA hairpin). This engineered design of the Cascade complex transforms the complex into a flexible tool for transcriptional activation of genes (CASCADEa), where targeted genes are achieved by selecting guide sequences that direct the binding of the Cascade complex to one or more regulatory elements (e.g., TSS) of a selected gene. Example 12 describes the design of an E. coli IE type cp-Cas7 protein fused to the VP64 activation domain to confer transcriptional activation activity to the Cascade complex. Transcriptional activators include, but are not limited to: homology domain proteins, zinc finger proteins, winged helical (forkhead) proteins, leucine zipper proteins, helical-loop helical proteins, heterodimeric transcription factors, activation domains, and transcription factors that bind enhancers (see, e.g., Molecular Cell Biology, Harvey Lodish, et al., WH Freeman & Co.; (2002) ISBN 978-0849394805).
[0225] Additionally, Cascade complexes comprising a type I CRISPR-Cas subunit fused to a transcriptional repressor can be used to repress gene expression. The target site may contain a transcriptional regulatory element. In one embodiment, the Cascade subunit protein may be linked to a KRAB domain via a linker polypeptide. Cascade complexes comprising a Cascade subunit protein / KRAB domain fusion can transform the complex into a flexible tool for gene transcriptional repression (CASCADEi), where targeted genes are achieved by selecting a guide sequence that directs the binding of the Cascade complex to one or more regulatory elements of a selected gene. Transcriptional repressors include, but are not limited to: passive transcriptional repressors, the bzip transcription factor family, sp1-like transcriptional repressors, and active transcriptional repressors (e.g., transcriptional repression via histone deacetylases, histone deacetylation, and recruitment by bispecific repressors (see, for example, Thiel, G., et al., Eur. J. Biochem. 271: 2855–2862 (2004); Nicola Reynolds, N., et al., Development 140: 505-512 (2013); Gaston, K., et al., Cell Mol. Life Sci., 60: 721-741 (2003)).
[0226] In another implementation, the Cascade subunit protein can be fused to the affinity tag.
[0227] In other embodiments of the invention, type I CRISPR-Cas guide polynucleotides can be modified by inserting selected polynucleotide elements or by making nucleotide changes at selected sites within the guide polynucleotide (e.g., fundamentally different changes from DNA to RNA, and other changes to the guide polynucleotides described above). Such embodiments include, but are not limited to, type I CRISPR-Cas guide polynucleotides at the 5', 3', or internally fused to one or more nucleotide effector domains (e.g., MS2 or MS2-P65-HSF1 binding RNA or recruiting transcription factors). Figure 9 The diagram shows type I CRISPR guide polynucleotides and the relative positions of the Cas proteins in the Cascade complex. Figure 9 Cas7, Cas5, Cas8, Cse2, and Cas6; the dashed box around Cas6 indicates its interaction with the crRNA hairpin; cRNA is shown as the black line encompassing the hairpin within the dashed box. Figure 9 In the crRNA, there is also an RNA aptamer hairpin introduced into the 3' hairpin of the guide polynucleotide. Figure 9 (as indicated by the arrow).
[0228] The length of the type I CRISPR-Cas guide can also be modified, usually by lengthening or shortening the binding regions of the Cas7 and Cse2 subunits. Figure 10A The Cascade complex, consisting of three Cas7 subunits, one Cse2 subunit, and a shortened crRNA, is shown. Figure 10A Cas7, Cas5, Cas8, Cse2, and Cas6; the dashed box around Cas6 indicates its interaction with the crRNA hairpin; cRNA is shown as a black line including the hairpin. Figure 10B The Cascade complex, consisting of nine Cas7 subunits, three Cse2 subunits, and an elongated crRNA, is shown. Figure 10B Cas7, Cas5, Cas8, Cse2, and Cas6; the dashed box around Cas6 indicates its interaction with the crRNA hairpin; cRNA is shown as a black line including the hairpin.
[0229] Example 16 describes the generation and testing of type I CRISPR-Cas guide crRNA modifications, and the suitability of the modified guide for constructing modified type I CRISPR-Cas effector complexes.
[0230] In a third aspect, the present invention relates to nucleic acid sequences encoding one or more modified cascade components, and expression cassettes, vectors, and recombinant cells comprising nucleic acid sequences encoding one or more modified cascade components. Some embodiments of the third aspect of the invention include one or more polypeptides (e.g., Cse2, Cas5, Cas6, Cas7, and Cas8 proteins, and one or more homology guides), wherein the components are capable of forming effector complexes. Typically, when more than one homology guide is expressed, these guides have different spacer region sequences to direct binding to different nucleic acid target sequences. Such embodiments include, but are not limited to, expression cassettes, vectors, and recombinant cells.
[0231] In one embodiment, the present invention relates to one or more expression cassettes comprising one or more nucleic acid sequences encoding one or more modified cascade components. The expression cassette typically contains regulatory sequences involving one or more of the following: transcriptional regulation, post-transcriptional regulation, or translational regulation. The expression cassette can be introduced into a variety of organisms, including but not limited to bacterial cells, yeast cells, plant cells, and mammalian cells (including human cells). The expression cassette typically contains functional regulatory sequences corresponding to the organism into which it is introduced.
[0232] Other embodiments of the invention relate to vectors comprising one or more nucleic acid sequences encoding one or more modified Cascade components, including expression vectors. The vector may also include sequences encoding selectable or screenable tags. Furthermore, nuclear-targeting sequences may be added, for example, to Cascade subunit proteins. The vector may also include polynucleotides encoding protein tags (e.g., polyHis tags, hemagglutinin tags, fluorescent protein tags, and bioluminescent tags). Sequences encoding such protein tags may be fused with one or more nucleic acid sequences, for example, encoding Cascade subunit proteins.
[0233] General methods for constructing expression vectors are known in the art. Expression vectors for host cells are commercially available. Several commercial software products are designed to facilitate the selection and construction of appropriate vectors, such as insect cell vectors for insect cell transformation and gene expression in insect cells, bacterial plasmids for bacterial transformation and gene expression in bacterial cells, yeast plasmids for cell transformation and gene expression in yeast and other fungi, mammalian vectors for mammalian cell transformation and gene expression in mammalian cells or mammals, and viral vectors (including, but not limited to, lentiviruses, retroviruses, adenoviruses, herpes simplex virus type I or II, parvoviruses, reticuloendothelial proliferator-associated virus (AAV) vectors) for cell transformation and gene expression, as well as methods that readily allow the cloning of such polynucleotides.
[0234] AAV-based vectors (rAAV) are one example of viral vectors that can be used to implement the methods of this invention. AAV is a single-stranded DNA member of the Parvoviridae family and is a naturally occurring replication-defective virus. AAV vectors are the most commonly used viral vectors for gene therapy. Twelve human AAV serotypes (AAV serotypes 1 [AAV-1] to AAV-12) and more than 100 serotypes from non-human sources are known.
[0235] Lentiviral vectors are another example of viral vectors that can be used to implement the methods of this invention. Lentivirals are members of the Retroviridae family and are single-stranded RNA viruses that can infect both dividing and non-dividing cells, and can provide stable expression by integrating into the genome. To increase the safety of lentiviral vectors, the components necessary for generating the viral vector are aliquoted into multiple plasmids. Transfer vectors are typically non-replicating and may contain additional deletions in the 3' LTR, which causes the virus to inactivate itself after integration. Packaging and envelope plasmids are typically used in conjunction with transfer vectors. For example, packaging plasmids may encode a combination of Gag, Pol, Rev, and Tat genes. Transfer plasmids may contain viral LTRs and psi packaging signals. Envelope plasmids typically contain envelope proteins (typically vesicular stomatitis virus glycoprotein, VSV-GP, due to its broad infectivity).
[0236] Exemplary plant transformation vectors include those derived from Ti plasmids of *Agrobacterium tumefaciens* (see Lee, LY, et al., *Plant Physiology* 146:325-332 (2008)). Similarly, *Agrobacterium rhizogenes* plasmids are also useful and known in the art. For example, SNAPGENETM (GSL Biotech LLC, Chicago, IL; snapgene.com / resources / plasmid_files / your_time_is_valuable / ) provides a wide list of vectors, individual vector sequences, vector diagrams, and many commercial sources of such vectors.
[0237] To express and purify recombinant Cascade in a bacterial expression system, vectors encoding Cascade subunit proteins and minimal CRISPR arrays containing target guide sequences can be designed. Therefore, one aspect of the invention includes such an expression system. In one embodiment, the Cascade complex is expressed by three different plasmid vectors that collectively encode the following components: the Cas8 protein; Cse2, Cas7, Cas5, and Cas6 proteins; and a CRISPR RNA. In some embodiments, the expression plasmid encoding Cas8 contains a native gDNA gene sequence, and in other embodiments, the expression plasmid may encode a codon-optimized Cas8 for expression in selected cell types. Similarly, expression plasmids encoding Cse2, Cas7, Cas5, and Cas6 may contain native gDNA gene sequences or may contain gene sequences that have been codon-optimized for expression in selected cell types. In some embodiments, the entire Cascade subunit protein encoding an operon may be located downstream of a single transcription promoter, such that different proteins are translated from a single polycistronic transcript. In another implementation, the genes encoding the Cascade subunit protein can be separated from each other, with a transcription terminator and a promoter in between.
[0238] Expression plasmids encoding crRNA can contain as few duplicates as possible flanking a single spacer sequence and downstream of an appropriate transcription promoter, or they can contain numerous duplicates of the same guide sequence or multiple different guide sequences flanking multiple spacer sequences. Co-expression of CRISPR and the Cascade subunit, particularly the Cas6 subunit, results in the processing of longer precursor crRNAs into mature-length crRNAs, each containing a single duplicate fragment at the 5' and 3' ends of the crRNA, and a single spacer sequence in the middle.
[0239] An alternative strategy for expressing the complete Cascade complex in *E. coli* uses two plasmids: one plasmid encodes the entire Cas8–Cse2–Cas7–Cas5–Cas6 operon on a single expression plasmid, and another plasmid encodes a CRISPRRNA. In this case, the 5' end of the Cse2 gene, which typically overlaps with the 3' end of the Cas8 gene, is spatially separated from the 3' end of the Cas8 gene to allow for the attachment of a multinucleotide sequence encoding an affinity tag and / or protease recognition sequence.
[0240] Example 2 describes two types of bacterial expression plasmid systems for the Cascade protein: the first type contains two plasmids, the first encoding the Cas8 protein and the second encoding all four subunits of the CasBCDE complex (cse2–cas7–cas5–cas6 operons); and the second type contains expression plasmids encoding all five subunits of the Cascade complex (cas8–cse2–cas7–cas5–cas6 operons). Homologous CRISPR arrays are also described.
[0241] To facilitate the purification of the Cascade complex, an affinity tag, such as an N-terminal Strep-II tag or a His6 tag, can be attached to the Cse2 subunit. Alternatively, an amino acid sequence recognized by a protease, such as TEV or HRV3C, can be inserted between the affinity tag and the native N-terminus of the Cse2 subunit, thereby releasing the affinity tag from the final recombinant Cascade complex after initial purification by biochemical cleavage of the sequence by the protease. The affinity tag can also be placed on other subunits or remain on the Cse2 subunit and combine with additional affinity tags on other subunits. Exemplary Cascade subunit proteins containing affinity tags are shown in Examples 1, 2, 3A, 3B, and 3C.
[0242] For the IE-type Cascade system, *E. coli* strains can be transformed with a plasmid encoding CRISPRRNA and the cse2–cas7–cas5–cas6 gene to induce protein expression and generate a Cascade complex lacking the Cas8 subunit. This Cascade complex is commonly referred to as the Cas8-negative Cascade complex, or alternatively as the CasBCDE complex (see, for example, Jore, M., et al., *Nat. Struct. Mol. Biol.* 18:529-536 (2011)). This purified complex can be combined with separately purified Cas8 biochemistry to reconstruct the complete Cascade (see, for example, Sashital, DG, et al., *Mol. Cell 46:606-615 (2012)).
[0243] Table 6 shows exemplary sequences of bacterial expression plasmids encoding minimal CRISPR arrays, including the cas8, cse2–cas7–cas5–cas6 constructs, and the cas8–cse2–cas7–cas5–cas6 construct, each containing different tags and designs. Plasmids encoding the Cascade complex and Cascade complexes from the homologous type I system can be designed similarly to the exemplary expression plasmid sequences of type IE found in *E. coli* K-12MG1655, following the guidance in this specification. Table 6 also contains sequences of expression plasmids expressing the Cas8–Cse2–Cas7–Cas5–Cas6 protein, and FokI fusions with the cas8 or cas6 genes to generate nuclease-Cascade fusions for gene editing experiments.
[0244]
[0245] Table 7 contains bacterial expression plasmids with single polypromoters encoding all five subunit proteins, as well as the crRNA sequences from single bacterial expression plasmids. In this design, each gene is separated from other genes flanking it, upstream and downstream, that have transcription promoters and terminators. Additional sequences encoding affinity tags and / or protease recognition tags, as well as fusions with nuclease proteins, can be introduced to generate Cascade-nuclease fusions for gene editing.
[0246]
[0247] Based on the design criteria outlined in this paper, additional bacterial expression plasmids encoding homologous Cascade complexes from other type I subtypes and other bacteria or archaea can be designed. These expression plasmids can be designed using the gDNA sequence of the Cascade gene, or they can be designed using gene sequences that have been codon-optimized for expression in *E. coli* or other bacterial strains.
[0248] To express Cascade or fuse it to Cascade effectors in mammalian cells such as human cells, eukaryotic expression plasmid vectors have been designed to enable the expression of associated protein and RNA components via eukaryotic transcription and translation mechanisms. In one embodiment, Cascade can be generated in mammalian cells by encoding each protein component on a separate expression vector driven by a eukaryotic promoter (e.g., cytomegalovirus (CMV) promoter) and by encoding crRNA on a separate expression vector driven by an RNA polymerase III promoter (e.g., human U6 promoter). CRISPRRNA can be encoded using a minimal CRISPR array containing at least two repeat sequences flanking one or more spacer sequences, which serve as guide portions for the mature crRNA. Constructs that generate CRISPRRNA can be designed with additional sequences flanking the outermost repeat sequences in the minimal array. Processing of the CRISPRRNA precursor is achieved via the RNA processing subunit of the Cascade complex (Cas6 subunit protein), which can be expressed from a separate plasmid. Processing of the precursor CRISPRRNA is enabled by the RNA processing subunit of the Cascade complex (Cas6 subunit protein), which can be expressed from a separate plasmid.
[0249] Table 8 contains the sequences of individual eukaryotic expression plasmids for each protein of the *E. coli* IE type Cascade complex. The Cas8 subunit can be fused to additional effector nuclease domains, such as FokI nuclease (Examples 1, 3A, 3B, and 3C). Table 8 also contains the sequences of expression plasmids for the crRNA component of Cascade, encoding two separate crRNAs, with three repeat sequences flanking two spacer regions. Each protein-coding gene can be attached to a polynucleotide sequence that attaches a nuclear localization signal (NLS), an affinity tag, and a linker sequence connecting those tags. Other fusions with any Cascade subunit protein can be encoded by additional polynucleotide sequences typically attached to the 5' or 3' coding sequence, including additional polynucleotide sequences encoding amino acid linkers connecting the Cascade subunit protein to additional target polypeptide sequences. Examples of candidate fusion proteins are described herein.
[0250]
[0251] To express components of the Cascade complex on fewer expression vectors, polycistronic expression vectors can be constructed, allowing a single promoter (e.g., the CMV promoter) to simultaneously drive the expression of multiple coding sequences separated by the Thata asigna virus 2A sequence. The 2A viral peptide sequence induces ribosome jumping, enabling multiple protein-coding genes to be tandemly linked in a single polycistronic construct for expression in eukaryotic cells. Therefore, polycistronic vectors can be designed to encode four or five protein subunits of the Cascade complex on a single transcript driven by a single promoter. Table 9 contains sequences of eukaryotic polycistronic expression plasmids that can be combined with CRISPR RNA expression plasmids to generate functional Cascade in mammalian cells.
[0252]
[0253] In some implementations, the CRISPR RNA is encoded in the 3' untranslated region (UTR) of a protein-coding gene, and its expression is driven by an RNA polymerase II promoter (e.g., a CMV promoter) to produce transcripts. In such implementations, a minimal CRISPR array is designed to be located downstream of a protein-coding gene such as Cas6, Cas7, or a reporter gene (e.g., enhanced green fluorescent protein, eGFP), and separated from the protein-coding sequence by a MALAT1 triplet sequence previously shown to confer stability to the upstream transcript. The minimal CRISPR array is processed by the RNA processing subunit of Cascade (typically expressed using a different plasmid)—an endonuclease that cleaves the minimal CRISPR array—introducing breaks in the transcript, and the triplet sequence protects the 3' end of the upstream protein-coding gene from premature exonuclease degradation. Table 10 shows sequences containing three polynucleotide sequences, where the CRISPR sequence is cloned downstream of Cas6, Cas7, or eGFP, and the expression of the entire fusion sequence is driven by a CMV promoter.
[0254]
[0255]
[0256] In some implementations, the CRISPRRNA array is encoded on the same vector as the polycistronic construct expressing the five 5-Cascade subunit proteins; the combination of these two elements produces an all-in-one vector that generates all the functional subunits (protein and RNA) of the Cascade complex, as well as any nuclease or effector domain fused to a Cascade subunit. Table 11 contains two representative sequences of these all-in-one polynucleotide sequences that encode all their respective components to generate functional FokI-Cascade RNPs in mammalian cells.
[0257]
[0258] Examples 3A, 3B, and 3C describe expression systems using separate plasmids and minimal CRISPR arrays expressing each Cascade subunit protein, expression systems in which multiple Cascade subunit protein coding sequences are expressed from a single promoter, and expression systems in which a single plasmid Cascade expression system is constructed to express the entire cas8–cse2–cas7–cas5–cas6 operon and minimal CRISPR array for use in mammalian cells.
[0259] Following the guidance of this specification, those skilled in the art can design additional mammalian expression vectors that encode other Cascade complexes similar to those provided in the Escherichia coli IE type Cascade complex.
[0260] In a fourth aspect, the present invention relates to generating a modified type I CRISPR-Cas effector complex by introducing a plasmid encoding one or more components of a modified type I CRISPR-Cas effector complex into a host cell. Transformed host cells (or recombinant cells) or cell progeny transfected using recombinant DNA technology may contain one or more nucleic acid sequences encoding one or more components of a modified type I CRISPR-Cas effector complex. Methods for introducing polynucleotides (e.g., expression vectors) into host cells are known in the art and are generally selected according to the type of host cell. Such methods include, for example, viral or bacteriophage infection, transfection, binding, electroporation, calcium phosphate precipitation, polyethyleneimine-mediated transfection, DEAE-glucan-mediated transfection, protoplast fusion, lipid transfection, liposome-mediated transfection, particle gun technology, microparticle bombardment, direct microinjection, and nanoparticle-mediated delivery. In one embodiment of the invention, a polynucleotide encoding a component of a modified type I CRISPR-Cas effector complex is introduced into a bacterial cell (e.g., *Escherichia coli*).
[0261] Examples 4A and 4B describe methods for introducing and expressing the coding sequence of the Cas8 protein, as well as the coding sequence of a modified type I CRISPR-Cas effector complex component, for the bacterial production of such complexes using an E. coli expression system.
[0262] The various exemplary host cells disclosed herein can be used to generate recombinant cells using the modified Cascade effector complex. Such host cells include, but are not limited to, plant cells, yeast cells, bacterial cells, insect cells, algal cells, and mammalian cells.
[0263] For ease of discussion, the term "transfection" will be used below to refer to any method of introducing polynucleotides into host cells.
[0264] In some embodiments, host cells are transiently or non-transiently transfected using nucleic acid sequences encoding one or more components of a type I CRISPR-Cas effector complex. In some embodiments, cells are transfected as naturally occurs in the object. In some embodiments, the transfected cells are first removed from the object, such as primary or progenitor cells. In some embodiments, primary or progenitor cells are cultured and / or returned to the same or different objects after in vitro transfection.
[0265] Expression and purification of the modified type I CRISPR-Cas effector complex is labor-intensive; therefore, a higher-throughput plasmid-based delivery system was designed to facilitate screening on a large number of guide polynucleotide or effector complex variants. Each of the five Cas genes was codon-optimized for human use and cloned into a CMV-driven expression plasmid as an N-terminal NLS fusion. A minimal CRISPR array of paired gRNAs containing the TRAJ27 exon targeting the T-cell receptor α site (UCSC genome browser, hg38) was cloned into a sixth plasmid downstream of the human U6 promoter (Example 3A). Figure 35 ).exist Figure 35 In the image, the components are arranged from left to right as follows: hu6 starter, a gray rectangle with a rhombus end; repeating area 1, a hollow rhombus (white); interval area 1, a gray waffle rectangle; repeating area 2, a gray rhombus; interval area 2, a gray dotted rectangle; and repeating area 3, a black rhombus. Figure 35 In the brackets, the regions encoding two gRNAs are indicated. In some embodiments, the two guide RNAs may be the same (e.g., targeting the same nucleic acid target sequence), and in other embodiments, the two guide RNAs may be different (e.g., targeting two different nucleic acid target sequences).
[0266] In most type I systems, gRNAs are naturally catalyzed by Cas6 ribonucleases present in Cascade (see, for example, Brouns, SJ, et al., Science 321:960-964 (2008); Hochstrasser, M., et al., Trends Biochem. Sci. 40:58–66 (2015), avoiding the need for multiple promoter methods with paired gRNAs as shown herein. Therefore, one embodiment of the invention includes a vector containing paired guide polynucleotides operatively linked to regulatory elements to provide expression of guide polynucleotides (e.g., gRNAs). At the TRAJ27 site, six-plasmid co-transfection yields up to ~3% editing, and removal of any component negates genome editing except for Cas11. The *E. coli* Cascade effector complex does not absolutely require Cas11 for DNA binding (see, for example, Westra, E., et al., RNA). Biol.9:1134-1138(2012)).
[0267] In another embodiment of the invention, a minimal CRISPR array, typically containing two guide sequences, is introduced into a cell or biochemical reaction as a DNA template. The DNA template is generated by PCR amplification (e.g., Figure 42A (Example 20A). Such a minimal CRISPR array can be introduced into cells containing one or more plasmids encoding Cascade complex protein components. In some embodiments, both the minimal CRISPR array containing paired guide polynucleotides and the vector can be introduced into cells or biochemical reactions. When using two Cascade RNP complexes (e.g., a method of binding a nucleic acid target sequence or a method of cleaving a nucleic acid target sequence; see, for example...), Figure 15A , Figure 15B , Figure 15C In the method, the minimal CRISPR array can encode two different guide RNAs. Therefore, in some implementations, the two guide RNAs can be different (e.g., targeting two different nucleic acid target sequences). In the method using a single Cascade RNP complex (e.g., when using a type I CRISPR-Cas effector complex associated with the mCas3 protein or a type I CRISPR-Cas effector complex in which the Cas3 fusion protein is associated with the complex; for example, see, for example...), Figure 16A , Figure 17B , Figure 17C , Figure 21A , Figure 21B , Figure 21C , Figure 21DThe minimal CRISPR array can encode two copies of the same guide RNA sequence. Therefore, in some implementations, the two guide RNAs can be identical (e.g., targeting the same nucleic acid target sequence).
[0268] In yet another embodiment, a polynucleotide encoding a guide sequence that also contains a sequence and structure recognized by the Cas6 protein for processing precrRNA nucleotides into mature guide RNAs can be introduced into a cell or biochemical reaction. In other embodiments, mature guide polynucleotides that do not require processing can be used for the assembly of cascade complexes. Such mature guides may contain sequence modifications (e.g., phosphate thioester bonds at the 5' and / or 3' ends to help protect the guides from digestion by nucleases, such as RNases). Additional guide modifications include those described herein for nucleotide sequences (e.g., nucleotide analogs, etc.).
[0269] Examples 9A, 9B, 9C, and 9D illustrate the design and delivery of an *E. coli* type IE Cascade complex containing the FokI fusion protein to facilitate genome editing in human cells. Example 9B describes the delivery of a plasmid vector expressing components of the Cascade complex into eukaryotic cells. In a fifth aspect, the present invention relates to the purification of modified type I CRISPR-Cas effector complexes from cells, and the application of such complexes. The modified type I CRISPR-Cas effector complexes are generated in host cells. The modified type I CRISPR-Cas effector complexes (in this case, the Cascade RNP complex) are purified from cell lysates.
[0270] Examples 5A and 5B describe the purification of *E. coli* type IE Cascade RNP complexes generated by overexpression in bacteria, as described in Example 4B. The method uses immobilized metal affinity chromatography followed by size exclusion chromatography (SEC). Examples 5A and 5B describe methods that can be used to evaluate the quality of purified Cascade RNP products. Examples provided illustrate the purification of Cas8, Cas7, Cas6, Cas5, and Cse2 Cascade RNP complexes, Cascade complexes containing Cas7, Cas6, Cas5, and Cse2 proteins, and the FokI-Cas8 fusion protein.
[0271] The purified, modified type I CRISPR-Cas effector complex can also be used directly for biochemical assays (e.g., binding and / or cleavage assays). Examples 6A, 6B, and 6C describe the generation of dsDNA target sequences for in vitro DNA binding or cleavage assays. Example 6 describes three methods for generating target sequences, including annealing of synthetic ssDNA oligonucleotides, PCR amplification of selected nucleic acid target sequences from gDNA, and cloning of the nucleic acid target sequences into bacterial plasmids. The dsDNA target sequences were used in Cascade binding or cleavage assays.
[0272] Where necessary, site-specific binding and / or cleavage by one or more modified type I CRISPR-Cas effector complexes can be verified using electrophoretic mobility variation assays (see, for example, Garner, M., et al., Nucleic Acids Res. 9:3047-3060 (1981); Fried, M., et al., Nucleic Acids Res. 9:6505-6525 (1981); Fried, M., Electrophoresis 10:366-376 (1989); Fillebeen, C., et al., J.Vis.Exp. (94), e52230, doi:10.3791 / 52230 (2014)), or the biochemical cleavage assays described in Example 7.
[0273] The data shown in Example 7 demonstrate that the modified type I CRISPR-Cas effector complex can exhibit near-quantitative DNA cleavage, as confirmed by transforming supercoiled, circular plasmid substrates into linear cleavage forms. Following the demonstration of robust biochemical activity using the modified type I CRISPR-Cas effector complex (e.g., a fusion protein containing the FokI-Cascade component), genome editing was performed in cells.
[0274] Examples 8A, 8B, 8C, and 8D illustrate the design and delivery of E. coli type IE Cascade complexes containing the Cas subunit protein-FokI fusion protein into human cells. Data from Example 8D demonstrate efficient genome editing through the delivery of pre-assembled Cascade RNPs to target cells and human cells.
[0275] Purified and modified type I CRISPR-Cas effector complexes can be directly introduced into cells. Methods for introducing components into cells include electroporation, lipid transfection, particle gun technology, and particle bombardment.
[0276] Figure 36A , Figure 36B , Figure 36C and Figure 36D Comparative data on genome editing in human cells using plasmid-based delivery with modified Cascade-RNP and modified type I CRISPR-Cas complexes are provided. Figures 36A-36D , Figure 36A In this study, purified RNPs were transfected into HEK293 cells, and then the edited sites were analyzed using next-generation sequencing (NGS). Figure 36A As shown in (RNP transfection), the FokI-Cascade RNP complex, targeting two adjacent sites, is... Figure 36A (Shown above the straight line on the left side of the figure) Nuclear transfection into HEK293 cells ( Figure 36A (Star-shaped, gray on the left of the figure) to induce DNA cutting and genome editing. Editing efficiencies at 16 unique genomic target sites were calculated (see Example 6C, Table 31, human double Hsa1-16) (n=1). TRACs are constant regions of T cell receptors. When T cell receptors are generated, they include splice junctions (i.e., “variable” regions and “connecting” regions). Some TRAC guides described herein target connecting regions (e.g., TRAJ27). The spacing between each target region is shown below the figure ( Figure 36A From left to right, 25, 30, 35, 40, 45 base pairs (bp). Figure 36A In the middle, the vertical axis represents percentage editing efficiency ( Figure 36A Editing efficiency (%), the horizontal axis represents targets 1-16, and below the horizontal axis are brackets indicating the length of the interval in base pairs (bp).
[0277] Figure 36B Provided Figure 36A Representative DNA repair results for target 7. Figure 36B In the figure, the relative positions of the half-sites targeted by paired gRNAs and their associated PAM sites are shown at the top. The intervening intervals are shown in the top row. The expected cleavage sites are also shown at the top of the figure. Figure 36B The position "0" is shown as a vertical black midline and the bp distance (-50 to 50). Each horizontal gray line represents a different class of sequencing reads observed at the target site. The indicators for these lines are as follows: gray area = sequence match; horizontal black line = deletion; and hollow box = insertion. To the right of each line in the figure is a circle: a black circle is a wild-type read; and a white hollow circle is a mutant read. The expected wild-type read is shown in the first gray box ("Ref"; i.e., the reference sequence). The wild-type read is shown in the second gray bar (the second gray bar; Figure 36B (in black circles). The next 11 lines show the mutant reads ( Figure 36B(Hollow circle). The insertion length, given by the number of base pairs, is displayed in the column to the right of the circle. The total percentage of reads is displayed in the next column to the right, and the total read length is shown in the last column to the right.
[0278] like Figure 36C The 6-plasmid transfection system shown in the image transfected HEK293 cells with 6 different plasmids. Figure 36C (Star-shaped, gray, on the left of the image), five plasmids encoding the Cas protein ( Figure 36C Plasmids indicating FokI-Cas8, Cas11, Cas7, Cas5, and Cas6, and a plasmid encoding paired gRNAs, are located under the control of the CMV and human U6 (hU6) promoter. Figure 36C (gRNA), followed by NGS analysis of the editing sites. The FokI-Cascade RNP complex is illustrated below the dashed line. The values from... Figure 36A Editing efficiency at target 7 (n=2) Figure 36A (black bars in the image), and includes plasmid mixtures lacking a single component ( Figure 36C Below the horizontal axis, including the gray boxes containing - / +, serve as a reference. Figure 36C (The hollow strip in the diagram).
[0279] like Figure 36D As shown in the (2-plasmid transfection system), paired gRNA expression plasmids ( Figure 36D (gRNA plasmid) and expression plasmids encoding polycistronic proteins separated by the T2A "ribosome jumping" sequence peptide ( Figure 36D HEK293 cells were transfected with CMV-Cas7-2A-Cas11-2A-Cas5-2A-Cas6-2A-FokI-Cas8. Figure 36D (Star-shaped, gray, left side of the figure), followed by NGS analysis of the edited sites. The FokI-Cascade RNP complex is illustrated below the dashed line. For those from... Figure 37C 2-plasmid system transfection ( Figure 36D Transfection with a 6-plasmid system (n=3) using hollow frames and hollow frames (n=3) Figure 36D (black bar), calculated Figure 36A The editing efficiency at the 16 targets shown in the image. Figure 36D In the diagram, the vertical axis represents the percentage editing efficiency ("Editing efficiency (%)"), the horizontal axis represents targets 1-16, and below the horizontal axis are brackets indicating the length of the interval interval in base pairs (bp). Figure 36D (From left to right, 25, 30, 35, 40, 45bp).
[0280] The experiments were performed by nuclear transfection of HEK293 cells with purified Cascade-RNPs containing nuclear localization signal sequences on FokI and Cas6. Next-generation sequencing of PCR amplicon obtained from gDNA demonstrated an editing efficiency of up to ~4%, and in the 16 target sites tested, editing was typically located at sites containing a 30 bp interval (...). Figure 36A Careful examination of the spectrum of the repaired results revealed that the insertions and deletions clustered in the middle of the interstitial regions. Figure 36B This design is consistent with that of the type I CRISPR-Cas complex. Therefore, in one embodiment of the invention, the modified type I CRISPR-Cas complex is directly introduced into the cell. For plasmid delivery experiments ( Figure 36C The assembled plasmid mixture contained 420 ng of each plasmid except for one, and then water was added as a negative control after nuclear transfection, or 700 ng of the missing plasmid was added. For the initial FokI-EcoCascade polycistronic 2-plasmid delivery experiment ( Figure 36D Cells were electroporated with 500 ng of each plasmid or 500 ng of paired gRNA expression plasmids and 2.5 μg of polycistronic plasmids (total 3 μg for each condition). In one embodiment, all five cas genes were constructed in a single polycistronic expression vector tandemly linked by a T2A “ribosomal skipping” sequence (see, e.g., Kim, J., et al., PLoS ONE 6, e18556 (2011); Liu, Z., et al., Sci. Rep. 7:2193 (2017)). Figure 36D Surprisingly, the editing efficiency and DNA repair results produced by co-transfection with polycistronic plasmids and paired gRNA expression plasmids were similar to those observed using the 6-plasmid method (Example 9A) and the direct RNP delivery method (Examples 8A, 8B, 8C, and 8D), supporting the conclusion that biochemically modified type I CRISPR-Cas effector complexes can be assembled and delivered to the nucleus of human cells. In summary, these experiments validate a greatly simplified expression system that can reconstruct delicate 11-subunit RNA-guided nucleases in eukaryotic cells using only two molecular components similar in size to widely used Cas9 and sgRNA plasmids.
[0281] Data from modified type I CRISPR-Cas complexes (Escherichia coli (EcoCascade), Pseudomonas S-6-2 (PseCascade), and Pseudomonas aeruginosa (SthCascade)) indicate that most target sites will be unique, as they must include two half-sites, the necessary spacer interval, and the permitted PAM. Modified Cascade homologues from EcoCascade, PseCascade, and SthCascade were selected for more detailed characterization.
[0282] Figure 37A , Figure 37B , Figure 37C and Figure 37D Editing efficiency related to FokI connectors, interval length, and Cascade homologues is shown. Figure 37A The editing efficiency of FokI-EcoCascade is displayed as the length of the FokI-Cas8 connector ( Figure 37A Hollow circle, bottom line 10aa; hollow circle top line, 20aa; black circle, 17aa; and gray circle, 30aa (connecting sub-length) and interval spacing function. Figure 37A In the diagram, the vertical axis represents editing efficiency (%), and the horizontal axis represents intervals in bp. Each data point represents an average of 3–4 unique target sites.
[0283] Figure 37B FokI-Cascade nucleases with 30-aa linkers were provided. FokI-Cas8 linkers with 12 IE-type Cascade variants were generated, and genome editing at 4–7 target sites was tested. Each data point represents a single genomic site, and the bars show the mean and standard deviation (SD) between sites. Targets contain AAG (… Figure 37B (gray bar) or GAA ( Figure 37B (White bar) PAM sequence and 30bp spacer interval, where the species on the horizontal axis are as follows: Eco, Escherichia coli; Pse, Pseudomonas S-6-2; Sen, Salmonella enterica; Geo, Geothermal bacillus EPR-M; Mar, Methanococcus oryzae; Ahe, Arantibacter argentiformis; Oceanus HL-35; Pae, Pseudomonas aeruginosa; Sth, Streptococcus thermophilus; Str, Streptococcus spp.; Kpn, Klebsiella pneumoniae; Lba, Trichophyton family bacteria.
[0284] exist Figure 37C The image shows FokI-PseCascade data, where the vertical axis represents percentage editing efficiency. Figure 37CEditing efficiency (%) is plotted, and the horizontal axis represents the length of intervals in base pairs (bp). The FokI-Cas8 linker is 17 amino acids long. Each data point represents a single genomic site, and the bars show the mean and sd between 7–8 sites.
[0285] Figure 37D Data on FokI-PseCascade editing efficiency as a function of the PAM sequence is provided, with the vertical axis representing percentage editing efficiency. Figure 37D Editing efficiency (%), and the horizontal axis corresponds to the PAM sequence ( Figure 37D From left to right: CCG, CGC, AAG, AAA, ATG, AAC, AGG, ATA, GAG, and AAT. Each genomic locus contains one AAG PAM and a variable PAM at the second half of the locus, as shown on the horizontal axis. Each data point represents a single genomic locus, and the bars show the mean and sd across 6–15 loci.
[0286] Figure 37E Provides FokI-EcoCascade editing efficiency ( Figure 37E The data, denoted by the vertical axis and editing efficiency (%), are functions of the PAM sequence. Target sites contain fixed AAG PAM and variable PAM at the second half-site, as shown on the horizontal axis. Figure 37E From left to right: CCG, CGC, AAG, AGG, ATG, GAG, AAA, AAC, ATA, and AAT. Each point represents a single target site in HEK293 cells, and each PAM tested 6–15 sites (n = 1 / site). The bar chart shows the mean and sd.
[0287] Figure 37F Provides FokI-SthCascade efficiency ( Figure 37F The data, denoted by the vertical axis and editing efficiency (%), are functions of the PAM sequence. Target sites contain a fixed GAA PAM and a variable PAM at the second half-site, as shown on the horizontal axis. Figure 37F (From left to right: CC, AA, GA, TA, and CA). Each point represents a single target site in HEK293 cells, and each PAM tested 18–33 sites (n = 1 / site). The bar chart shows the mean and sd.
[0288] Figure 37G The provided heatmap shows the heat from Figure 37C and Figure 37DThe graph shows the insertion / deletion category frequencies at 40 genomic sites with high editing efficiency (10–53%). The top bar chart shows the percentage editing efficiency from 0–60. The heatmap shown in the middle figure shows insertion lengths of 1–8 bp, and the heatmap in the bottom figure shows deletion lengths of 1–50 bp. The 40 genomic target sites are indicated on the horizontal axis (1–40). Figure 37G , target). Single bp insertions are separated by nucleotide identity, and the gray intensity bar at the bottom of the figure corresponds to the percentage of insertion frequency ( Figure 37G Ins Freq (%), scale from 0 to greater than or equal to 20) and missing frequency percentage ( Figure 37G Del Freq (%), scale from 0 to ≥20). The bar chart on the right shows the average frequency for each insertion / missing category. Figure 37G (Scale bar from 0 to 20). The pie chart on the right shows the fraction of 2–4 bp inserts generated from the putative templated repair. Figure 37G The black area of the pie chart (represented by the cut-off site) is defined here as containing repeating sequences adjacent to the cut-off site. "Other" is represented by the gray area of the pie chart.
[0289] The study investigated the most closely associated sites of five of the most highly edited FokI-PseCascade target sites (~20-48% edit) in the human genome, constrained only by a 30-33 bp spacer interval requirement. No <22 mismatch sites were identified in either of the two halves of any of the five targets. Experiments were conducted on FokI-EcoCascade FokI-Cas8 connector types and spacer intervals. Figure 37A Cells were nuclearly transfected with 2.4 μg of FokI-EcoCascade polycistronic plasmid and ~0.5-3.5 μg of paired gRNA expression plasmids.
[0290] For screening of FokI-Cascade homologues ( Figure 37B Cells were nuclearly transfected with 1.5 μg of FokI-Cascade polycistronic plasmid and ~0.4-2.2 μg of paired gRNA expression plasmids. Throughout the homologue, 4-7 sites were targeted, with sites exhibiting high editing efficiency of FokI-EcoCascade selected. Editing experiments were conducted on the homologue variant FokI-Cas8 linker type and spacer spacing. Figure 37C and Figures 41A to 41C Cells were nuclearly transfected with 5 μg of polycistronic plasmid and ~100-400 ng of oligomer-templated paired gRNA expression amplicones. For this experiment, the gRNA concentration in each well or homologous variant was not normalized. Additionally, for… Figures 41A to 41CCells were nuclearly transfected with FokI-PseCascade gRNA at an average of ~1.5x or more compared to FokI-EcoCascade or FokI-SthCascade gRNA.
[0291] This article describes oligomer-templated PCR amplification (e.g., Example 20A). Figure 42A and Figure 42B The image shows the effect of the human U6 (hU6) promoter in mammalian cells. Figure 42A A PCR strategy for templated oligomers to generate amplicon for paired gRNA expression (420). In short, reverse internal oligonucleotides ( Figure 42A (424) encodes two gRNA sequences and is modified for novel target sites (also known as unique primers encoding "repetition-spacer-repetition-spacer-repetition" sequences). Figure 42A 421: Repeat region, hollow rectangle; Spacer region 1, gray rectangle; Repeat region, hollow rectangle; Spacer region 2, gray rectangle; Repeat region, hollow rectangle), while the remaining primers remain unchanged. Figure 42A : Forward external primer, 422; Forward internal primer, 423; Reverse external primer, 425). Figure 42B The figure shows the editing efficiency at target 7 after co-transfection of HEK293 cells with a polycistronic plasmid encoding the FokI EcoCascade RNP complex and paired gRNA expression plasmids or paired gRNA expression amplicons (see Figure 1). Figure 36B ).exist Figure 42B In the graph, the vertical axis represents editing efficiency (%) and the horizontal axis represents paired gRNA cassettes (ng). Data points are as follows: FokI-EcoCascade RNP complex (ng), paired gRNA plasmids, paired gRNA amplicon; 375, hollow triangle, hollow circle; 750, black triangle, black circle; 1,500, gray triangle, gray circle; 3,000, black triangle with white line, black circle with white line. Figure 42B The data demonstrate that paired gRNA expression amplicones have comparable, if not higher, editing efficiency compared to paired gRNA expression plasmids.
[0292] For PAM screening ( Figure 37D , Figure 37E , Figure 37F , Figures 39A-39D , Figure 40C and Figure 40FTypically, cells are nuclear transfected with 3 μg of FokI-Cascade polycistronic plasmid and 150 ng (FokI-PseCascade and FokI-EcoCascade) or ~80-120 ng (FokI-SthCascade) of oligomer-templated paired gRNA expression amplicon (unless otherwise specified).
[0293] In order to perform specific analysis ( Figures 38A to 38C Cells were nuclear transfected with 3 μg of polycistronic Cascade and 150 ng of oligomeric templated paired gRNA expression amplicon, and harvested 5 days after nuclear transfection. Figure 38A At the top, horizontal lines represent the spacing between regions, scissors represent the expected cut sites, and half-sites of the genomic target are displayed together with their corresponding PAM regions. Figure 38A (Rectangular boxes with contrasting ends). The relationship between the shown half-site and the target is shown by dashed lines. For each target, 32 base pairs are shown, and the PAM region is shown as the neighboring seed sequence. Figure 38A Pairs of gRNAs are provided, designed to contain mismatches with one or both halves of a genomic target, as shown by the filled boxes (excluding PAM sites) in the grid. Note that, for simplicity, both halves are shown in the same orientation. Figure 38B The relative editing efficiency at 70 genomic targets for each combination of mismatched paired gRNAs is provided, plotted as a percentage of editing efficiency for perfectly matched gRNAs. Figure 38B In the middle, the top row represents the target ( Figure 38B Target 70), the next line represents the guide ( Figure 38B (gRNA1 and gRNA2), the next line identifies the sets that do not match ( Figure 38B The next row shows the FokI-Cascade RNP complex (mm set 1 and mm set 2). The left column shows data relative to editing wizard 1-mm set 1 / wizard 2-mm set 2, and the right column shows data relative to editing efficiency percentages. Figure 38B The relative edit eff (%) (scale 0-100) means that the left column shows the data for gRNA1 and gRNA2 with mismatch (mm) sets 1 and 2, and the right column shows the data for the same target but with mismatch (mm) sets (n=1) with exchanges between gRNA1 and gRNA2. Figure 38C It provides editing efficiency at 73 target points (n=1), such as Figure 38B As shown in the image.
[0294] After developing a scalable method that eliminates the need for labor-intensive cloning steps by generating paired gRNA expression cassettes via oligomer-templated PCR amplification (as described herein), a set of 96 genomic targets for each homologous variant were re-screened for reduced FokI linker and DNA spacer interval lengths. Using the 17-aa linker, FokI-PseCascade consistently produced an average editing efficiency of ~15-25% within a spacer interval window of approximately 30–33 bp, with some targets showing insertion / deletion rates as high as ~40-50%. Figure 37C Similar trends were observed using other homologues. PAM requirements were investigated by targeting genomic sites containing one homologous PAM and a second mutant PAM. In vitro, it has been shown that PAM recognition is much more heterogeneous than that of rigid 5'-GG-3' Streptococcus pyogene (see, for example, Szczelkun, M., et al., Proc. Natl. Acad. Sci. USA 111:9798–9803 (2014); Hayes, R., et al., Nature 530:499–503 (2016); Westra, E., et al., Mol. Cell. 46:595–605 (2012); Fineran, P., et al., Proc. Natl. Acad. Sci. USA 111:E1629–E1638 (2014); Leenay, R., et al., Mol. Cell. 62:137–147 (2016)). Surprisingly, in vitro data showed that a significant amount of PAMs were indeed allowed to be active, exhibiting a clear hierarchical bias. Figure 37D ; Figures 39A to 39D Conversely, when the mutated PAM represents a “self” target from the CRISPR array, the editing is completely abolished.
[0295] exist Figures 39A-39D In each of these, the vertical axis corresponds to editing efficiency (editing efficiency (%)), and the horizontal axis corresponds to the PAM sequence associated with the target. Figure 39A The FokI-PseCascade editing efficiency as a PAM sequence function is provided. The genomic site contains a fixed ATG PAM and a variable PAM at the second half of the site, as shown on the horizontal axis. The bars show the mean and sd (6–14 sites per variable PAM, n = 1 / target site). Note that... Figure 37D The data for FokI-PseCascade is described, where one PAM is fixed at AAG, and another PAM is variable among a set of PAMs including ATG. Therefore, a subset of those PAMs is AAG-ATG. Figure 39A The data for FokI-PseCascade is described, where one PAM is fixed at ATG, and the other PAM is variable in a set of PAMs including AAG. Figure 39A The horizontal axis, from left to right, represents AAG, AAC, AAA, ATG, GAG, ATA, AAT, and AGG. Therefore, a subset of those PAMs is still AAG-ATG, and is... Figure 37D The same AAG-ATG site in the sample.
[0296] Figure 39B FokI-EcoCascade editing is provided as a function of PAM sequences. Figure 39B The horizontal axis, from left to right, shows CCG, CGC, AAG, AGG, ATG, GAG, AAA, AAC, ATA, and AAT. The fixed PAM is AAG, and the bars show the mean and sd (6-15 sites per variable PAM, n = 1 / target site). Figure 39C ( Figure 39C The horizontal axis, from left to right, provides AAG, ATG, AAC, AAA, AGG, GAG, AAT, and ATA as an example. Figure 39B Similar analysis was shown, but the first PAM was fixed to ATG (6-14 sites per variable PAM, n = 1 / target site). Figure 39B The ATG column corresponding to AAG-ATG pairs (mean value ~3) and Figure 39C The AA column corresponding to the AAG-ATG pair (with a mean of ~3) is the same. Note that the vertical axis has a different scale. Figure 39D FokI-SthCascade editing is provided as a function of PAM sequences. Figure 39D The horizontal axis, from left to right, shows CC, AA, GA, TA, and CA. The fixed PAM is GAA, and the bars show the mean and sd (18-33 sites per variable PAM; n = 1 / target site).
[0297] Figure 40A , Figure 40B , Figure 40C , Figure 40D , Figure 40E and Figure 40F Data relating to exemplary variations in the editing efficiency of the modified type I CRISPR-Cas complex are shown. [The data obtained are missing from the original text.] Figure 40A (FokI-PseCascade) and Figure 40D The percentage editing efficiency (vertical axis) shown in (FokI-SthCascade) relative to the interval spacing in bps (horizontal axis) is essentially as described in Example 20C for... Figure 41A and Figure 41C The data shown is described. Figure 40A and Figure 40D In the graph, the horizontal axis represents the 23-34 bp spacer interval, and the bars from left to right represent the FokI-Cas8 peptide linker lengths of 17 amino acids (light gray bars), 20 amino acids (dark gray bars), and 30 amino acids (white bars). Basically, according to... Figure 39B The acquisition mentioned Figure 40C and Figure 40F The data shown. Figure 40C and Figure 40F Provides editing for FokI-PseCascade and FokI-SthCascade. Figure 40C , Figure 40F Vertical axis, editing efficiency (%) as PAM sequence ( Figure 40C From left to right: CCG, CGC, AAG, AAA, ATG, AAC, AGG, ATA, GAG, and AAT; Figure 40F From left to right, the functions are CC, AA, GA, TA, and CA. Figure 40B The FokI-PseCascade RNP complex is shown. The fixed PAM of FokI-PseCascade is AAG ( Figure 40B AAG PAM), and another PAM in a group of PAMs ( Figure 40B In the variable PAM), it is variable. Figure 40E The FokI-SthCascade RNP complex is shown. The fixed PAM of FokI-SthCascade is GAA ( Figure 40B GAAPAM), and another PAM in a group of PAMs ( Figure 40E The preference for PAMs is variable in the FokI-PseCascade (variable PAM). The linker and spacer region preferences were rescreened, and data showed nearly 50% editing. PAM preference was also examined. From this data, the in vitro rank preference for PAMs was determined. Essentially, the same analysis was performed on variants of Streptococcus thermophilus. Editing was lower in the Streptococcus thermophilus system. However, the data presented in this paper suggest that, in vivo, in human cells, the Streptococcus thermophilus system exhibits a highly heterogeneous preference for PAMs. The fact that a single A upstream of the pre-spacer region (i.e., the target sequence) allows for editing generally provides an increased number of potential target sequences within the gene (e.g., relative to the number of potential type II CRISPR-Cas9 PAM-related target sites within the same gene). Furthermore, the in vivo data presented in this paper correlate with the in vitro PAM preference demonstrated by Sinkunas, T., et al., EMBO J.32:385-394 (2013).
[0298] Accumulated NGS data across hundreds of edited genomic sites provided the ability to characterize DNA repair outcomes of DSBs introduced via FokI-PseCascade. Focusing on 40 unique sites with insertion / deletion frequencies >10%, the frequencies of deletions and insertions were analyzed as a function of the total mutant read length within a 50 bp window around the predicted cleavage site. Insertions of 2–4 bp were highly enriched and present in the vast majority of sites examined. Figure 37E Detailed examination revealed that ~90% of these insertions contained perfect repeats of sequences adjacent to the cut site. While not wishing to be limited by any particular theory, this repeatability is likely a result of templated repair of the staggered cuts introduced by the dimer FokI.
[0299] The specificity of FokI-PseCascade was evaluated by editing two highly efficient target sites using a large number of mismatched paired gRNAs. Figure 38A Previous studies by Cascade have highlighted the ~8-nt PAM proximal seed sequence and mismatches at every 6th position in the 32-nt guide gRNA, as these bases are flipped out from the RNA-DNA heteroduplex structure formed after target binding (see, for example, Jung, C., et al., Cell 170:35–47 (2017); Mulepati, S., et al., Science 345:1479–1484 (2014); Fineran, P., et al., Proc. Natl. Acad. Sci. USA 111: E1629–E1638 (2014); Semenova, E., et al., Proc. Natl. Acad. Sci. USA 108:10098–10103 (2011)). Mismatches in the proximal seed region of PAM are highly detrimental to genome editing, while mismatches in the distal region of PAM are well tolerated, resulting in near-wild editing efficiency. Figure 38B ; Figure 38C However, when mismatched blocks were present at both halves of the site, editing in the tested whole-pair gRNAs decreased sharply. Figure 38B , Figure 38C Based on PAM data and FokI-PseCascade-mediated genome editing and interval data ( Figure 38C ; Figure 37D One advantage of the modified type I CRISPR-Cas complex of the present invention is that the targetable sites can appear once every ~20 to ~30 bp in the human genome, while editing at potential off-target sites is impossible.
[0300] Therefore, in one embodiment of the invention, the potential target sites or "target density" of a given modified FokI-Cascade system is a function of its effective spacing distance and PAM preference, and will have some variability among homologs. In some embodiments, the target density of FokI-PseCascade, FokI-EcoCascade, and FokI-SthCascade in the human genome can be calculated using the following criteria (extrapolating data to calculate predicted target density).
[0301] The following motifs can be used to calculate the target density of FokI-PseCascade:
[0302] 5'–[half-site 1–PAM1]–[interval]–[PAM2–half-site 2]–3'.
[0303] Here, [half-site 1–PAM1] represents the inverse complement of the target sequence and PAM of the half-site 1 gRNA1 target strand, and [half-site 2–PAM2] represents the half-site 2 gRNA2 non-target strand PAM and target sequence. This is based on the distribution of interval lengths supporting editing with FokI-PseCascade (see, for example,...). Figure 37D The effective interval length is approximately 30-33 bp. PAMs are defined as belonging to set 1 (AAG, AAA, ATG, AAC) which gives the highest edit, or to set 2 (AAG, AGG, ATG, GAG, AAA, AAC, AAT, ATA) if they contain any test PAMs that show activity (see, for example). Figure 39A ; Figure 40B Accordingly, potential target sites that satisfy the preferred interval length of the two PAMs belonging to set 1 or set 2 will appear once on average every 33.4 bp or 9.2 bp, respectively.
[0304] Similarly, target density was determined for FokI-EcoCascade, except that the interval length was defined as 31–33, and PAMs were defined as belonging to set 1 (AAG, AGG, ATG, GAG, AAA) of the highest editing level, or to set 2 (AAG, AGG, ATG, GAG, AAA, AAC, AAT, ATA) if they contained any PAMs that showed activity (see, for example). Figure 39C ; Figure 39D Based on this, potential target sites were calculated using either ensemble 1PAMs or ensemble 2PAMs, which appeared on average once every 30.4 bp and 12.2 bp, respectively.
[0305] Similarly, the density of human genomic targets in FokI-SthCascade was determined, except that the interval length was defined as 29–31 bp and PAMs were defined as NNAs (see, for example). Figure 39D Based on this, the average occurrence of potential target sites was calculated to be once every 4 bp.
[0306] Therefore, as described herein, the modified type I CRISPR-Cas complex provides a method for providing a variety of potential target sites by offering a number of PAM-adjacent target sequences that can be used for genome editing. Accordingly, one embodiment of the invention relates to a method for providing an increased number of available target sequences within a gene using PAM sequences associated with the modified type I CRISPR-Cas complex (e.g., relative to the number of available target sequences associated with PAM sequences of type II or V CRISPR-Cas systems). Application of this method involves using a modified type I CRISPR-Cas complex, which may include, but is not limited to, binding and / or cleaving of target sequences, mutations in target sequences, transcriptional regulation associated with target sequences or their regulatory elements, and target sequences mediated by the modified type I CRISPR-Cas complex described herein, as well as intentional modifications, alterations, and / or significantly different structural changes (e.g., in gene products) mediated by the modified type I CRISPR-Cas complex described herein.
[0307] In some implementations, modifications, alterations, and / or mutations in gDNA can be generated by site-specific introduction of selected polynucleotide sequences (e.g., a subset of donor polynucleotides) at DNA target sites in the genome, using the modified type I CRISPR-Cas effector complex described herein to produce non-human transgenic organisms. The transgenic organisms can be animals or plants.
[0308] Transgenic animals are typically produced by introducing a modified type I CRISPR-Cas effector complex into fertilized egg cells. The basic technique described in the preparation of transgenic mice (see, for example, Cho, A., et al., “Generation of Transgenic Mice,” Current Protocols in Cell Biology, CHAPTER. Unit-19.11 (2009)) involves five basic steps: first, preparing a system as described herein, comprising suitable donor polynucleotides; second, harvesting donor fertilized eggs; third, microinjecting the system into mouse fertilized eggs; fourth, implanting the microinjected fertilized eggs into pseudopregnant recipient mice; and fifth, performing genotyping and analyzing the modifications to the gDNA established in the first-generation mice. The first-generation mice will pass on the genetic modifications to any offspring. First-generation mice are typically heterozygous for the transgene. Mating these mice will produce mice that are homozygous for the transgene for 25% of the time.
[0309] Methods for producing transgenic plants are well known and can be applied using modified type 1I CRISPR-Cas effector complexes. For example, transgenic plants produced using Agrobacterium-mediated transformation typically contain a single transgene inserted into one chromosome. By allowing isolated transgenic plants containing a single transgene to sexually interbreed with themselves (i.e., self-pollination), transgenic plants that are homozygous relative to the transgene can be produced. Typical conjugation assays include, but are not limited to, single nucleotide polymorphism assays and thermoamplification assays to distinguish homozygotes from heterozygotes.
[0310] In a sixth aspect, the present invention relates to the generation of substrate channels using a modified type I CRISPR-Cas effector complex. In some embodiments, fusion proteins comprising substrate channel elements and Cas7 subunits are constructed. These Cas7 fusion proteins are then assembled into a modified type I CRISPR-Cas effector complex (e.g., comprising fusions of Cse2, Cas5, Cas6, Cas7 substrate channel elements, and Cas8). In some embodiments, the crRNA of the modified type I CRISPR-Cas effector complex may be extended to accommodate additional Cas7 subunits (see, for example, Luo, M., et al., Nucleic Acids Res. 44:7385-7394 (2016)). Different substrate elements may be fused to Cas7 and then mixed at desired stoichiometry. When these various Cas7 subunits are assembled into a complete type I CRISPR-Cas effector complex, the co-localization of the substrate elements can enhance the efficacy of the substrate channel action.
[0311] In some implementations, an RNA scaffold is constructed that allows multiple Cas7 substrate channel element fusion complexes to bind in the absence of other type I CRISPR-Cas effector complex components.
[0312] The substrate channel element can be fused to the N-terminus and / or the C-terminus of the Cas7. Alternatively, a cyclic arrangement of Cas7s can be fused to the substrate channel element.
[0313] Figure 11A and Figure 11B The diagram illustrates a substrate channel composed of three successive enzymes in the pathway. The substrate channel facilitates the direct delivery of intermediate metabolites to the active sites of the enzymes in the metabolic pathway chain, without releasing them into additional channel space. Figure 11A A typical arrangement of the modified substrate channels is shown. Enzymes E1, E2, and E3 interact covalently or non-covalently with the scaffold protein matrix (S1, S2, S3). Double-headed arrows represent interactions between the enzymes and the scaffold proteins (e.g., affinity interactions). The substrate (X) is then processed into the product (Y) without being released into the additional channel space. Figure 11B One embodiment of the invention is shown, comprising a modified type I CRISPR-Cas effector complex carrying enzymes E1, E2, and E3 as fusion proteins (i.e., covalently interacting) with the Cas7 subunit protein, thereby generating a substrate channel. The cpCas7 protein and the backbone formed by the cpCas7 protein may also be useful in the practice of this aspect of the invention.
[0314] In other embodiments, substrate channel elements can be fused to Cas6. The Cas6 subunit of the Cascade complex recognizes specific RNA hairpin structures. RNA scaffolds composed of multiple cascaded Cas6 RNA hairpin structures can be constructed. Cas6 peptides from different Cascade complexes have different recognition sequences. Therefore, RNA scaffolds can be constructed from multiple orthogonal Cas6 RNA hairpins. By fusing different substrate channel elements to orthogonal Cas6 peptides, substrate channel complexes can be assembled in specific stoichiometry.
[0315] Substrate channel elements can be fused to the N-terminus and / or the C-terminus of the Cas6. Alternatively, cascaded Cas6 elements can be fused to the substrate channel elements.
[0316] In some implementations, the target heterologous metabolic pathway can be expressed in model organisms, such as *E. coli*. When a gene is heterologously expressed, codon optimization can be performed to express the gene more efficiently.
[0317] In one embodiment, the target metabolic pathway is the mevalonate pathway from Saccharomyces cerevisiae. The substrate channel elements of this pathway include, but are not limited to, acetyl-CoA-thioase (AtoB), hydroxymethylglutaryl-CoA synthase (HMGS), and hydroxymethylglutaryl-CoA reductase (HMGR).
[0318] In another embodiment, the target metabolic pathway is the glycerol synthesis pathway from Saccharomyces cerevisiae. The substrate channel elements of this pathway include, but are not limited to, glycerol 3-phosphate dehydrogenase (GPD1) and glycerol-3-phosphate phosphatase (GPP2).
[0319] In yet another embodiment, the target metabolic pathway is the starch hydrolysis pathway from Clostridium stercorarium. The substrate channel elements of this pathway include, but are not limited to, CelY and CelZ.
[0320] In another embodiment, the target metabolic pathway is the glucose phosphotransferase pathway from *E. coli*. The substrate channel elements of this pathway include, but are not limited to, trehalose-6-phosphate synthase (TPS) and trehalose-6-phosphate phosphatase (TPP).
[0321] In a seventh aspect, the present invention relates to the targeted recruitment of functional domains fused to Cascade subunit proteins to sites comprising a type II Cas9 protein and a nucleic acid-targeting nucleic acid (NATNA). These functional domains are disclosed herein, and include, but are not limited to, protein domains having enzymatic functions capable of transcriptional activation or transcriptional repression. Examples 13A and 13B describe methods for modifying type II CRISPR sgRNA, crRNA, tracrRNA, or crRNA and tracrRNA sequences with type I type I CRISPR repeat stem sequences to allow the recruitment of one or more Cascade subunit proteins to a type II CRISPR Cas protein / guide RNA complex binding site.
[0322] Figure 12A , Figure 12B and Figure 12C This shows a generalized view of the functional domains fused to the Cascade subunit protein being directionally recruited to the target site by the dCas9:NATNA complex site. It includes the spacer region sequence ( Figure 12A Type II CRISPRNATNA (101) Figure 12A ,102) through the linker nucleic acid sequence ( Figure 12A ,103) covalently linked to a type I CRISPR repeat stem sequence ( Figure 12A ,104). Covalently linked to a type I CRISPR repeat stem sequence ( Figure 12AThe type II CRISRP NATNA (105) can bind to type II dCas9 ( Figure 12A ,106) and type I Cascade subunit proteins (e.g., Cas6; Figure 12A ,107), which is achieved by connecting subsequences ( Figure 12A ,108) are fused to functional protein domains (e.g., enzyme domains, transcriptional activation or repression domains); Figure 12A ,109), thus forming an RNP complex. This RNP complex ( Figure 12B ,110) can target sequences containing the spacer region of type II cristonavirus (CRISPRNATNA) Figure 12A ,101) complementary target sequences ( Figure 12B Double-stranded DNA (112) Figure 12B ,111). Target recognition of the RNP complex leads to spacer sequence ( Figure 12A ,101) and target sequence ( Figure 12B Hybridization between , 112) Figure 12B , 113). Localizing Cascade subunit functional domain fusion proteins to DNA allows for DNA modification via functional domains of neighboring genes or transcriptional regulation. Figure 12C ,114).
[0323] In an eighth aspect, the present invention relates to compositions comprising a modified type I CRISPR-Cas effector complex, a modified guide polynucleotide, and combinations thereof. In some embodiments, the modified type I CRISPR-Cas effector complex comprises an associated Cas3 fusion protein. Wild-type type I CRISPR-Cas systems require the synergistic action of a Cascade effector complex for DNA targeting and a Cas3 helicase-nuclease for progressive DNA degradation. In one embodiment of the invention, the type I CRISPR-Cas effector complex is modified to prepare precise DSBs by fusing the complex to a nuclease domain (e.g., a nonspecific FokI endonuclease domain). This method uses paired guide polynucleotides that target two half-site DNA sequences separated by an intermediate sequence (i.e., a spacer interval).
[0324] Embodiments of this aspect of the invention relate to compositions comprising two modified type I CRISPR-Cas effector complexes, each of which comprises a spacer region and comprises a Cas subunit and a nuclease (e.g., FokI; see, for example) Figure 2A , Figure 2B and Figure 2C A fusion protein of the Cascade complex, wherein at least two parameters are varied to modulate genome editing efficiency. Such parameters include:
[0325] The length of the linker polypeptide used to generate the fusion protein containing the Cas subunit protein and a nuclease (e.g., FokI); and
[0326] The length of the spacer interval between nucleic acid target sequences that the spacer region can bind to.
[0327] This article provides guidance on amino acid compositions and sequence linker peptides.
[0328] One embodiment of this aspect of the invention is a composition comprising:
[0329] The first modified type I CRISPR-Cas effector complex comprises:
[0330] The first Cse2 subunit protein, the first Cas5 subunit protein, the first Cas6 subunit protein, and the first Cas7 subunit protein.
[0331] A first fusion protein comprising a first Cas8 subunit protein and a first FokI, wherein the N-terminus or C-terminus of the first Cas8 subunit protein is covalently linked to the C-terminus or N-terminus of the first FokI via a first linker polypeptide, and wherein the first linker polypeptide has a length of approximately 10 amino acids to approximately 40 amino acids.
[0332] Contains a first guide polynucleotide capable of binding to a first spacer region of a first nucleic acid target sequence; and
[0333] The second modified type I CRISPR-Cas effector complex comprises:
[0334] Second Cse2 subunit protein, second Cas5 subunit protein, second Cas6 subunit protein, and second Cas7 subunit protein.
[0335] A second fusion protein comprising a second Cas8 subunit protein and a second FokI, wherein the N-terminus of the second Cas8 subunit protein or the C-terminus of the second Cas8 protein is covalently linked to the C-terminus or N-terminus of the second FokI via a second linker polypeptide, and wherein the second linker polypeptide has a length of approximately 10 amino acids to approximately 40 amino acids.
[0336] The second guide polynucleotide contains a second spacer region capable of binding to a second nucleic acid target sequence, wherein the pre-spacer region adjacent motif (PAM) of the second nucleic acid target sequence and the PAM of the first nucleic acid target sequence have a spacer interval of about 20 to about 42 base pairs.
[0337] Examples of such a first modified type I CRISPR-Cas effector complex binding to a first nucleic acid target sequence and a second modified type I CRISPR-Cas effector complex binding to a second nucleic acid target sequence are shown in Figure 2A , Figure 2B and Figure 2C middle.
[0338] In some embodiments, the length of the first linker polypeptide and / or the second linker polypeptide is from about 15 amino acids to about 30 amino acids, or from about 17 amino acids to about 20 amino acids. In one embodiment, the first linker polypeptide and the second linker polypeptide are of the same length.
[0339] The first and second Cas8 subunits can each contain the same amino acid sequence as the Cas8 subunit.
[0340] Similarly, the first Cse2 subunit protein and the second Cse2 subunit protein may each contain the same amino acid sequence of the Cse2 subunit protein, the first Cas5 subunit protein and the second Cas5 subunit protein may each contain the same amino acid sequence of the Cas5 subunit protein, the first Cas6 subunit protein and the second Cas6 subunit protein may each contain the same amino acid sequence of the Cas6 subunit protein, the first Cas7 subunit protein and the second Cas7 subunit protein may each contain the same amino acid sequence of the Cas7 subunit protein, and combinations thereof.
[0341] Typically, the N-terminus of the first Cas8 subunit protein is covalently linked to the C-terminus of the first FokI via a first linker polypeptide, the C-terminus of the first Cas8 subunit protein is covalently linked to the N-terminus of the first FokI via a first linker polypeptide, the N-terminus of the second Cas8 subunit protein is covalently linked to the C-terminus of the second FokI via a second linker polypeptide, the C-terminus of the second Cas8 subunit protein is covalently linked to the N-terminus of the second FokI via a second linker polypeptide, and combinations thereof.
[0342] Embodiments of this aspect of the invention include embodiments in which the length of the spacer interval between the second nucleic acid target sequence and the first nucleic acid target sequence is about 22 base pairs to about 40 base pairs, about 26 base pairs to about 36 base pairs, about 29 base pairs to about 35 base pairs, or about 30 base pairs to about 34 base pairs.
[0343] The first FokI and the second FokI can be monomeric subunits that can combine to form homodimers, or different subunits that can combine to form heterodimers.
[0344] In a preferred embodiment, the guide polynucleotide comprises RNA.
[0345] In some implementations, the gDNA contains a PAM of a second nucleic acid target sequence and a PAM of a first nucleic acid target sequence.
[0346] In some embodiments, the modified type I CRISPR-Cas effector complex is based on a type I CRISPR-Cas effector complex selected from one or more organisms: *Salmonella enterica*, *Geothermal aerobicans* (strain EPR-M), *Methanococcus insica* MRE50, *Pseudomonas aeruginosa* (e.g., *Pseudomonas aeruginosa* (strain ND07), *Pseudomonas* S-6-2, and *Escherichia coli*). In a preferred embodiment, the modified type I CRISPR-Cas effector complex is based on a type I CRISPR-Cas effector complex of *Pseudomonas aeruginosa* (e.g., *Pseudomonas aeruginosa* (strain ND07), *Pseudomonas* S-6-2, and / or *Escherichia coli*). *Pseudomonas* S-6-2 induces ~10-fold higher editing efficiency compared to its *E. coli* homologue, and other homologues tested generally showed activity equivalent to that of *E. coli*, demonstrating that modified type I CRISPR-Cas effector complexes from different type I systems can be functionally used for genome editing in human cells.
[0347] The data shown in Examples 18A, 18B, 18C, 18D, 20A, 20B, and 20C demonstrate that altering the length of the linker polypeptide used to generate a fusion protein containing the Cas subunit protein and FokI and / or altering the spacing length between the spacer regions that the spacer regions can bind promotes the regulation of genome editing efficiency in cells.
[0348] In yet another embodiment, the present invention relates to a modified type I CRISPR-Cas effector complex comprising a first fusion protein containing a Cascade subunit protein (e.g., a Cas8 subunit protein) and a first functional domain (e.g., FokI), and a second fusion protein containing a dCas3* protein and a second functional domain (e.g., FokI). Figure 13A Cas7, Cas5, Cas8, Cse2, and Cas6; the dashed box around Cas6 indicates its interaction with the crRNA hairpin; cRNA is shown as a black line including the hairpin. Contains a first functional domain (e.g., FokI) ( Figure 13A A modified type I CRISPR-Cas effector complex (Cas8-linker 1-FP1 fusion) can bind to DNA and then recruit dCas3*-second domain (e.g., FokI) fusion proteins. Figure 13A dCas3*-connector 2-FP2). In which the first functional domain ( Figure 13ACas8-connector 1-FP1 fusion) and second functional domain ( Figure 13A In the case where the dCas3*-linker 2-FP2 contains a subunit of a dimer protein, the dCas3*-second domain (e.g., FokI) fusion protein binds to a modified type I CRISPR-Cas effector complex containing a first domain (e.g., FokI), promoting dimerization of the first and second domains. Figure 13A ). Figure 14A This shows a polypeptide containing a linker ( Figure 14A Linker 1) is linked to the Cas subunit protein ( Figure 14A The first functional area of the striped frame Figure 14A FD1) and the linker polypeptide associated with the Cascade complex ( Figure 14A Connector 2) connects to the second functional domain ( Figure 14A Modified type I CRISPR-Cas effector complex of dCas3* (FD2) Figure 14A The binding of dsDNA to the Cascade complex facilitates the proximity of FD1 and FD2 and promotes their interaction. The binding of the Cascade complex involves a single PAM sequence (…). Figure 14A PAM (hollow frame). In Figure 14A In the diagram, dsDNA is shown as paired horizontal dashed lines. In the case of a dimer endonuclease (e.g., FokI), proximity of FD1 and FD2 favors the formation of a functional dimer.
[0349] One advantage of the embodiments of the present invention is that, compared to using two FokI-Cascade complexes (which will... Figure 14A and Figure 2A , Figure 2B and Figure 2C (For comparison), a single Cascade complex (recognizing a single PAM sequence) can be used to cleave double-stranded nucleic acid target sequences. Using two FokI-Cascade complexes requires two properly oriented PAM sequences (…). Figure 2A , Figure 2B and Figure 2C This may limit the selection of proximal nucleic acid target sequences.
[0350] The length and / or composition of the linker polypeptide used to generate a fusion protein containing a Cas subunit protein and a nuclease (e.g., FokI), and the length and / or composition of the linker polypeptide used to generate a fusion protein containing a dCas3* protein and a nuclease, can be varied to modulate genome editing efficiency. Examples 21A, 21B, 21C, and 21D describe the design and testing of various Cas3-FokI linker compositions and lengths, as well as FokI-Cas8 linker compositions and lengths for modulating genome editing efficiency.
[0351] Another embodiment of this aspect of the invention includes a modified type I CRISPR-Cas effector complex ( Figure 13B Cas7, Cas5, Cas8, Cse2, and Cas6; the dashed box around Cas6 indicates its interaction with the crRNA hairpin; cRNA is shown as the black line containing the hairpin, and contains polypeptides via linkers ( Figure 13B dCas3* protein linked by a linker Figure 13B ,dCas3*) and functional domains ( Figure 13B A fusion protein containing a dCas3*-domain (e.g., cytidine deaminase). The modified type I CRISPR-Cas effector complex can bind to DNA and recruit dCas3*-domain (e.g., cytidine deaminase) fusion proteins. This implementation can facilitate site-specific targeting of nucleic acid target sequences for modification via or interaction with functional domains. In the case of cytidine deaminase, the modified type I CRISPR-Cas effector complex and the fusion protein containing dCas3* protein and cytidine deaminase can be used for site-specific base editing of nucleic acid target sequences. Figure 14B The modified type I CRISPR-Cas effector complex is shown. Figure 14B Examples of Cascade, which contain fusion proteins comprising linker polypeptides ( Figure 14B Connector) and functional domain ( Figure 14B ,FD) linked to dCas3* protein ( Figure 14B dCas3*), in which the complex binds to dsDNA ( Figure 14B (pairs of horizontal dashed lines). In Figure 14B In this process, the contact between the functional domain and dsDNA is facilitated. Binding of the Cascade complex involves a single PAM sequence ( Figure 14B PAM (hollow frame). Figure 14C The modified type I CRISPR-Cas effector complex is shown. Figure 14C Another example of Cascade is a fusion protein comprising a linker polypeptide ( Figure 14CConnector) and functional domain ( Figure 14C ,FD) linked to dCas3* protein ( Figure 14C dCas3*), in which the complex binds to dsDNA ( Figure 14C (Pairs of horizontal dashed lines). The binding of the Cascade complex involves a single PAM sequence ( Figure 14C PAM (hollow frame). In Figure 14C In this process, the contact between functional domains and ssDNA is promoted.
[0352] Additional functional domains and proteins that can be used to construct fusion proteins with type I CRISPR-Cas subunit proteins are described in this specification and examples. The length of the linker peptide composition and the Cas3-linker peptide-functional domain fusion protein can be evaluated by following the guidance of Examples 21A to 21D and this specification to assess the impact on functional domain performance.
[0353] Some embodiments of the present invention may use a modified type I CRISPR-Cas effector complex and the mCas3 protein, wherein the mCas3 protein contains downregulated helicase activity (e.g., mCas3 protein—a progressive mutant of Cas3 protein with reduced movement along DNA relative to wild-type type I CRISPR-Cas3 protein), or the mCas3 protein lacks helicase activity (e.g., the mCas3 protein is no longer a progressive nuclease such as wtCas3 protein, but the mCas3 protein retains cleavage activity). The modified type I CRISPR-Cas effector complex can bind to DNA and then recruit the mCas3 protein. This embodiment can facilitate site-specific cleavage of genomic DNA.
[0354] Table 48 describes a large number of mCas3 proteins, among which mutations in Cas3 proteins affect the ATP-binding / hydrolysis region of the helicase domain or the conserved region of the ssDNA pathway in the helicase domain. Figure 44 A linear plot of the functional domains of the EcoCas3 protein and the relative positions of mutations created within the Cas3 coding sequence are shown. Figure 44The data indicates the HD nuclease domain (amino acids 1-272), helicase domain (RecA1 region, amino acids 273-521; RecA2 region, amino acids 522-737), linker (amino acids 738-794), and C-terminal domain (CTD, amino acids 795-888). Huo, Y., et al., Nat. Struct. Mol. Biol. 9:771-777 (2014) disclosed the use of strain HD8 (Q53VY2; SEQ ID NO: 1869) from *Thermobifida fusca* (accession code: Q47PJ0; SEQ ID NO: 1869), *Saccharomonospora viridis* (C7MTA6; SEQ ID NO: 1870), *Thermobifora curvata* (D1A6Q2; SEQ ID NO: 1922), *Streptomyces avermitilis* (Q825B5; SEQ ID NO: 1925), *Streptomyces botropensis* (M3DI13; SEQ ID NO: 1923), and *Thermus thermophiles*. Sequence conservation analysis was performed by sequence alignment of the Cas3 family proteins from *E. coli* (NO: 1924; P38036; SEQ ID NO: 1844). Twenty-four different *EcoCas3* protein variants (Examples 23A to 23C) with mutations in the ATP-binding portion of the helicase domain or ssDNA loop-binding domain were screened. Seven mutants showed significantly more deletion categories and / or positional variations in the amplicon window; this finding supports the progressive reduction of mCas3 proteins relative to wtCas3.
[0355] Examples 23A to 23C describe such an mCas3 protein, wherein the average deletion induced by the mCas3 protein is shorter than the average deletion produced using the corresponding wtCas3 protein. Such an mCas3 protein can be used for genome editing (e.g., in human cells). Figure 45A , Figure 45B , Figure 45C and Figure 45DData indicating the mCas3 protein, when introduced and expressed in human cells, shows a shorter mean deletion length compared to the wtCas3 protein associated with the Cascade RNP complex when associated with the Cascade RNP complex. Referring to the teachings of this specification, those skilled in the art can perform similar mutations in the corresponding regions of Cas3 proteins obtained from bacterial species other than *E. coli*.
[0356] Examples 26A through 26C provide further examples of mCas3 proteins that can be used to generate genomic deletions, wherein the average mCas3 protein-induced deletion is shorter than the average deletion generated using the corresponding wtCas3 protein. The data shown in the examples support the use of an ATPase / helicase-deficient variant of Cas3 from Pseudomonas S-6-2 (mPseCas3 protein) in conjunction with the PseCascade RNP complex to generate deletions at the intended cleavage site (i.e., deletions localized to the cleavage site).
[0357] The activity of the wtPseCas3 protein / PseCascade was further characterized. Additional experiments were performed using target-enriched probes that enable the detection of large genomic deletions. Specifically, HEK293 cells were transfected with a DNA template encoding the PseCascade RNP complex, the wtPseCas3 protein, and a minimal CRISPR array targeting the TRAC site, essentially as described in Examples 26A through 26C. Target-enriched probes were used to isolate and sequence genomic fragments; while in Example 26C, an amplicon window was used to identify the presence of deletions. The target enrichment / sequencing approach provides unbiased observation of larger deletions that cannot be provided by using an amplicon window to identify deletions. Overall, deletions assessed using target enrichment and genomic fragment sequencing were found to be largely unidirectional, starting upstream of the wtPseCas3 protein initiation site. Deletions ranged from 1 bp to nearly 250 kb. In addition to providing a method for cutting genomic DNA and providing deletions of a given length, this method can also be used to generate large, random subsets of deletions at defined locations to probe regulatory / promoter regions of genes.
[0358] The mCas3 protein can contain one or more mutations (e.g., combinations of mutations as described in Table 48).
[0359] Control of deletion length for several mCas3 proteins has been demonstrated. In some embodiments, the mCas3 protein of the present invention associated with the Cascade complex containing the guide polynucleotide can provide an average deletion length of about 1 to about 600 base pairs, about 1 to about 500 base pairs, about 1 to about 400 base pairs, about 1 to about 300 base pairs, preferably about 1 to about 250 base pairs, about 1 to about 200 base pairs, or about 1 to about 100 base pairs.
[0360] In some embodiments, the wtCas3 or mCas3 protein can be fused to various subunits of the Cascade complex to further control the average Cas3 deletion length. Constraining the Cascade complex can restrict or prevent the movement of the Cas3 or mCas3 protein along the DNA because it will be fixed to the Cascade complex binding site. Typically, the wtCas3 or mCas3 protein can be fused to the N-terminal or C-terminal domain of the protein component of the Cascade complex using a linker polypeptide (e.g., for the EcoCascade complex, fusion can be with EcoCas8, EcoCas6, or EcoCas5). An NLS sequence can also be attached to the N-terminus of the fusion protein. Examples of such constructs of E. coli Cascade protein components are shown in Table 12. These EcoCas3 fusion proteins also have an NLS sequence attached to their N-terminus.
[0361]
[0362]
[0363] *The protein sequence is the sequence encoding a polycistronic protein.
[0364] Embodiments of the present invention include a modified type I CRISPR mCas3 protein capable of reducing migration along DNA relative to the wild-type type I CRISPR Cas3 protein (wtCas3 protein). In some embodiments, the mCas3 protein contains about 90% or more, preferably about 95% or more, more preferably about 98% or more sequence identity with the corresponding wtCas3 protein. The coding sequence of the mCas3 protein may contain a covalently linked nuclear localization signal at the N-terminus, C-terminus, or both. The mCas3 protein may contain one or more mutations that downregulate helicase activity, wherein the modified mCas3 protein retains nuclease activity (or at least a portion thereof) relative to the corresponding wtCas3 protein. Typically, the DNA is dsDNA containing a target region containing a nucleic acid target sequence. When the wtCas3 protein is associated with the corresponding Cascade nucleoprotein complex (“Cascade NP complex / wtCas3 protein”; for example, the Cascade RNP complex), and the Cascade NP complex contains a guide with a spacer region complementary to the nucleic acid target sequence, the binding of the Cascade NP complex / wtCas3 protein to the nucleic acid target sequence favors cleavage in the DNA target region, typically resulting in deletion in the target region; and when the mCas3 protein is associated with the Cascade NP complex (“Cascade NP complex / mCas3 protein”; for example, the Cascade RNP complex / mCas3 protein) and binds to the nucleic acid target sequence, it favors cleavage in the DNA target region and results in a shorter average deletion length relative to the average deletion length of wtCas3.
[0365] In some embodiments, one or more mutations in the mCas3 protein are amino acid substitutions relative to the wtCas3 protein. In other embodiments, one or more deletions include amino acid deletions or insertions in the mCas3 protein coding sequence relative to the wtCas3 protein. One or more mutations may be located in the RecA1 or RecA2 region of the helicase domain. In one embodiment, one or more mutations downregulate the binding of the mCas3 protein to ssDNA relative to the wtCas3 protein (e.g., mutations affecting ssDNA loop binding and / or mutations in conserved regions of the ssDNA pathway in the helicase domain). In another embodiment, one or more mutations downregulate ATP hydrolysis via the mCas3 protein relative to the wtCas3 protein, or downregulate ATP binding to the mCas3 protein relative to the wtCas3 protein. In other embodiments, the mCas3 protein comprises a combination of one or more mutations that downregulate the binding of the mCas3 protein to ssDNA relative to the wtCas3 protein, downregulate ATP hydrolysis via the mCas3 protein relative to the wtCas3 protein, or downregulate ATP binding to the mCas3 protein.
[0366] Other embodiments include the coding sequence of an mCas3 protein covalently linked to the amino or carboxyl terminus of a Cas protein coding sequence to a Cascade nucleoprotein complex (e.g., the Cascade RNP complex). Such a Cas protein may be selected from Cse2, Cas8, Cas7, Cas6, and Cas5 proteins.
[0367] In some embodiments, the wtCas3 protein is an Escherichia coli type 1 CRISPRCas3 protein. In other embodiments, the wtCas3 protein is a wtCas3 protein selected from Pseudomonas S-6-2, Thermolyticus brownii, Saccharomyces viride, Curvularia thermomonas, Streptomyces avermanii, Streptomyces bozzolani, Thermophilus thermophilus, Vibrio cholerae, Salmonella enterica, Geothermal bacillus EPR-M, Methanogens inoculum MRE50, and Pseudomonas aeruginosa (strain ND07).
[0368] For the E. coli type 1 CRISPRwtCas3 protein, one or more mutations may include, but are not limited to, D452H, A602V, or D452H and A602V.
[0369] In other embodiments, the cell contains DNA, and the cell may be a eukaryotic cell (e.g., a human cell).
[0370] In another embodiment, the present invention includes a polynucleotide containing a coding sequence of the mCas3 protein, an expression cassette containing a coding sequence of the mCas3 protein, a plasmid containing a coding sequence of the mCas3 protein, and a Cascade nucleoprotein complex containing the mCas3 protein.
[0371] In a ninth aspect, the present invention relates to a method using a modified type I CRISPR-Cas effector complex.
[0372] In some embodiments, the present invention includes a method for binding a nucleic acid target sequence in a polynucleotide (e.g., dsDNA), comprising providing one or more modified type I CRISPR-Cas effector complexes for induction into a cellular or biochemical reaction, and introducing the modified type I CRISPR-Cas effector complexes into the cellular or biochemical reaction to facilitate contact between the modified type I CRISPR-Cas effector complexes and the polynucleotide. Contact between the complexes and the polynucleotides results in the binding of the modified type I CRISPR-Cas effector complexes to the nucleic acid target sequence in the polynucleotide.
[0373] In one implementation, the modified type I CRISPR-Cas effector complex includes a guide complementary to a nucleic acid target sequence in a polynucleotide. The modified type I CRISPR-Cas effector complex binds to the nucleic acid target sequence in the polynucleotide.
[0374] In other embodiments, the first modified type I CRISPR-Cas effector complex includes a guide complementary to a first nucleic acid target sequence in the polynucleotide, and the second modified type I CRISPR-Cas effector complex includes a guide complementary to a second nucleic acid target sequence in the polynucleotide. The first modified type I CRISPR-Cas effector complex binds to the first nucleic acid target sequence, and the second modified type I CRISPR-Cas effector complex binds to the second nucleic acid target sequence in the polynucleotide.
[0375] In yet another embodiment, the modified type I CRISPR-Cas effector complex includes a guide complementary to a nucleic acid target sequence in a polynucleotide, and also includes a dCas3* fusion protein capable of binding to the complex. The modified type I CRISPR-Cas effector complex binds to the nucleic acid target sequence in the polynucleotide, and the effector complex includes a dCas3* fusion protein that binds to the complex.
[0376] Such methods of binding nucleic acid target sequences can be performed in vitro (e.g., in a biochemical reaction or in cultured cells; in some embodiments, the cultured cells are human cultured cells that remain in the culture and are not introduced into the human body); in vivo (e.g., in the cells of a living organism, with the condition that, in some embodiments, the organism is a non-human organism); or ex vivo (e.g., cells removed from a subject, with the condition that, in some embodiments, the subject includes a human subject, and in other embodiments, the subject is a non-human subject).
[0377] Various methods are known in the art for assessing and / or quantifying interactions between nucleic acid sequences and nucleotides, including, but not limited to, immunoprecipitation (ChIP) assays, DNA electrophoretic mobility shift assays (EMSA), DNA pull-down assays, and microplate capture and detection assays. Commercial kits, materials, and reagents are available for performing many of these methods and are, for example, from vendors such as Thermo Scientific (Wilmington, DE), Signosis (Santa Clara, CA), Bio-Rad (Hercules, CA), and Promega (Madison, WI). A common method for detecting interactions between peptide and nucleic acid sequences is EMSA (see, for example, Hellman LM, et al., Nature Protocols 2:1849-1861 (2007)).
[0378] In another embodiment, the invention includes a method for cleaving a nucleic acid target sequence in a polynucleotide (e.g., a single-strand cleavage or a double-strand cleavage in dsDNA), comprising providing one or more modified type I CRISPR-Cas effector complexes for use in a cellular or biochemical reaction, and introducing the modified type I CRISPR-Cas effector complexes into the cellular or biochemical reaction to facilitate contact between the modified type I CRISPR-Cas effector complexes and the polynucleotide.
[0379] In one embodiment, a first modified type I CRISPR-Cas effector complex comprising a guide and a first nuclease domain (e.g., FokI) complementary to a first nucleic acid target sequence in a polynucleotide is included. Figure 15A Cascade1, a box outlined in solid line, linked by a linker polypeptide, a black curve, to the first nuclease domain, represented as a circular sector), and a second modified type I CRISPR-Cas effector complex containing a guide and a second nuclease domain (e.g., FokI) complementary to a second nucleic acid target sequence in a polynucleotide. Figure 15A Cascade 2 (the dashed outline of the box, linked by the linker polypeptide, the black curve, to the second nuclease domain, represented as a circular sector) is introduced into cells or biochemical reactions. The first modified type I CRISPR-Cas effector complex ( Figure 15B Cascade1) binds to dsDNA ( Figure 15B The first nucleic acid target sequence in dsDNA is represented by paired horizontal black lines, and the first nuclease domain cleaves the first strand of dsDNA. Figure 15C Cascade1), and the second modified type I CRISPR-Cas effector complex ( Figure 15BThe modified type I CRISPR-Cas effector complex binds to the second nucleic acid target sequence in dsDNA, and the second nuclease domain cleaves the second strand of dsDNA. Binding of the modified type I CRISPR-Cas effector complex results in the cleavage of the nucleic acid target sequence in the polynucleotide (e.g., dsDNA) by the modified type I CRISPR-Cas effector complex.
[0380] In another embodiment, a first modified type I CRISPR-Cas effector complex comprising a guide complementary to a first nucleic acid target sequence in a polynucleotide, a second modified type I CRISPR-Cas effector complex comprising a guide complementary to a second nucleic acid target sequence in a polynucleotide, and a Cas3 cleavage enzyme (e.g., an ATPase-deficient Cas3 variant with only cleavage enzyme activity) are introduced into a cell or biochemical reaction. The first modified type I CRISPR-Cas effector complex binds to the first nucleic acid target sequence in dsDNA, the Cas3 cleavage enzyme protein is associated with the first complex, and cleaves the first strand of dsDNA; and the second modified type I CRISPR-Cas effector complex binds to the second nucleic acid target sequence in dsDNA, the Cas3 cleavage enzyme protein is associated with the second complex, and cleaves the second strand of dsDNA. The binding of the modified type I CRISPR-Cas effector complex to the associated Cas3 cleavage enzyme protein results in the cleavage of the nucleic acid target sequence in the polynucleotide (e.g., dsDNA) by the modified type I CRISPR-Cas effector complex. Data presented in Examples 25A, 25B, and 25C demonstrate that the Cascade RNP complex, containing a Cas3 ATPase-deficient mutant protein, can induce targeted genomic deletions via paired nicks. These paired nicks can promote targeted deletions in the host cell (e.g., human cell) genome.
[0381] In another embodiment, a modified type I CRISPR-Cas effector complex comprising a guide and a first nuclease domain (e.g., FokI) complementary to the nucleic acid target sequence in the polynucleotide is used. Figure 16A Cascade; a dashed outline of a box connected by a linker polypeptide, a black curve leading to the first nuclease domain (represented as a circular sector), and a dCas3*-second nuclease domain (e.g., FokI) fusion protein capable of binding to the complex. Figure 16A dCas3; the solid-line outline of the box, linked by the linker polypeptide, the black curve, to the second nuclease domain, represented as a circular sector) is introduced into cells or biochemical reactions. The modified type I CRISPR-Cas effector complex ( Figure 16B Cascade) binds to the nucleic acid target sequence in dsDNA ( Figure 16B (paired horizontal black lines) and cut the first strand of dsDNA ( Figure 16CCascade), and the dCas3* fusion protein is associated with the Cascade RNP complex ( Figure 16B ,dCas3*), and cleave the second strand of dsDNA ( Figure 16C ,dCas3*).
[0382] In other embodiments, a modified type I CRISPR-Cas effector complex comprising a guide complementary to a target region containing a nucleic acid target sequence in a polynucleotide and a Cas3 protein (e.g., Cas3 protein or mCas3 protein) capable of binding to the complex is introduced into a cellular or biochemical reaction. The modified type I CRISPR-Cas effector complex binds to the nucleic acid target sequence in dsDNA, the Cas3 protein (e.g., Cas3 protein or mCas3 protein) is associated with the complex, and cleaves at least one strand of the dsDNA in the target region. In some embodiments, cleavage of the dsDNA by the mCas3 protein results in a deletion in the dsDNA target region. This method can be used to create long-range deletions of a specific length and can be used to generate gene knockout or knock-in. In some embodiments, the Cas3 protein (e.g., Cas3 protein or mCas3 protein) can be fused to a Cascade complex subunit protein (e.g., Cas7 protein, Cas8 protein, Cas5 protein, Cse2 protein). Examples 23A through 23C describe embodiments of the mCas3 protein.
[0383] In another embodiment, the present invention relates to the use of a type I CRISPR-Cas effector complex, wherein a nuclease domain is fused to a Cascade complex protein (see, for example, Example 11A, Table 38) or a dCas3* protein (e.g., a dCas3* protein fused to a DNase) to delete a nucleic acid target sequence. This method can be used to create cuts and deletions in dsDNA target regions and can be used to generate gene knockouts. In some embodiments, the nuclease domain can be fused to a Cascade complex subunit protein such as Cas7, Cas8, Cas5, or Cse2.
[0384] Methods for cleaving nucleic acid target sequences in polynucleotides may also include introducing donor polynucleotides into cells to facilitate the incorporation of at least a portion of the donor polynucleotide into the cell's gDNA.
[0385] Figure 17A This demonstrates a guide that can be included to complement the first nucleic acid target sequence in a polynucleotide. Figure 17A The first modified type I CRISPR-Cas effector complex of the first nuclease domain (e.g., Cascade1) and the first nuclease domain (e.g., FokI) Figure 17AThe linker polypeptide, shown as a curved line connecting Cascade1 and a gray circular fan, and a guide containing a second nucleic acid target sequence complementary to the polynucleotide. Figure 17A The second modified type I CRISPR-Cas effector complex of the second nuclease domain (e.g., Cascade 2) and the second nuclease domain (e.g., FokI) Figure 17A The linker polypeptide, shown as a curved line connecting Cascade2, and gray circular fan-shaped segments, represents the two strands of dsDNA cleaved together. Figure 17A Examples of (pairs of black horizontal lines). Figure 17B This illustrates a homologous arm containing a DNA sequence complementary to the adjacent double-strand cleavage site. Figure 18B Donor polynucleotide (donor, dashed line) Figure 17B (The paired dashed lines are shown above Cascade2). Figure 17C This illustrates a portion of the donor polynucleotide incorporated into the double-strand cleavage site region. Figure 17C (The paired dashed lines connecting the paired black horizontal lines representing dsDNA). Donor polynucleotide incorporation is mediated by cellular DNA repair mechanisms (e.g., HDR). Figures 17B to 17C (The arrow points downwards; the vertical arrow represents the cellular DNA repair mechanism).
[0386] In other embodiments, a modified type I CRISPR-Cas effector complex comprising a guide and a first nuclease domain complementary to a first nucleic acid target sequence in a polynucleotide can pair with a second component comprising a second nuclease domain, wherein the second component is capable of binding to a second nucleic acid target sequence of the polynucleotide. Examples of such a second component include transcription activator-like effector nucleases (TALENs) comprising a second nuclease domain, zinc finger nucleases (ZFNs) comprising a second nuclease domain, or dCas9 / NATNA complexes comprising a second nuclease domain.
[0387] In one implementation, a combination of a Cascade complex containing a guide complementary to a first nucleic acid target sequence in the target polynucleotide and a dCas9 / NATNA complex can be used to delete a region of a target polynucleotide (e.g., gDNA), wherein NATNA contains a spacer sequence complementary to a second nucleic acid target sequence in the target polynucleotide. The first and second nucleic acid target sequences are selected to be located flanking the nucleic acid target sequence targeted for deletion. A Cas3 protein containing active endonuclease activity binds to the Cascade complex, and then progressively deletes a single strand of dsDNA containing the nucleic acid target sequence targeted for deletion. When the Cas3 protein collides with the dCas9 / NATNA complex (i.e., a “roadblock”), the Cas3 nuclease activity can be terminated at the second nucleic acid target sequence by the dCas9 / NATNA complex. Figures 21A-21DAn example of Cas3 deletion in a nucleic acid target sequence is shown. Figure 21A The diagram shows nucleic acid target sequence 1, which contains nucleic acid target sequences located on both sides of the nucleic acid target sequence targeted for deletion. Figure 21A NATS1) and nucleic acid target sequence 2 ( Figure 21A dsDNA of NATS2 Figure 21A (pairs of horizontal black lines). Figure 21A The Cascade complex, which contains a wizard complementary to NATS1, is shown. Figure 21A Cascade; a rectangle with a black outline), Cas3 protein ( Figure 21A Cas3; gray circular sector) and the dCas9 / NATNA complex containing a spacer region complementary to NATS2 ( Figure 21A dCas9; a rectangle with a dashed frame. Figure 21B The binding of the Cascade complex to NATS1, the association between the Cas3 protein and the Cascade complex, and the binding of the dCas9 / NATNA complex to NATS2 were shown. Figure 21C This demonstrates progressive deletion of Cas3 by targeting the single-stranded nucleic acid target sequence for deletion. Figure 21D The dissociation of the Cas3 protein from dsDNA at the site of binding to the dCas9 / NATNA complex of NATS2 is shown. The data illustrated in Examples 24A-24D support the use of protein roadblocks to control the length of deletions mediated by the Cas3 protein associated with the Cascade nucleoprotein complex; thus providing a method for promoting the formation of deletions of defined length in gDNA of cells (e.g., human cells) using the Cas3 protein associated with the Cascade nucleoprotein complex.
[0388] In another embodiment, a combination of a first Cascade complex containing a guide complementary to a first nucleic acid target sequence in the target polynucleotide and a second Cascade complex containing a guide complementary to a second nucleic acid target sequence in the target polynucleotide can be used to delete a region of a target polynucleotide (e.g., gDNA). The first and second nucleic acid target sequences are selected to be located flanking the nucleic acid target sequence targeted for deletion. A Cas3 protein containing active endonuclease activity binds to each Cascade complex and then progressively deletes both strands targeting the nucleic acid target sequence targeted for deletion. Cas3 nuclease activity can be terminated at the first and second nucleic acid target sequences via the Cascade complex when each Cas3 protein collides with one Cascade complex. Figures 22A-22D An example of Cas3 deletion in both strands of a nucleic acid target sequence is shown. Figure 22A The diagram shows nucleic acid target sequence 1, which is located on either side of the nucleic acid target sequence targeted for deletion. Figure 22ANATS1) and nucleic acid target sequence 2 ( Figure 22A dsDNA of NATS2 Figure 22A (Pairs of horizontal black lines). Figure 22A The first Cascade complex is shown, which contains a wizard complementary to NATS1. Figure 22A Cascade1; a rectangle with a black outline), Cas3 protein ( Figure 22A Cas3; gray circular fan), and a second Cascade complex containing a guide complementary to NATS2 ( Figure 22A , Cascade2; a rectangle with a dashed frame. Figure 22B The binding of the Cascade complex to NATS1 and NATS2, as well as the association between the Cas3 protein and the Cascade complex, were shown. Figure 22C The progressive deletion is shown to result from movement along the DNA and degradation by Cas3 nucleases targeting the two strands of the nucleus for the deleted nucleic acid target sequence. Figure 22D The dissociation of Cas3 protein from dsDNA is shown at the location of the Cascade complex that binds to NATS1 and NATS2.
[0389] In other embodiments, the Cascade complex can be modified to prevent it from binding to the Cas3 protein, and such a modified Cascade complex can substantially bind to... Figures 21A-21D The same mechanism shown is used as a roadblock to prevent progressive DNA degradation by catalytically activating Cas3 associated with the Cascade RNP complex. Other site-specific binding proteins (e.g., transcription activator-like effectors (TAL) or zinc finger (ZnF) DNA-binding proteins) can act as roadblocks in a similar manner.
[0390] In some embodiments, the nucleic acid target sequence is dsDNA (e.g., genomic) DNA. In some embodiments, the nucleic acid target sequence is double-stranded, and one or both strands are cleaved. Such methods of cleaving nucleic acid target sequences can be performed in vitro, in vivo, or ex vivo.
[0391] As described above, in some embodiments, the present invention relates to introducing one or more modified type I CRISPR-Cas effector complexes into host cells to facilitate the cleavage of a nucleic acid target sequence in dsDNA in the presence of a donor polynucleotide, wherein one or more modified type I CRISPR-Cas effector complexes generate a cleavage site (or a cleavage site and associated deletion) in a target region containing the nucleic acid target sequence of the host cell DNA, thereby facilitating the insertion of at least a portion of the donor polynucleotide into the target region. In some embodiments, the cleavage site is a double-strand break in the target region (e.g., when using two modified type I CRISPR-Cas effector complexes of each containing a spacer region and a fusion protein containing a Cas protein and a nuclease (e.g., FokI) or two modified type I CRISPR-Cas effector complexes of each containing a spacer region associated with a Cas3 protein or mCas3 protein). In some embodiments, the cleavage site is a single-strand break in the target region (e.g., when using a type I CRISPR-Cas effector complex associated with an mCas3 protein). In other implementations, the cleavage site is a deletion in the target region (e.g., when using a type I CRISPR-Cas effector complex associated with the Cas3 or mCas3 protein).
[0392] To demonstrate homology-directed repair (HDR), a minimal CRISPR array was designed to target the FokI-PseCascade RNP complex at four sites (WDR92, B2M, CCR5, and TRAC) in the human genome. Using three oligonucleotides (SEQ ID NO: 1513 to SEQ ID NO: 1515; Example 20A) and unique primers encoding a “repeat-spacer-repeat-spacer-repeat” sequence, a minimal CRISPR array was generated using PCR-based assembly. The first and second spacers guide the FokI-PseCascade RNP complex to adjacent nucleic acid target sequences, enabling FokI dimerization and genome cleavage (i.e., generating cleavage sites).
[0393] For each HDR insertion site in the target region containing the cleavage site—in this case overlapping with the cleavage site—cells are transfected with: 3 μg of a vector encoding a FokI-PseCascade complex protein component containing a fused FokI fused to the N-terminus of Cas8 with NLS linked to the N-terminus of FokI, 150 ng of a minimal CRISPR array, and 0–60 pmol of a single-stranded oligodeoxynucleotide (ssODN) template donor polynucleotide for HDR. The ssODN contains homologous arms, each 75 nucleotides long, and the two arms are symmetrically positioned around the cleavage site. The donor polynucleotide also contains a phosphate thioester bond at the 3' end of the homologous arm to reduce or prevent cellular degradation of the donor polynucleotide. At the 5' end of the phosphate thioester bond, the donor polynucleotide also contains a “TAATAAT” insertion sequence to insert two stop codons and increase the interstitial spacing in the repaired chromosome, thereby preventing re-cleavage of the FokI-PseCascade RNP complex.
[0394] Transfection was performed in HEK293 cells essentially as described in Example 20B, except that ssODN was included in the mixture to enable HDR. Several days after transfection, gDNA was purified from the cells, treated with an exonuclease to remove any residual ssODN that could contaminate subsequent PCRs, and then used as a template for amplification to measure donor insertion. Deep sequencing analysis was performed essentially as described in Example 20C. Table 13 shows the percentage of mutant reads in the total reads from this experiment (the first column is pmol of ssODN):
[0395]
[0396] The percentage of mutant reads indicates the mutant reads that include insertion-deletion mutations that result from non-homologous end-joins and the insertion of the “TAATAAT” HDR sequence.
[0397] Table 14 shows the percentage of HDR reads containing only the "TAATAAT" insertion sequence among the total mutant reads from this experiment (the first column is the ssODN of pmol):
[0398]
[0399] As can be seen from the data, dsDNA cleavage via the Cascade RNP complex enables HDR and incorporation of donor polynucleotides at multiple loci in the human genome.
[0400] In yet another embodiment, the invention includes a method for modifying one or more nucleic acid target sequences in a polynucleotide (e.g., DNA) during a cellular or biochemical reaction, comprising providing one or more modified type I CRISPR-Cas effector complexes (e.g., cas subunit protein-cytidine deaminase fusion proteins) for introduction into the cellular or biochemical reaction, and introducing the modified type I CRISPR-Cas effector complexes into the cellular or biochemical reaction thereby promoting contact between the modified type I CRISPR-Cas effector complexes and the polynucleotide, resulting in binding of the modified type I CRISPR-Cas effector complexes to the nucleic acid target sequences in the polynucleotides in a manner favorable to mutations in the nucleic acid target sequences (e.g., C to T, G to A, A to G, and T to C). Figures 18A-18D This demonstrates the use of the Cascade complex, which contains a Cas subunit protein-linker polypeptide-cytidine deaminase fusion protein (Cascade / CD complex), to mutate target nucleotides in cellular gDNA. Figure 18A Examples of paired black horizontal lines ("C" for cytosine and "G" for guanine). The Cascade / CD complex ( Figure 18A The "Cascade" with a linker polypeptide is shown as a curved line connecting Cascade and cytidine deaminase ("CD" is indicated by a gray circular fan shape) introduced into the cell. The Cascade / CD complex contains a guide (…) complementary to the DNA target sequence of the adjacent target cytosine. Figure 18B (“C”). In Figure 18B In this process, the Cascade / CD complex binds to the DNA target sequence, and cytidine deaminase converts cytosine (… Figure 18B "C" is converted to cytosine ( Figure 18C Then, the cell repair mechanism can repair cytosine into thymine and change mismatched guanidine into adenine (“U”). Figures 18C-18D (The arrow points downwards; the vertical arrow represents the cellular DNA repair mechanism).
[0401] In yet another embodiment, the invention includes a method for regulating in vitro or in vivo transcription, for example, the transcription of a gene containing a regulatory element sequence. Such a method includes providing one or more modified type I CRISPR-Cas effector complexes (e.g., Cas subunit protein-transcription factor fusion proteins) for introduction into a cell or biochemical reaction, and introducing the modified type I CRISPR-Cas effector complex into the cell or biochemical reaction to facilitate contact between the modified type I CRISPR-Cas effector complex and the regulatory element sequence, resulting in the binding of the modified type I CRISPR-Cas effector complex to the regulatory element sequence, thereby facilitating the regulation of in vitro or in vivo transcription of a gene containing the regulatory element sequence.
[0402] Figure 19A and Figure 19B A general illustration of an example of transcriptional activation of a universal gene (“gene 1”) is shown. Figure 19A This provides an overview of transcriptional regulation of endogenous genes in eukaryotic cells. Figure 19A In the image, two parallel black lines represent double-stranded DNA, indicating gene 1 (…). Figure 19A The location of gene 1, and the transcription start site associated with gene 1 ( Figure 19A (TSS). Figure 19A The first image shows the transcription factors required for transcriptional activation of gene 1. Figure 19A TF) and polymerase II ( Figure 19A Pol II is shown as not yet associated with gene 1-TSS. The second figure shows the association of TF with its homologous TSS. Then TF recruits transcription activating proteins ( ). Figure 19A ,TP), which subsequently recruits RNA polymerase II ( Figure 19A Pol II). Typically, in eukaryotes, TF factors and TP form complexes containing a variety of proteins and possibly other molecules. The third figure illustrates the transcription of gene 1 via Pol II. Figure 19A (The curved arrow at the end of gene 1 indicates the direction of transcription). This type of transcriptional activation typically depends on the TF specific to gene expression. Figure 19B The illustration shows one embodiment of the invention, in which the Cascade complex is modified to include a protein or factor (…). Figure 19B CASCADEa), which attracts one or more components (transcription activators) responsible for transcriptional activation in cells; Figure 19B ,TA). An example of such a protein or factor is the protein VP64. CASCADEa contains a guide capable of binding at or near the TSS ( ). Figure 19B (TSS). Figure 19B In the image, two parallel black lines represent double-stranded DNA, indicating gene 1 (…). Figure 19B The location of gene 1, and the transcription start site (TSS) associated with gene 1. Figure 19B In the first image, CASCADEa and polymerase II ( Figure 19B Pol II is shown as not yet associated with gene 1-TSS. The second figure shows the association between CASCADEa and its target TSS. CASCADEa then recruits transcriptional activating proteins (…). Figure 19B TA), which then recruits RNA polymerase II ( Figure 19B The third figure shows the transcription of gene 1 produced by Pol II. Figure 19B(The curved arrow at the end of gene 1 indicates the direction of transcription). One advantage of this embodiment of the invention is that the transcriptional activation of the gene does not depend on endogenous transcription factors that bind to the gene's TSS, but can be targeted to the gene's TSS by selecting an appropriate Cascade guide.
[0403] Figure 20A and Figure 20B This demonstrates the use of a fusion of the Cas subunit protein-KRAB domain and a regulatory sequence associated with gene 1. Figure 20A (promoter) complementary guide ( Figure 20A The Cascade complex, having a conjoint polypeptide (shown as a curved line connecting Cascade and a circular element representing the KRAB domain), is effective for universal genes ( Figure 20A A general illustration of an example of transcriptional repression of gene 1. The binding of CASCADEi to regulatory sequences ( Figure 20B This leads to transcriptional repression of gene 1. Figure 20B (The black line ending with an X represents transcriptional repression).
[0404] The modified type I CRISPR-Cas effector complex as described herein can be integrated into the kit. In some embodiments, the kit includes a package having a container holding one or more kit elements, said kit elements as one or more separate compositions, or, optionally, as a mixture if compatibility of the components allows. In some embodiments, the kit also contains one or more excipients of: buffers, buffer reagents, salts, sterile aqueous solutions, preservatives, and combinations thereof. Exemplary kits may contain one or more modified type I CRISPR-Cas effector complexes and one or more excipients, or one or more nucleic acid sequences encoding one or more components of a modified type I CRISPR-Cas effector complex.
[0405] In addition, the kit may include instructions for using the modified type I CRISPR-Cas effector complex composition.
[0406] Another aspect of the invention relates to a method for preparing or producing one or more modified type I CRISPR-Cas effector complexes or components thereof. In one embodiment, the preparation or production method includes generating the modified type I CRISPR-Cas effector complex in cells and purifying the modified type I CRISPR-Cas effector complex from cell lysate.
[0407] Modified type I CRISPR-Cas effector complex compositions may also contain detectable tags, such as portions that can provide a detectable signal. Examples of detectable tags include, but are not limited to, enzymes, radioisotopes, multiple specific binding pairs, fluorophores (FAMs), fluorescent proteins (green fluorescent protein (GFP), red fluorescent protein, mCherry, tdTomato), DNA or RNA aptamers with suitable fluorophores (enhanced GFP (eGFP), "Spinach"), quantum dots, antibodies, etc. A wide variety of suitable detectable tags are well known to those skilled in the art.
[0408] In some embodiments, the modified type I CRISPR-Cas effector complex (i.e., nucleoprotein particles) can be introduced into cells by methods including, but not limited to, nuclear transfection, gene gun delivery, sonar perforation, cell extrusion, lipid transfection, or the use of other chemicals, cell-penetrating peptides, etc. In other embodiments, a vector system, an expression cassette containing a DNA sequence encoding one or more components, and one or more RNA molecules (e.g., mRNA) containing an expression cassette encoding an RNA sequence encoding one or more components can be used to introduce the modified type I CRISPR-Cas effector complex and the coding sequences of one or more components of an associated protein into cells.
[0409] One embodiment of the present invention relates to the generation of recombinant cells (e.g., modified lymphocytes) using a modified type I CRISPR-Cas effector complex. The method typically includes facilitating the contact of dsDNA containing a target region of a host cell containing a nucleic acid target sequence with one or more modified type I type I CRISPR-Cas effector complexes of the present invention. Contact between the modified type I type I CRISPR-Cas effector complex and the nucleic acid target sequence results in the modified type I type I CRISPR-Cas effector complex binding to the target region containing the nucleic acid target sequence, cleaving the target region containing the nucleic acid target sequence, and modifying the dsDNA in the target region, thereby generating recombinant cells. In some embodiments, the dsDNA contains more than one nucleic acid target sequence, and a modified type I type I CRISPR-Cas effector complex containing a spacer region sequence complementary to each nucleic acid target sequence is used to bind, cleave, and modify each nucleic acid target sequence. In some embodiments, the modification of the target region is an insertion, deletion, or a combination of insertion and deletion. The above describes methods for cleaving nucleic acid target sequences in polynucleotides (e.g., single-stranded cleavages or double-stranded cleavages in dsDNA), including providing one or more modified type I CRISPR-Cas effector complexes for introduction into cells.
[0410] Embodiments of the present invention include generating recombinant cells using one or more modified type I CRISPR-Cas effector complexes, wherein the gDNA of the recombinant cells contains knockout mutations (e.g., those in the B2M gene and / or the PDCD1 gene), knock-in mutations (e.g., editing at the TRAC site and integration of a CAR from a donor polynucleotide), or a combination thereof. In some embodiments, at least a portion of the donor polynucleotide is incorporated into the nucleic acid target sequence after cleavage at the nucleic acid target sequence in the TRAC gene of the gDNA. The donor polynucleotide may comprise a CAR construct wherein the CAR is inserted into the nucleic acid target sequence.
[0411] The recombinant cells prepared by the method of this invention can be used for adoptive cell transfer (ACT). ACT is a rapidly emerging immunotherapy that uses transplanted immune cells to treat cancer. ACT involves transferring cells into a patient. Most commonly, the immune cells are derived from the immune system and are intended to enhance immune function. In autologous cancer immunotherapy, immune cells or stem cells are harvested from a patient, expanded to large numbers through in vitro culture, and then returned to the patient. Immune cells or stem cells can be modified in culture in various ways (e.g., using genome editing to integrate CARs into the genome of T cells). In some embodiments, lymphocytes for modification are isolated from the subject, modified, and then reintroduced into the same subject. This technique is called autologous lymphocyte therapy. In allogeneic cancer immunotherapy, cultured and expanded immune cells or stem cells derived from a single donor can provide treatment for a large number of patients. Such immune cells or stem cells can also be modified in culture in various ways. In some embodiments, lymphocytes can be isolated, modified, and introduced into different subjects. This technique is called allogeneic lymphocyte therapy.
[0412] In some embodiments, this immunotherapy method may utilize lymphocytes, including but not limited to T cells, natural killer cells (NK cells), B cells, tumor-infiltrating lymphocytes (TILs), chimeric antigen receptor T cells (CAR-T cells), T cell receptor-modified T cells (TCRs), TCR CAR-T cells, CAR TIL cells, CAR-NK cells, modified NK cells, or lymphocyte-generating hematopoietic stem cells. In other embodiments, the cells are stem cells, dendritic cells, etc. The genome of such cells can be modified using one or more of the modified type I Cascade effector complexes of this invention (e.g., the generation of insertions and / or deletions in the lymphocyte genome).
[0413] Lymphocytes can be isolated from objects such as human objects, for example from blood, or from solid tumors, such as in the case of TILs, or from lymphoid organs such as the thymus, bone marrow, lymph nodes, and mucosa-associated lymphoid tissue, for modification. Techniques for isolating lymphocytes are well known in the art. For example, lymphocytes can be isolated from peripheral blood mononuclear cells (PBMCs), which can be separated from whole blood using, for example, ficoll, a hydrophilic polysaccharide and density gradient centrifugation to separate blood layers. Typically, an anticoagulated or defibrinated blood sample is plated on top of a ficoll solution and centrifuged to form distinct cell layers. The bottom layer consists of erythrocytes (red blood cells), which are collected or aggregated by the ficoll medium and settle completely to the bottom layer. The next layer mainly contains granulocytes, which also migrate downwards through the ficoll-paque solution. The next layer includes lymphocytes, as well as monocytes and platelets, with lymphocytes typically located at the interface between the plasma and the ficoll solution. To isolate the lymphocytes, this layer is recovered, washed with a saline solution to remove platelets, ficoll, and plasma, and then centrifuged again. Alternatively, centrifugation techniques (e.g., using...) can be employed. (Haemonetrics, Braintree, MA) machines or Lovo automated cell processing systems (Fresenius Kabi USA, LLC, Lake Zurich, IL) separate cells from donor blood.
[0414] Other techniques used for lymphocyte isolation include biopanning, which separates cell populations from solution by binding target cells to an antibody-coated plastic surface. Unwanted cells are then removed by treatment with specific antibodies and complement. Additionally, fluorescence-activated cell sorting (FACS) analysis can be used to detect and count lymphocytes. FACS analysis uses flow cytometry, which separates labeled cells based on differences in light scattering and fluorescence.
[0415] For TILs, lymphocytes are isolated from tumors and grown, for example, in high doses of IL-2, and selected using a co-culture assay for cytokine release against autologous tumors or HLA-matched tumor cell lines. Cultures showing evidence of increased specific reactivity compared to allogeneic, non-MHC-matched controls are rapidly expanded and then introduced into subjects for cancer treatment (see, for example, Rosenberg, S., et al., Clin. Cancer Res. 17:4550-4557 (2011); Dudly, M., et al., Science 298:850-854 (2002); Dudly, M., et al., J. Clin. Oncol. 26:5233-5239 (2008); Dudley, M., et al., J. Immnother. 26:332-342 (2003)).
[0416] After separation, lymphocytes can be characterized based on specificity, frequency, and function. Commonly used assays include the ELISPOT assay, which measures the frequency of T-cell responses.
[0417] In some implementations, CD4+ and CD8+ T cells are isolated from donor peripheral blood mononuclear cells (PBMCs). Those skilled in the art can isolate T cells or other lymphoid cells using the various methods described above. Such cells can also be isolated through differentiation from iPSC cells.
[0418] After isolation, lymphocytes can be activated using techniques known in the art to promote proliferation and differentiation into specialized effector lymphocytes. Activated T cell surface markers include, for example, CD3, CD4, CD8, PD1, IL2R, and others. Activated cytotoxic lymphocytes can kill target cells after binding to homologous receptors on the surface of target cells. NK cell surface markers include, for example, CD16, CD56, and others.
[0419] After isolation and optional activation, lymphocytes can be modified to provide desired properties. One or more modified type I Cascade effector complexes of the present invention can be used to introduce genomic modifications, including but not limited to introducing coding sequences to be expressed and / or inactivating endogenous gene expression. In some embodiments, one or more modified type I Cascade effector complexes of the present invention can be used to edit the TRAC gene (encoding the T cell receptor α constant), the B2M gene (encoding β2 microglobulin), and / or the PDCD1 gene (encoding programmed cell death protein 1; also known as PD-1).
[0420] T cells and NK cells are examples of lymphocytes that can be modified by the methods of the present invention. In some embodiments, one or more modified type I Cascade effector complexes of the present invention can introduce cleavage sites in the target region of a gene in the presence of a donor polynucleotide containing a CAR, wherein the CAR is integrated into the target region of the lymphocyte genome. In other embodiments, one or more modified type I Cascade effector complexes of the present invention can be used to introduce cleavage sites in the target region of a gene to promote the generation of knockout mutations to prevent gene expression.
[0421] In another embodiment, the modified type I Cascade effector complex of the present invention can be used to introduce genomic modifications into human iPSCs. In some embodiments, one or more modified type I Cascade effector complexes of the present invention can be used to edit the TRAC gene, B2M gene, and / or PDCD1 gene. In other embodiments, the modified type I Cascade effector complex, together with donor polynucleotides, can be used to introduce genomic modifications and coding sequences, such as CARs or cytokines (e.g., IL2, IL15, etc.). The modified iPSCs can then be further differentiated into mature cell types containing T cells and NK cells or dendritic cells. In some embodiments, the modified iPSCs can differentiate into CAR-T cells and CAR-NK cells.
[0422] In some embodiments of the method of the present invention, the donor polynucleotide comprises a polynucleotide encoding a CAR. The CAR can be targeted to insert via homologous recombination (“knock-in”) into the target region of a gene containing a cleavage site (e.g., the TRAC gene). One advantage of this approach is that it can also provide targeted knockout of the TRAC gene; that is, disable the TRAC gene. Examples of extracellular antigen recognition domains that can be incorporated into the CAR construct have been described above (see Table 2). In one embodiment, the extracellular antigen recognition domain comprises a CD19 binding moiety (e.g., anti-CD19 scFv). In another embodiment, the extracellular antigen recognition domain comprises a B-cell maturation antigen (BCMA) binding moiety (e.g., anti-BCMA scFv).
[0423] In embodiments of the method of the present invention, including the generation of cleavage sites in a target region of DNA, the method may further include introducing a donor polynucleotide into a modified cell, thereby facilitating the insertion of at least a portion of the donor polynucleotide into the target region containing the cleavage site of the modified cell. The donor polynucleotide may be introduced directly into the modified cell. In some embodiments, a vector is used to introduce the donor polynucleotide. General methods for constructing vectors are known in the art. Examples of viral vectors include, but are not limited to, lentiviruses, retroviruses, adenoviruses, herpes simplex virus I or II, parvoviruses, reticuloendothelial proliferative virus, and AAV vectors.
[0424] Other embodiments of the method of the present invention include introducing a mutation into the B2M gene. In a preferred embodiment, the mutation is a knockout mutation in the B2M gene.
[0425] Other embodiments of the method of the present invention include introducing a mutation into the PDCD1 gene. In a preferred embodiment, the mutation is a knockout mutation in the PDCD1 gene.
[0426] Genome modification facilitated by one or more modified type I Cascade effector complexes of the present invention can be carried out by simultaneously or sequentially introducing a modified Cascade complex, polynucleotides (e.g., plasmids or expression cassettes), or a mixture thereof into host cells (e.g., lymphocytes).
[0427] After generating modified lymphocytes, the lymphocytes can be screened to select cells that express (e.g., express desired cell surface receptors) or do not express (e.g., cell surface proteins whose expression has been inactivated by genome editing using one or more modified type I Cascade effector complexes) using methods such as high-throughput screening techniques, including but not limited to FACS, microfluidic-based screening platforms. These techniques are known in the art (see, for example, Wojcik, M., et al., Int. J. Mol.. Sci. 16: 24918-24945 (2015)).
[0428] Once modified lymphocytes are generated, they can be formulated into pharmaceutical compositions for delivery to a subject requiring treatment. The compositions of the present invention comprise modified lymphocytes and one or more pharmaceutically acceptable excipients. Exemplary excipients include, but are not limited to, carbohydrates, inorganic salts, antimicrobial agents, antioxidants, surfactants, buffers, acids, bases, and combinations thereof. Suitable excipients for injectable compositions include water, alcohols, polyols, glycerol, vegetable oils, phospholipids, and surfactants. Carbohydrates such as sugars, derivatized sugars such as sugar alcohols, aldonic acids, esterified sugars, and / or sugar polymers can be provided as excipients. Specific carbohydrate excipients include, for example: monosaccharides, such as fructose, maltose, galactose, glucose, D-mannose, sorbitol, etc.; disaccharides, such as lactose, sucrose, trehalose, cellobiose, etc.; polysaccharides, such as raffinose, melitriose, maltodextrin, dextran, starch, etc.; and aldose alcohols, such as mannitol, xylitol, maltitol, lactitol, xylitol, sorbitol (glucol), pyranosylsorbitol, inositol, etc. Excipients may also include inorganic salts or buffers, such as citric acid, sodium chloride, potassium chloride, sodium sulfate, potassium nitrate, sodium dihydrogen phosphate, disodium hydrogen phosphate, and combinations thereof. Refrigerants (e.g., (BioLife Solutions Inc., Bothell, WA) CS2, CS5, or CS10 cryo-mediums can be used to freeze cells for storage and transport.
[0429] The pharmaceutical compositions of the present invention may further include antimicrobial agents for preventing or inhibiting the growth of microorganisms. Non-limiting examples of antimicrobial agents suitable for use in the present invention include benzalkonium chloride, benzyl chloride, benzyl alcohol, cetylpyridinium chloride, chlorobutanol, phenol, phenethyl alcohol, phenylmercuric nitrate, thimerosal, and combinations thereof.
[0430] Antioxidants may also be present in the pharmaceutical composition. Antioxidants are used to prevent oxidation, thereby preventing the deterioration of lymphocytes or other components in the formulation. Suitable antioxidants for use in this invention include, for example, ascorbyl palmitate, butylated hydroxyanisole, butylated hydroxytoluene, hypophosphite, monothioglycerol, propyl gallate, sodium bisulfite, sodium formaldehyde sulfoxylate, sodium metabisulfite, and combinations thereof.
[0431] Surfactants can be present as excipients. Exemplary surfactants include: polysorbates, such as TWEEN 20 and TWEEN 80, and Pluronic compounds, such as F68 and F88 (BASF, Mount Olive, New Jersey); sorbitol esters; lipids, such as phospholipids, such as lecithin and other phosphatidylcholine, phosphatidylethanolamine (although preferably not in liposomal form), fatty acids and fatty esters; steroids, such as cholesterol; chelating agents, such as EDTA; and zinc and other such suitable cations.
[0432] Acids or bases may be present as excipients in pharmaceutical compositions. Non-limiting examples of acids that may be used include those selected from the following: hydrochloric acid, acetic acid, phosphoric acid, citric acid, malic acid, lactic acid, formic acid, trichloroacetic acid, nitric acid, perchloric acid, phosphoric acid, sulfuric acid, fumaric acid, and combinations thereof. Suitable examples of bases include, but are not limited to, those selected from the following: sodium hydroxide, sodium acetate, ammonium hydroxide, potassium hydroxide, ammonium acetate, potassium acetate, sodium phosphate, potassium phosphate, sodium citrate, sodium formate, sodium sulfate, potassium sulfate, potassium fumarate, and combinations thereof.
[0433] The number of lymphocytes (or other recombinant cells) in the composition will vary depending on a number of factors, but it will be optimally therapeutically effective when the composition is in unit dose form or in a container (e.g., a bag). The therapeutically effective dose can be experimentally determined by repeatedly administering increasing amounts of the composition to determine which amount produces the clinically desired endpoint.
[0434] The amount of any single excipient in the composition will vary depending on the nature and function of the excipient and the specific needs of the composition. Typically, the optimal amount of any single excipient is determined through routine experiments, i.e., by preparing compositions containing varying amounts of excipient (ranging from low to high), examining stability and other parameters, and then determining the range that yields optimal performance without significant side effects. However, typically, the amount of excipient present in the composition is from about 1% to about 99% by weight, preferably from about 5% to about 98% by weight, more preferably from about 15% to about 95% by weight, and a concentration of less than 30% by weight is most preferred. These aforementioned pharmaceutical excipients, along with other excipients, are described in "Remington: The Science & Practice of Pharmacy," current edition, Williams & Williams; the "Physician's Desk Reference," current edition, Medical Economics, Montvale, NJ; and Kibbe, AH, Handbook of Pharmaceutical Excipients, current edition, American Pharmaceutical Association, Washington, DC.
[0435] The pharmaceutical composition can be contained in syringes, implantable devices, etc., according to the intended method of delivery and use. Preferably, the amount of composition present is suitable for a single dose, in a pre-measured or pre-packaged form.
[0436] The pharmaceutical compositions described herein may optionally include one or more additional agents, such as other drugs used to treat the target cancer of the patient or to treat known side effects of the treatment. For example, T cells release cytokines into the bloodstream, which can lead to dangerously high fever and a sharp drop in blood pressure. This condition is called cytokine release syndrome (CRS). In many patients, CRS can be managed with standard supportive care, including steroids and immunotherapy, such as tocilizumab (Actemra™, Genentech, South San Francisco, CA), which blocks IL-6 activity.
[0437] At least one therapeutically effective cycle of treatment with the modified lymphocyte composition will be administered to the subject. "Therapeutically effective cycle" means a treatment cycle, when administered, that elicits a positive therapeutic response with respect to the individual's target disease. "Positive therapeutic response" means that the individual treated according to the invention exhibits improvement in one or more symptoms of the disease, including improvements such as tumor reduction and / or reduced need for lymphocyte therapy.
[0438] In some embodiments, multiple therapeutically effective doses of compositions comprising lymphocytes or other drugs will be administered. The compositions of the present invention, although not necessarily, are generally administered by injection (e.g., subcutaneous, intradermal, intravenous, intraarterial, intramuscular, intraperitoneal, intramedullary, intratumoral, intranodular), by infusion, or locally. The pharmaceutical formulation may be in the form of a liquid solution or suspension immediately prior to administration. The foregoing is intended to be exemplary, as other routes of administration are also contemplated. The pharmaceutical compositions may be administered using the same or different routes of administration according to any medically acceptable method known in the art.
[0439] The actual dose administered will vary depending on the subject's age, weight, and general condition, as well as the severity of the condition being treated, the judgment of the healthcare professional, and the specific lymphocytes administered. The effective therapeutic dose can be determined by those skilled in the art and will be adjusted according to the specific requirements of each particular case.
[0440] Typically, the therapeutically effective dose of lymphocytes per patient will range from approximately 1 x 10^6 cells. 5 One to approximately 1 x 10 10 One or more lymphocytes, such as 1 x 10^12 6 One to approximately 1 x 10 10 One, for example, 1 x 10 7 One to 1x 10 9 One, such as 5 x 10 7 5 x 10 8 One, or any quantity within these ranges. Other dosage ranges may be 1 x 10 per kg / body weight.4 One to 1x10 10 The total number of lymphocytes can be administered in a single large dose or in two or more doses, such as one day or more apart. The amount of compound administered will depend on the potency of the specific lymphocyte composition, the disease being treated, and the route of administration.
[0441] Additionally, the dosage may include a mixture of lymphocytes, such as a mixture of CD8+ and CD4+ cells. If a mixture of CD8+ and CD4+ cells is provided, the ratio of CD8+ to CD4+ cells may be, for example, 1:1, 1:2 or 2:1, 1:3 or 3:1, 1:4 or 4:1, 1:5 or 5:1, etc.
[0442] Modified lymphocytes can be administered before, simultaneously with, or after other agents. If provided simultaneously with other agents, the modified lymphocytes can be provided in the same or different compositions. Therefore, lymphocytes and other agents can be provided to an individual in a concurrent treatment manner. By "concurrent treatment," it means administering the drug to a subject such that the combination of substances produces a therapeutic effect in the treated subject. For example, depending on a particular dosing regimen, concurrent treatment can be achieved by administering a combination of a pharmaceutical composition containing a therapeutically effective dose of modified lymphocytes and a pharmaceutical composition containing at least one other agent, such as another chemotherapeutic agent. Similarly, modified lymphocytes and a therapeutic agent can be administered in at least one therapeutic dose. Individual pharmaceutical compositions can be administered simultaneously or at different times (e.g., consecutively, in any order, on the same day or on different days), as long as the combination of substances produces a therapeutic effect in the treated subject.
[0443] As described herein, the modified type I Cascade effector complex of this invention provides a genome editing tool. Experiments demonstrating functional remodeling of class-1 CRISPR-Cas systems in mammalian cells for genome editing show that such improved plasmid designs can allow the use of other class-1 CRISPR-Cas systems, including those exhibiting fewer protein components and unique PAM requirements, and possibly even RNA and DNA-targeting effector complexes from type III CRISPR-Cas systems (see, for example, Hille, F., et al., Cell 172:1239-1259 (2018); Tamulaitis, G., et al., Trends Microbiol. 25:49-61 (2017)). The multi-subunit nature of the Cascade complex provides the possibility of multivalent and / or spatially precise recruitment of effector fusions, such as synthetic transcription factors, epigenome modifiers, and base editors. Additionally, heterologous expression of the complete DNA interference pathway from a type I system—namely, Cascade-mediated recruitment of Cas3 helicase-nucleases to genomic target sites—can be used to generate large DNA deletions, exposing longer ssDNA bundles for homology-guided repair and / or mechanically disrupting protein-DNA roadblocks at identified genomic sites. Therefore, in one embodiment of the invention, a modified type I CRISPR-Cas system can be used to generate large deletion regions and can introduce donor polynucleotides (e.g., suitable homologous arms) into the cell, thereby facilitating the insertion of at least a portion of the donor polynucleotide into the region.
[0444] The embodiments of the present invention include, but are not limited to, the following.
[0445] Implementation Scheme 1. A composition comprising:
[0446] The first modified type I CRISPR-Cas effector complex comprises:
[0447] The first Cse2 subunit protein, the first Cas5 subunit protein, the first Cas6 subunit protein, and the first Cas7 subunit protein.
[0448] A first fusion protein comprising a first Cas8 subunit protein and a first FokI, wherein the N-terminus or C-terminus of the first Cas8 subunit protein is covalently linked to the C-terminus or N-terminus of the first FokI via a first linker polypeptide, and wherein the first linker polypeptide has a length of approximately 10 amino acids to approximately 40 amino acids.
[0449] Contains a first guide polynucleotide capable of binding to a first spacer region of a first nucleic acid target sequence; and
[0450] The second modified type I CRISPR-Cas effector complex comprises:
[0451] Second Cse2 subunit protein, second Cas5 subunit protein, second Cas6 subunit protein, and second Cas7 subunit protein.
[0452] A second fusion protein comprising a second Cas8 subunit protein and a second FokI, wherein the N-terminus or C-terminus of the second Cas8 subunit protein is covalently linked to the C-terminus or N-terminus of the second FokI via a second linker polypeptide, and wherein the second linker polypeptide has a length of approximately 10 amino acids to approximately 40 amino acids.
[0453] The second guide polynucleotide contains a second spacer region capable of binding to a second nucleic acid target sequence, wherein the pre-spacer region adjacent motif (PAM) of the second nucleic acid target sequence and the PAM of the first nucleic acid target sequence have a spacer interval of approximately 20 bp to approximately 42 bp.
[0454] Implementation Scheme 2. The composition as described in Implementation Scheme 1, wherein the first linker polypeptide has a length of about 15 amino acids to about 30 amino acids.
[0455] Implementation Scheme 3. The composition as described in Implementation Scheme 2, wherein the first linker polypeptide has a length of about 17 amino acids to about 20 amino acids.
[0456] Implementation Scheme 4. The composition of any one of Implementation Schemes 1-3, wherein the second linker polypeptide has a length of about 15 amino acids to about 30 amino acids.
[0457] Implementation Scheme 5. The composition as described in Implementation Scheme 4, wherein the second linker polypeptide has a length of about 17 amino acids to about 20 amino acids.
[0458] Implementation Scheme 6. The composition as described in any of the foregoing embodiments, wherein the first linker polypeptide and the second linker polypeptide are of the same length.
[0459] Implementation Scheme 7. The composition as described in any of the foregoing implementation schemes, wherein each of the second nucleic acid target sequence and the first nucleic acid target sequence has a spacer interval of about 22 bp to about 40 bp.
[0460] Implementation Scheme 8. The composition as described in Implementation Scheme 7, wherein each of the second nucleic acid target sequence and the first nucleic acid target sequence has a spacer interval of about 26 bp to about 36 bp.
[0461] Implementation Scheme 9. The composition as described in Implementation Scheme 8, wherein each of the second nucleic acid target sequence and the first nucleic acid target sequence has a spacer interval of about 29 bp to about 35 bp.
[0462] Implementation Scheme 10. The composition as described in Implementation Scheme 9, wherein each of the second nucleic acid target sequence and the first nucleic acid target sequence has a spacer interval of about 30 bp to about 34 bp.
[0463] Implementation Scheme 11. The composition as described in any of the foregoing embodiments, wherein the first FokI and the second FokI are monomer subunits capable of binding to form a homodimer.
[0464] Implementation Scheme 12. The composition of any one of Implementation Schemes 1-10, wherein the first FokI and the second FokI are different monomer subunits capable of combining to form a heterodimer.
[0465] Implementation Scheme 13. The composition as described in any of the foregoing embodiments, wherein the N-terminus of the first Cas8 subunit protein is covalently linked to the C-terminus of the first FokI via a first linker polypeptide.
[0466] Implementation Scheme 14. The composition of any one of Implementation Schemes 1-12, wherein the C-terminus of the first Cas8 subunit protein is covalently linked to the N-terminus of the first FokI via a first linker polypeptide.
[0467] Implementation Scheme 15. The composition as described in any of the foregoing embodiments, wherein the N-terminus of the second Cas8 subunit protein is covalently linked to the C-terminus of the second FokI via a second linker polypeptide.
[0468] Implementation Scheme 16. The composition of any one of Implementation Schemes 1-14, wherein the C-terminus of the second Cas8 subunit protein is covalently linked to the N-terminus of the second FokI via a second linker polypeptide.
[0469] Implementation Scheme 17. The composition as described in any of the foregoing embodiments, wherein each of the first Cas8 subunit protein and the second Cas8 subunit protein comprises the same amino acid sequence.
[0470] Implementation Scheme 18. The composition as described in any of the foregoing embodiments, wherein each of the first Cse2 subunit protein and the second Cse2 subunit protein contains the same amino acid sequence, each of the first Cas5 subunit protein and the second Cas5 subunit protein contains the same amino acid sequence, each of the first Cas6 subunit protein and the second Cas6 subunit protein contains the same amino acid sequence, and each of the first Cas7 subunit protein and the second Cas7 subunit protein contains the same amino acid sequence.
[0471] Implementation Scheme 19. The composition as described in any of the foregoing embodiments, wherein the first guide polynucleotide comprises RNA.
[0472] Implementation Scheme 20. The composition as described in any of the foregoing embodiments, wherein the second guide polynucleotide comprises RNA.
[0473] Implementation Scheme 21. The composition as described in any of the foregoing embodiments, wherein the genomic DNA comprises a PAM of a second nucleic acid target sequence and a PAM of a first nucleic acid target sequence.
[0474] Implementation Scheme 22. A cell comprising: any of the compositions described in the preceding embodiments.
[0475] Implementation Scheme 23. The cell as described in Implementation Scheme 22, wherein the cell’s genomic DNA contains a second nucleic acid target sequence PAM and a first nucleic acid target sequence PAM.
[0476] Implementation scheme 24. Cells as described in implementation scheme 22 or 23, wherein the cells are prokaryotic cells.
[0477] Implementation Scheme 25. Cells as described in Implementation Scheme 22 or 23, wherein the cells are eukaryotic cells.
[0478] Implementation Scheme 26. One or more nucleic acid sequences encoding the first Cse2 subunit protein, the first Cas5 subunit protein, the first Cas6 subunit protein, the first Cas7 subunit protein, the first fusion protein, and the first guide polynucleotide as described in any one of Implementation Schemes 1-21.
[0479] Implementation Scheme 27. One or more nucleic acid sequences encoding the second Cse2 subunit protein, the second Cas5 subunit protein, the second Cas6 subunit protein, the second Cas7 subunit protein, the second fusion protein, and the second guide polynucleotide as described in any one of Implementation Schemes 1-21.
[0480] Implementation scheme 28. One or more expression cassettes comprising one or more nucleic acid sequences as described in implementation scheme 26, implementation scheme 27, or both implementation scheme 26 and implementation scheme 27.
[0481] Implementation Scheme 29. One or more carriers comprising one or more expression boxes as described in Implementation Scheme 28.
[0482] Implementation Scheme 30. A method for incorporating a polynucleotide comprising a first nucleic acid target sequence and a second nucleic acid target sequence, said method comprising:
[0483] Provide a composition according to any one of embodiments 1-21 for use in introduction into cells or biochemical reactions; and
[0484] Introducing the composition into cells or a biochemical reaction thereby promoting contact between a first modified type I CRISPR-Cas effector complex and a first nucleic acid target sequence, and between a second modified type I CRISPR-Cas effector complex and a second nucleic acid target sequence, resulting in the binding of the first modified type I CRISPR-Cas effector complex to the first nucleic acid target sequence, and the binding of the second modified type I CRISPR-Cas effector complex to the second nucleic acid target sequence in a polynucleotide.
[0485] Implementation Scheme 31. The method as described in Implementation Scheme 30, wherein the genomic DNA comprises polynucleotides.
[0486] Implementation Scheme 32. A method for cleaving a polynucleotide containing a first nucleic acid target sequence and a second nucleic acid target sequence, the method comprising:
[0487] Provide a composition according to any one of embodiments 1-21 for use in introduction into cells or biochemical reactions; and
[0488] Introducing the composition into cells or a biochemical reaction thereby promoting contact between a first modified type I CRISPR-Cas effector complex and a first nucleic acid target sequence, and between a modified second type I CRISPR-Cas effector complex and a second nucleic acid target sequence, resulting in the first nucleic acid target sequence being cleaved by the first modified type I CRISPR-Cas effector complex and the second nucleic acid target sequence being cleaved by the second modified type I CRISPR-Cas effector complex.
[0489] Implementation Scheme 33. The method as described in Implementation Scheme 32, wherein the genomic DNA comprises polynucleotides.
[0490] Implementation Scheme 34. A kit comprising: the composition of any one of Implementation Schemes 1-21; and a buffer.
[0491] Implementation Scheme 35. A kit comprising: one or more nucleic acid sequences as described in Implementation Scheme 26, Implementation Scheme 27, or Implementation Scheme 26 and Implementation Scheme 27; and a buffer.
[0492] Implementation Scheme 36. A composition comprising:
[0493] The modified type I CRISPR-Cas effector complex comprises:
[0494] Cse2 subunit protein, Cas5 subunit protein, Cas6 subunit protein, and Cas7 subunit protein.
[0495] A first fusion protein comprising a Cas8 subunit protein and a first FokI, wherein the N-terminus or C-terminus of the first Cas8 subunit protein is covalently linked to the C-terminus or N-terminus of the first FokI via a first linker polypeptide, respectively.
[0496] Guide polynucleotides containing spacer regions capable of binding to nucleic acid target sequences; and
[0497] The second fusion protein comprises a modified type I CRISPR-Cas3 fusion protein containing dCas3* protein and a second FokI, wherein the N-terminus or C-terminus of the dCas3* protein is covalently linked to the C-terminus or N-terminus of the second FokI via a second linker polypeptide, and wherein the first linker polypeptide has a length of approximately 10 to approximately 40 amino acids. The effector complex comprises...
[0498] Implementation Scheme 37. The composition as described in Implementation Scheme 36, wherein the first linker polypeptide has a length of about 5 amino acids to about 40 amino acids.
[0499] Implementation Scheme 38. The composition as described in Implementation Scheme 36, wherein the second linker polypeptide has a length of about 5 amino acids to about 40 amino acids.
[0500] Implementation Scheme 39. Cells comprising: the composition of any one of Implementation Schemes 36-38.
[0501] Implementation scheme 40. The cell as described in implementation scheme 39, wherein the cell is a prokaryotic cell.
[0502] Implementation Scheme 41. The cell as described in Implementation Scheme 39, wherein the cell is a eukaryotic cell.
[0503] Implementation Scheme 42. One or more nucleic acid sequences encoding the Cse2 subunit protein, Cas5 subunit protein, Cas6 subunit protein, Cas7 subunit protein, first fusion protein, and guide polynucleotide as described in any one of Implementation Schemes 36-38.
[0504] Implementation scheme 43. One or more nucleic acid sequences encoding the second fusion protein described in any one of implementation schemes 36-38.
[0505] Implementation scheme 44. One or more expression cassettes comprising one or more nucleic acid sequences of implementation scheme 42, implementation scheme 43, or implementation scheme 42 and implementation scheme 43.
[0506] Implementation scheme 45. One or more carriers comprising one or more expression boxes as described in implementation scheme 44.
[0507] Implementation Scheme 46. A method for incorporating a polynucleotide containing a nucleic acid target sequence, said method comprising:
[0508] Provide the composition according to any one of embodiments 36-38 for use in introduction into cells or biochemical reactions; and
[0509] Introducing the composition into cells or biochemical reactions promotes contact between the modified type I CRISPR-Cas effector complex and the nucleic acid target sequence, as well as contact between the second fusion protein and the modified type I CRISPR-Cas effector complex, resulting in the binding of the modified type I CRISPR-Cas effector complex and the second fusion protein to the nucleic acid target sequence in the polynucleotide.
[0510] Implementation Scheme 47. The method as described in Implementation Scheme 46, wherein the genomic DNA comprises polynucleotides.
[0511] Implementation Scheme 48. A method for cleaving a polynucleotide containing a nucleic acid target sequence, the method comprising:
[0512] Provide the composition according to any one of embodiments 36-38 for use in introduction into cells or biochemical reactions; and
[0513] The composition is introduced into cells or biochemical reactions to promote contact between a first modified type I CRISPR-Cas effector complex and a first nucleic acid target sequence, and between a modified second type I CRISPR-Cas effector complex and a second nucleic acid target sequence.
[0514] Introducing the composition into cells or biochemical reactions promotes contact between the modified type I CRISPR-Cas effector complex and the nucleic acid target sequence, as well as contact between the second fusion protein and the modified type I CRISPR-Cas effector complex, resulting in the cleavage of the nucleic acid target sequence by the modified type I CRISPR-Cas effector complex and the second fusion protein.
[0515] Implementation Scheme 49. The method as described in Implementation Scheme 48, wherein the genomic DNA comprises polynucleotides.
[0516] Implementation Scheme 50. A kit comprising: the composition of any one of Implementation Schemes 36-38; and a buffer.
[0517] Implementation Scheme 51. A kit comprising one or more nucleic acid sequences as described in Implementation Scheme 42, Implementation Scheme 43, or Implementation Scheme 42 and Implementation Scheme 43; and a buffer.
[0518] Implementation Scheme 52. A modified type I CRISPRCas3 mutant protein (“mCas3 protein”) capable of reducing migration along DNA compared to wild-type type I CRISPRCas3 protein (“wtCas3 protein”), said mCas3 protein comprising:
[0519] It shares approximately 95% or more sequence identity with the corresponding wtCas3 protein.
[0520] Nuclear localization signals are covalently linked at the amino terminus, carboxyl terminus, or both.
[0521] One or more mutations that can downregulate helicase activity, wherein the modified type I CRISPRCas3 mutant protein retains nuclease activity;
[0522] The DNA in this context is double-stranded DNA (dsDNA) containing a target region with a nucleic acid target sequence.
[0523] Specifically, when the wtCas3 protein is associated with the corresponding Cascade nucleoprotein complex (“Cascade NP complex / wtCas3 protein”), and the Cascade NP complex contains a guide with a spacer region complementary to the nucleic acid target sequence, the binding of the Cascade NP complex / wtCas3 protein to the nucleic acid target sequence facilitates cleavage in the DNA target region, resulting in deletion (“wtCas3 deletion”); and
[0524] When the mCas3 protein is associated with the Cascade NP complex (“Cascade NP complex / mCas3 protein”) and binds to the nucleic acid target sequence, it facilitates cleavage in the DNA target region, resulting in a shorter deletion compared to wtCas3- deletion.
[0525] Implementation Scheme 53. The mCas3 protein as described in Implementation Scheme 53, wherein one or more mutations are substitutions of amino acids.
[0526] Implementation Scheme 54. The mCas3 protein as described in any of the foregoing implementation schemes, wherein one or more mutations are located in the RecA1 or RecA2 region of the helicase domain.
[0527] Implementation Scheme 55. The mCas3 protein as described in any of the foregoing implementation schemes, wherein one or more mutations downregulate the binding of the mCas3 protein to single-stranded DNA (ssDNA) relative to the wtCas3 protein.
[0528] Implementation Scheme 56. The mCas3 protein as described in any of the foregoing implementation schemes, wherein one or more mutations downregulate the hydrolysis of adenosine triphosphate (ATP) by the mCas3 protein, or downregulate the ATP-binding protein of mCas3.
[0529] Implementation Scheme 57. The mCas3 protein as described in any of the foregoing embodiments, wherein the coding sequence of the mCas3 protein is covalently linked to the amino or carboxyl terminus of the coding sequence of the Cas protein of the Cascade NP complex.
[0530] Implementation scheme 58. The mCas3 protein as described in any of the foregoing implementation schemes, wherein one or more mutations downregulate the binding of the mCas3 protein to single-stranded DNA (ssDNA) relative to the wtCas3 protein.
[0531] Implementation Scheme 59. The mCas3 protein as described in any of the foregoing embodiments, wherein the coding sequence of the mCas3 protein is covalently linked to the amino or carboxyl terminus of the coding sequence of the Cas protein of the Cascade RNP complex.
[0532] Implementation Scheme 60. The mCas3 protein as described in any of the foregoing implementation schemes, wherein the Cas protein is selected from: Cse2, Cas8, Cas7, Cas6 and Cas5.
[0533] Implementation Scheme 61. The mCas3 protein as described in any of the foregoing implementation schemes, wherein the wtCas3 protein is a CRISPRCas3 protein of Escherichia coli type 1.
[0534] Implementation Scheme 62. The mCas3 protein as described in Implementation Scheme 61, wherein one or more mutations are selected from D452H, A602V, and D452H and A602V.
[0535] Implementation Scheme 63. The mCas3 protein as described in any of the foregoing implementation schemes, wherein DNA is present in the cell.
[0536] Implementation Scheme 64. The mCas3 protein as described in Implementation Scheme 63, wherein the cell is a eukaryotic cell.
[0537] Implementation Scheme 65. The mCas3 protein as described in Implementation Scheme 64, wherein the eukaryotic cell is a mammalian cell (e.g., a human cell).
[0538] Implementation Scheme 66. One or more polynucleotides encoding the mCas3 protein as described in any one of Implementation Schemes 52-65.
[0539] Implementation Scheme 67. A plasmid comprising a multinucleotide sequence operatively linked to a regulatory sequence encoding the mCas3 protein as described in any one of Implementation Schemes 52-65 for expression in mammalian cells.
[0540] Implementation Scheme 68. One or more plasmids comprising a polynucleotide sequence encoding the mCas3 protein of any one of Implementation Schemes 52-65, and one or more polynucleotides operatively linked to a regulatory sequence encoding a protein component of a corresponding type I CRISPR Cascade for expression in mammalian cells.
[0541] Implementation Scheme 69. One or more plasmids as described in Implementation Scheme 68, further comprising a plasmid encoding one or more guide polynucleotides operatively linked to a regulatory sequence for expression in mammalian cells.
[0542] Implementation scheme 70. Type I CRISPRCascade nucleoprotein complex, comprising the mCas3 protein as described in any one of embodiments 52-65.
[0543] Implementation Scheme 71. The type I CRISPR Cascade nucleoprotein complex as described in Implementation Scheme 70, wherein the nucleoprotein complex is an RNP.
[0544] Although preferred embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments have been provided by way of example only. From this specification and the embodiments, those skilled in the art can determine the essential features of the invention, and can make changes, substitutions, alterations, and modifications to adapt the invention to various uses and conditions without departing from the spirit and scope thereof. Such changes, substitutions, alterations, and modifications are also intended to fall within the scope of this disclosure.
[0545] Example
[0546] Aspects of the invention are illustrated in the following embodiments. Efforts have been made to ensure the accuracy of the figures used (e.g., quantities, concentrations, percentage changes, etc.), but some experimental errors and biases should be taken into account. Unless otherwise indicated, temperatures are in degrees Celsius and pressures are at or near atmospheric pressure. It should be understood that these embodiments are given by way of example only and are not intended to limit the scope of the invention.
[0547] Example 1
[0548] Computer design of polynucleotides encoding Cascade components
[0549] This embodiment provides a description of designing polynucleotide components encoding Cascade using genes, proteins, and CRISPR sequences derived from the IE-type CRISPR-Cas system.
[0550] Table 15 shows the polynucleotide DNA sequences of five proteins encoding Cascade IE type, specifically derived from the gene of *E. coli* strain K-12MG1655, and the amino acid sequences of the resulting protein components. The genome sequences were obtained from NCBI reference sequence NZ_CP014225.1. In Table 15, the polynucleotide sequences are either derived from *E. coli* gDNA amplification or from manufacturer-produced polynucleotides specifically codon-optimized for expression in *E. coli* and for expression in human cells.
[0551]
[0552] In addition, several fusion proteins containing the Cascade protein were designed. Table 16 shows the ...
Claims
1. An engineered Type I CRISPR Cas3 mutant protein ("mCas3 protein") that is capable of reducing movement along DNA relative to a wild-type Type I CRISPR Cas3 protein ("wtCas3 protein"), wherein the mCas3 protein is a Pseudomonas aeruginosa (PAO1) Cas3 protein having the amino acid sequence of SEQ ID NO: 1, and wherein the mCas3 protein comprises a mutation at amino acid position 448 of SEQ ID NO:
1. Pseudomonas sp. ) S-6-2 D448A mCas3 protein, and the nucleic acid localization signal is covalently attached to the amino terminus, the carboxy terminus, or both the amino terminus and the carboxy terminus of the mCas3 protein, and wherein the DNA is double-stranded DNA (dsDNA) comprising a target region comprising a nucleic acid target sequence.
2. The mCas3 protein of claim 1, wherein the mCas3 protein is covalently attached to the amino terminus or the carboxy terminus of a Cas protein of a Type I CRISPR PseCascade nucleoprotein (NP) complex.
3. The mCas3 protein of claim 1, wherein the DNA is within a cell.
4. The mCas3 protein of claim 3, wherein the cell is a eukaryotic cell.
5. A Type I CRISPR Cascade nucleoprotein complex comprising the mCas3 protein of any one of claims 1-4.
6. An engineered Type I CRISPR-Cas effector composition comprising: a Type I CRISPR-PseCascade subunit protein; a Type I guide polynucleotide; an engineered Type I CRISPR mCas3 protein that is Pseudomonas S-6-2 D448A mCas3 protein.
7. The engineered Type I CRISPR-Cas effector composition of claim 6, wherein the sequence of the Pseudomonas S-6-2 D448A mCas3 protein is set forth in SEQ ID NO: 1919.
8. The engineered Type I CRISPR-Cas effector composition of claim 6, further comprising a linker polypeptide that covalently links the engineered Type I CRISPR mCas3 protein to a Type I CRISPR-PseCascade subunit protein.
9. The engineered Type I CRISPR-Cas effector composition of claim 6, wherein the Type I CRISPR-PseCascade subunit protein is selected from the group consisting of a Cas8 protein, a Cas5 protein, and a Cas7 protein.
10. The engineered Type I CRISPR-Cas effector composition of claim 9, further comprising a Type I CRISPR Cas6 protein.
11. The engineered Type I CRISPR-Cas effector composition of claim 9, further comprising a Type I CRISPR Cse2 protein.
12. A cell comprising the engineered Type I CRISPR-Cas effector composition of any one of claims 6-11.
13. The cell of claim 12, further comprising: a second engineered Type I CRISPR-Cas effector composition comprising: a second Type I CRISPR-PseCascade subunit protein; a second Type I guide polynucleotide; and a second engineered Type I CRISPR mCas3 protein that is Pseudomonas S-6-2 D448A mCas3 protein.
14. The cell of claim 12, further comprising a donor polynucleotide.
15. The cell of claim 12, wherein the cell comprises a eukaryotic cell.
16. The cell of claim 15, wherein the eukaryotic cell comprises an induced pluripotent stem cell.
17. An in vitro or ex vivo method of nicking double-stranded DNA (dsDNA), comprising: contacting a first nick site in a first target region in the dsDNA with a first engineered Type I CRISPR-Cas effector composition, the first engineered Type I CRISPR-Cas effector composition comprising: a first Type I CRISPR-PseCascade subunit protein; a first Type I guide polynucleotide; and a first engineered Type I CRISPR mCas3 protein that is Pseudomonas S-6-2 D448A mCas3 protein; such that the first engineered Type I CRISPR-Cas effector composition binds the first target region, and the first engineered Type I CRISPR mCas3 protein binds and nicks the dsDNA at the first nick site.
18. The in vitro or ex vivo method of claim 17, further comprising: contacting a second nick site in a second target region in the dsDNA with a second engineered Type I CRISPR-Cas effector composition, the second engineered Type I CRISPR-Cas effector composition comprising: a second Type I CRISPR-PseCascade subunit protein; a second Type I guide polynucleotide; and a second engineered Type I CRISPR mCas3 protein that is Pseudomonas S-6-2 D448A mCas3 protein; such that the second engineered Type I CRISPR-Cas effector composition binds the second target region, and the second engineered Type I CRISPR mCas3 protein binds and nicks the dsDNA at the second nick site.
19. The in vitro or ex vivo method of claim 18, wherein the distance between the first nick site and the second nick site on the dsDNA is 1 to 120 base pairs.
Citation Information
Patent Citations
Gastric disease drug and its preparation method
CN1273182C
Compositions and methods of nucleic acid-targeting nucleic acids
US20140315985A1
Nucleic acids encoding chimeric T cell receptors
US7446190B2
Chimeric receptors with 4-1BB stimulatory signaling domain
US8399645B2
Modified cascade ribonucleoproteins and uses thereof
US9885026B2