Engineered proteins and methods of use thereof
By engineering the LbCas12a protein and introducing specific amino acid sequence mutations to form a heterologous polypeptide complex, the problem of the lack of target strand cleavage enzymes in type V CRISPR endonucleases was solved, thus improving the efficiency of target nucleic acid modification and editing.
Patent Information
- Application Number
- CN202480019848.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-01
- Filing Date
- 2024-03-01
- Publication Date
- 2025-11-04
Smart Images

Figure CN120897995A_ABST
Abstract
Description
[0001] Cross-referencing related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 487,643, filed March 1, 2023, the disclosure of which is incorporated herein by reference in its entirety.
[0003] Declaration regarding the electronic document of the sequence list
[0004] The XML text format sequence list entitled 1499-92_ST26.xml, which is 354,813 bytes in size, was generated on March 1, 2024 and filed with this application, is hereby incorporated by reference in this specification for its disclosure. Technical Field
[0005] This invention relates to engineered proteins (e.g., engineered enzymes) and methods of using such proteins. The invention further relates to compositions and systems for modifying or editing target nucleic acids. Background Technology
[0006] Type II CRISPR endonucleases, including the widely used SpCas9, share a common DNA cleavage mechanism. This family of enzymes contains two nuclease domains (HNH and RuvC), each of which cleaves a single DNA strand. When the Cas9-sgRNA complex (or the Cas9-crRNA-trRNA complex) binds to its target DNA sequence, the target DNA strand binds to the RNA spacer sequence, while the non-target DNA strand forms a single-stranded loop. The HNH domain of Cas9 cleaves the target DNA strand, and the RuvC domain cleaves the non-target strand (…). Figure 1 ).like Figure 1 As shown, for type II CRISPR endonucleases (e.g., Cas9), the target and non-target DNA strands are simultaneously cleaved by the HNH and RuvC domains, respectively, resulting in blunt-end double-strand breaks.
[0007] Unlike type II CRISPR endonucleases, type V CRISPR endonucleases (such as Cas12a) have only a single nuclease domain that sequentially cleaves both DNA strands, starting with the non-target strand. Figure 2 As shown, for type V endonucleases (e.g., Cas12a), the RuvC domain sequentially cleaves the non-target and target DNA strands, resulting in cross-strand breaks.
[0008] While Type II and Type V CRISPR endonucleases perform similar functions, their mechanisms and structures are highly divergent. The two different types are thought to have originated from different precursors, and only the RuvC domain shares any significant sequence or structural homology between the two types. Type V CRISPR endonucleases lack the HNH domain responsible for target strand cleavage in Type II enzymes. Instead, in Type V CRISPR endonucleases, the RuvC domain sequentially cleaves both DNA strands Figure 2 ), starting with the non-target strand. Thus, mutation of the catalytic residues of the RuvC domain prevents all nuclease activity and produces an inactivated enzyme, rather than a target strand nicking enzyme. One non-target strand nicking enzyme mutation has been identified in Cas12a; however, this mutation is outside of the RuvC domain and is thought to act by reducing the overall catalytic efficiency of the enzyme. No Type V CRISPR target strand nicking enzyme exists, and given the differences in structure and mechanism of action between Type V CRISPR endonucleases and Type II CRISPR endonucleases, there is no clear method to produce one. SUMMARY
[0009] A first aspect of the application relates to an engineered protein comprising: a first polypeptide that is a first portion of a modified protein, wherein the modified protein comprises an amino acid sequence that is at least 80%, 85%, 90%, or 95% identical to the amino acid sequence of SEQ ID NO: 180 (LbCas12a), and the modified protein comprises a mutation at one or more positions selected from the group consisting of N100, K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, N211, N263, T296, E330, K387, A404, D405, D423, E484, L498, N527, Q529, G532, D535, K538, E539, D541, Y542, Y553, Y554, D572, L585, K591, M592, K595, V596, S599, K600, K601, Y616, Y646, W649, and any combination thereof, with reference to the position numbering of SEQ ID NO: 180, relative to the amino acid sequence of SEQ ID NO: 180 (e.g., the optimal alignment with SEQ ID NO: 180); and a second polypeptide that is heterologous to the first polypeptide and is not a Type V CRISPR-Cas effector polypeptide, wherein the first polypeptide and the second polypeptide are different from each other, and wherein the engineered protein comprises the mutation. In some embodiments, the first polypeptide comprises the mutation.
[0010] Another aspect of the application relates to an engineered protein comprising: an amino acid sequence that is at least 80%, 85%, 90%, or 95% identical to the amino acid sequence of SEQ ID NO: 131 (SYN3298); a mutation at one or more positions in the amino acid sequence of SEQ ID NO: 131 (SYN3298) (e.g., in the optimal alignment with SEQ ID NO: 131) selected from the group consisting of N100, K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, N211, N263, T444, E478, K535, A552, D553, D571, E632, L646, N675, Q677, G680, D683, K686, E687, D689, Y690, Y701, Y702, D720, L733, K739, M740, K743, V744, S747, K748, K749, Y764, Y794, W797, and any combination thereof, with reference to the position numbering of SEQ ID NO: 131.
[0011] Further aspects of the application relate to a composition (e.g., a base editing composition) or system comprising: an engineered protein as described herein; a guide nucleic acid (e.g., a guide RNA); and optionally a deaminase, optionally wherein the engineered protein, the guide nucleic acid, and the optional deaminase form or are comprised in a complex.
[0012] Another aspect of the application relates to a complex comprising: an engineered protein as described herein; a guide nucleic acid (e.g., a guide RNA); and optionally a deaminase. In some embodiments, the deaminase can be fused and / or linked to the engineered protein.
[0013] Further aspects of the application relate to a nucleic acid molecule comprising a nucleotide sequence encoding an engineered protein as described herein.
[0014] Further aspects of the application relate to a method of modifying a target nucleic acid, the method comprising: contacting the target nucleic acid with: an engineered protein as described herein, and a guide nucleic acid (e.g., a guide RNA), optionally wherein the engineered protein and the guide nucleic acid form or are comprised in a complex, thereby modifying the target nucleic acid.
[0015] Another aspect of the invention relates to a method for increasing the efficiency of modifying a target nucleic acid, the method comprising: contacting the target nucleic acid with an engineered protein as described herein, and a guide nucleic acid (e.g., guide RNA), optionally wherein the engineered protein and the guide nucleic acid form a complex or are contained in a complex, thereby modifying the target nucleic acid.
[0016] The present invention further provides expression cassettes and / or vectors comprising the nucleic acid constructs of the present invention, and cells comprising the polypeptides, fusion proteins, and / or nucleic acid constructs of the present invention. Additionally, the present invention provides kits comprising the nucleic acid constructs of the present invention, and expression cassettes, vectors, and / or cells comprising said nucleic acid constructs.
[0017] It should be noted that aspects of the invention described in relation to one embodiment may be incorporated into different embodiments, even if not explicitly described therewith. That is, all embodiments and / or features of any embodiment may be combined in any manner and / or combination. The applicant reserves the right to amend any initially filed claim and / or accordingly file any new claim, including revising any initially filed claim to make it subordinate to and / or incorporated into any feature of any other claim or claims, even if not initially claimed in this manner. These and other objects and / or aspects of the invention will be set forth in detail in the description below. Further features, advantages, and details of the invention will be understood by those skilled in the art upon reading the accompanying drawings and detailed description of the preferred embodiments below, such descriptions being illustrative only. Attached Figure Description
[0018] Figure 1 It is a diagram depicting the mechanism of action of type II CRISPR endonuclease.
[0019] Figure 2 It is a diagram depicting the mechanism of action of type V CRISPR endonuclease.
[0020] Figure 3 This is the crystal structure of SpCas9 (PDB ID 4UN3) binding to a single-guide RNA (sgRNA) and target DNA. The domains shown are as follows: RuvC, bridging helix, Rec1, Rec2, HNH, and PAM interaction.
[0021] Figure 4 This is a diagram showing the Cas12a domain viewed facing the REC leaf. From this perspective, a portion of the crRNA / target DNA duplex is visible exposed on the surface of the Cas12a.
[0022] Figure 5is an overlay of the HNH domain of SpCas9 onto candidate insertion sites in LbCas12a.
[0023] Figure 6 is a graph depicting E. coli expression of HNH-3287, 3288, 3289, 3290, 3296, 3297, 3298, and 3299, cleaved and soluble fraction.
[0024] Figure 7 is a graph depicting nicking activity of purified HNH-3287, 3288, 3289, 3290, 3296, 3297, and 3298.
[0025] Figure 8 is an image of a gel showing that a nickase according to some embodiments of the application is expressed in soluble form in E. coli.
[0026] Figure 9 is an image of a gel showing that a nickase according to some embodiments of the application can nick a DNA substrate.
[0027] Figure 10 is an image of a gel showing that a nickase according to some embodiments of the application can be RNA dependent.
[0028] Figure 11 is an image of a gel showing that a nickase according to some embodiments of the application can act as a DNA nickase.
[0029] Figure 12 is a graph showing a labeled target strand.
[0030] Figure 13 is a graph showing a labeled non-target strand.
[0031] Figure 14 is an image of a gel including samples incubated with a labeled target strand.
[0032] Figure 15 is an image of a gel including samples incubated with a labeled non-target strand.
[0033] Figure 16 is a graph showing Figure 14 and Figure 15 lanes, and a control lane, of a whole gel.
[0034] Figure 17 is a graph showing editing efficiency of respective pairs of enzymes according to some embodiments of the application.
[0035] Figures 18-21 is a graph showing the percent C to T editing for each target region corresponding to the respective interval: FANCF Interval 1 Figure 18 ), FANCF Interval 2 Figure 19 ), AAVS1 Interval 1 Figure 20 ), and AAVS1 Interval 2 Figure 21 ).
[0036] Figures 22-23 is a graph showing the percent A to G editing for each target region corresponding to the respective interval: RNF2 Interval 1 Figure 22 ), and RNF2 Interval 2 Figure 23 ).
[0037] Figure 24 is a graph showing the indel generation frequency of RR-LbCas12a and three synthetic enzymes (3287RR, 3288RR, and 3298RR) at known non-native protospacer adjacent motif (PAM) sequences, each including G532R and K595R (RR) mutations with reference to position numbering of SEQ ID NO: 180 (LbCas12a).
[0038] Figure 25 is a graph showing the indel generation frequency of LbCas12a, RR-LbCas12a, 3287RR, 3288RR, and 3298RR at native TTV PAM sequences.
[0039] Figure 26 is a graph showing the indel generation frequency of RVR-LbCas12a (with G532R, K538V, and Y542R (RVR) mutations) and RVR 3287, RVR 3288, and RVR 3298 at known non-native PAM sequences, each including G532R, K538V, and Y542R (RVR) mutations with reference to position numbering of SEQ ID NO: 180 (LbCas12a).
[0040] Figure 27 is a graph showing the indel generation frequency of LbCas12a, RVR-LbCas12a (with G532R, K538V, and Y542R (RVR) mutations) and 3287RVR, 3288RVR, and 3298RVR at native TTV PAM sequences, each including G532R, K538V, and Y542R (RVR) mutations with reference to position numbering of SEQ ID NO: 180 (LbCas12a).
[0041] Figures 28-31is a chart showing biological replicates of indel generation frequency of LbCas12a, 3287, 3288, and 3298, and 3287D156R, 3288D156R, and 3298D156R at the native TTTV PAM sequence, including the D156R mutation with position numbering referenced to SEQ ID NO: 180 (LbCas12a).
[0042] Figure 32 is a chart showing indel generation frequency of RVR-LbCas12a, RVR 3287, RVR 3288, and RVR 3298 at known non-native PAM sequences.
[0043] Figure 33 is a chart showing indel generation frequency of LbCas12a, RVR-LbCas12a, 3287RVR, 3288RVR, and 3298RVR at the native TTTV PAM sequence. DETAILED DESCRIPTION
[0044] The present application will now be described below with reference to the drawings and examples, in which embodiments of the present application are shown. This description is not intended to be a detailed catalog of all the different ways in which the application can be implemented or of all the features that can be added to the application. For example, features illustrated with respect to one embodiment can be incorporated into other embodiments, and features illustrated with respect to a particular embodiment can be deleted from that embodiment. Thus, the present application contemplates that in some embodiments of the application, any feature or combination of features set forth herein can be excluded or omitted. Additionally, many modifications and additions can be made to the embodiments proposed herein by those of ordinary skill in the art given the benefit of this disclosure, without departing from the spirit and scope of the present application. Accordingly, the following description is intended for illustrating some specific embodiments of the application and is not intended to be exhaustive or to limit the application to the precise forms disclosed. It is therefore evident that the application can be altered or modified in various ways without departing from the scope and spirit of the application.
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0046] All publications, patent applications, patents and other references cited herein are incorporated by reference in their entireties for the teachings relevant to the sentence and / or paragraph in which the reference is presented.
[0047] Unless otherwise indicated herein, the various features of the present application described herein can be used in any combination. Moreover, the present application also contemplates that in some embodiments of the present application, any feature or combination of features set forth herein can be excluded. To illustrate, if the specification states that a composition comprises components A, B, and C, it is specifically intended that A, B, or C, or any combination thereof, can be omitted and / or excluded in any embodiment of the application.
[0048] As used in the description of the application and the appended claims, the singular forms "a", "an" and "the" are intended to include plural forms as well, unless the context clearly indicates otherwise.
[0049] Also as used in the description of the application and the appended claims, the term "and / or" means and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or").
[0050] As used herein, the term "about" when referring to a measurable value such as an amount or concentration, is intended to encompass variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the specified value, in addition to the specified value. For example, "about X", where X is a measurable value, is intended to include X and variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of X. Ranges of values provided herein are intended to include any and all subranges of the recited ranges of values, and / or individual values that fall within the recited ranges of values.
[0051] As used herein, phrases such as "between X and Y" and "between about X and Y" should be interpreted to include X and Y. As used herein, phrases such as "between about X and Y" mean "between about X and about Y," and phrases such as "from about X to Y" mean "from about X to about Y."
[0052] Unless otherwise indicated herein, the recitation of a range of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, and each individual value is incorporated in the specification as if it were individually recited herein. For example, if a range 10 to 15 is disclosed, then 11, 12, 13, and 14 are also disclosed.
[0053] As used herein, the terms "comprise", "comprises" and "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0054] As used herein, the transitional phrase “consisting essentially of’ means that the scope of a claim should be interpreted as encompassing the specified materials or steps recited in the claim and those that do not materially affect the basic and novel characteristic(s) of the claimed application. Thus, the term “consisting essentially of’ when used in the claims of the present application is not intended to be construed as equivalent to “comprising.”
[0055] As used herein, the terms “increase,” “increasing,” “enhance,” “enhancing,” “improve,” and “improving” (and grammatical variations thereof) describe an improvement of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 150%, 200%, 300%, 400%, 500%, or more, as compared to another measurable property or quantity (e.g., a control value).
[0056] As used herein, the terms “reduce,” “reduced,” “reducing,” “reduction,” “decrease,” and “decreasing” (and grammatical variations thereof) describe a decrease of, for example, at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%, as compared to another measurable property or quantity (e.g., a control value). In some embodiments, a reduction can result in no or substantially no (i.e., an amount that is minuscule, such as less than about 10% or even 5%) detectable activity or quantity.
[0057] A “heterologous nucleotide sequence” or “recombinant nucleotide sequence” is a nucleotide sequence that is not naturally associated with the host cell into which it is introduced, including non-naturally occurring multiple copies of a naturally occurring nucleotide sequence.
[0058] A “native” or “wild-type” nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence refers to a naturally occurring or endogenous nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence. Thus, for example, a “native nucleic acid” is a nucleic acid that naturally occurs in or is endogenous to a reference organism. A “homologous” nucleic acid sequence is a nucleotide sequence that is naturally associated with the host cell into which it is introduced.
[0059] As used herein, the terms “nucleic acid,” “nucleic acid molecule,” “nucleotide sequence,” and “polynucleotide” refer to linear or branched, single-stranded or double-stranded RNA or DNA or hybrids thereof. The terms also encompass RNA / DNA hybrids. When dsRNA is produced synthetically, less common bases such as inosine, 5-methylcytosine, 6-methyladenine, hypoxanthine, etc., may also be used for antisense, dsRNA, and ribozyme pairing. For example, polynucleotides containing C-5 propyne analogs of uridine and cytidine have been shown to bind RNA with high affinity and are potent antisense inhibitors of gene expression. Other modifications may also be made, such as modifications to the phosphodiester backbone or the 2'-hydroxyl group in the RNA ribose group.
[0060] As used herein, the term "nucleotide sequence" refers to a hybrid of nucleotides or the sequence of these nucleotides from the 5' to 3' ends of a nucleic acid molecule, and includes DNA or RNA molecules, including cDNA, DNA fragments or portions, genomic DNA, synthetic (e.g., chemically synthesized) DNA, plasmid DNA, mRNA, and antisense RNA, any of which may be single-stranded or double-stranded. The terms "nucleotide sequence," "nucleic acid," "nucleic acid molecule," "nucleic acid construct," "recombinant nucleic acid," "oligonucleotide," and "polynucleotide" are used interchangeably herein and refer to a hybrid of nucleotides. The nucleic acid molecules and / or nucleotide sequences provided herein are presented from left to right in a 5' to 3' orientation and use the standard code representation for nucleotide characters as set forth in U.S. Sequence Rules 37 CFR § § 1.831-1.835 and World Intellectual Property Organization (WIPO) Standard ST. 26. As used herein, "5' region" may refer to the polynucleotide region closest to the 5' end of a polynucleotide. Therefore, for example, elements in the 5' region of a polynucleotide can be located at any position from the first nucleotide at the 5' end of the polynucleotide to the nucleotide in the middle of the polynucleotide. As used herein, "3' region" can refer to the polynucleotide region closest to the 3' end of the polynucleotide. Therefore, for example, elements in the 3' region of a polynucleotide can be located at any position from the first nucleotide at the 3' end of the polynucleotide to the nucleotide in the middle of the polynucleotide.
[0061] As used herein, the term "gene" refers to a nucleic acid molecule capable of producing mRNA, antisense RNA, miRNA, antimicroRNA antisense oligodeoxyribonucleotides (AMO), etc. A gene may or may not be capable of producing a functional protein or gene product. A gene may include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and / or 5' and 3' untranslated regions).
[0062] A polynucleotide, gene, or polypeptide can be "isolated" in that the nucleic acid or polypeptide is substantially or essentially free from components which normally accompany or are present in association with the nucleic acid or polypeptide as found in its natural state. In some embodiments, such components include other cellular material, culture medium from recombinant production, and / or various chemicals used to synthesize the nucleic acid or polypeptide by chemical methods.
[0063] The term "mutation" refers to point mutations (e.g., missense or nonsense, or insertion or deletion of a single base pair that causes a frameshift), insertions, deletions, and / or truncations. When a mutation is a substitution of one residue for another within an amino acid sequence, or a deletion or insertion of one or more residues within a sequence, the mutation is typically described by identifying the original residue, followed by identifying the position of the residue within the sequence, and the identity of the new substituted residue.
[0064] As used herein, the term "complementary" or "complementarity" refers to the natural binding of polynucleotides via base pairing under the permitted salt and temperature conditions. For example, the sequence "A-G-T" (5' to 3') binds to the complementary sequence "T-C-A" (3' to 5'). Complementarity between two single-stranded molecules can be "partial," in which only some of the nucleotides bind, or it can be complete when total complementarity exists between the single-stranded molecules. The degree of complementarity between nucleic acid strands has significant effects on the efficiency and strength of hybridization between nucleic acid strands.
[0065] As used herein, "complementary" can mean 100% complementarity to a comparison nucleotide sequence, or can mean less than 100% complementarity (e.g., "substantially complementary," such as about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, etc. complementarity).
[0066] A "portion" or "fragment" of a nucleotide sequence or polypeptide (including a domain) is understood to mean a nucleotide sequence or polypeptide of reduced length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more residues (e.g., nucleotides or peptides) reduced relative to a reference nucleotide sequence or polypeptide, respectively) and comprising, consisting essentially of and / or consisting of contiguous residues of the nucleotide sequence or polypeptide that are the same as or nearly identical to (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical) the reference nucleotide sequence or polypeptide, respectively. In some embodiments, a portion of a reference nucleotide sequence or polypeptide is about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or more of the full-length reference nucleotide sequence or polypeptide. Such nucleic acid fragments or portions according to the application can be included in larger polynucleotides of which the nucleic acid fragment or portion is a component, where appropriate. As an example, a repeat sequence of a guide nucleic acid of the application can comprise a portion of a wild-type CRISPR-Cas repeat sequence (e.g., a wild-type type V CRISR Cas repeat sequence, such as a repeat sequence from a CRISPR Cas system (including, but not limited to, Casl2a (Cpfl), Casl2b, Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2g, Casl2h, Casl2i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Casl4a, Casl4b, and / or Casl4c, etc.). Similarly, a portion of a polypeptide can be included in a larger polypeptide of which the portion of the polypeptide is a component.
[0067] Different nucleic acids or proteins having homology are referred to herein as "homologs." The term homologs includes homologous sequences from the same and other species as well as orthologous sequences from the same and other species. "Homology" refers to the level of similarity between two or more nucleic acid and / or amino acid sequences, expressed as a percentage of positional identity (i.e., sequence similarity or identity). Homology also refers to the concept of similar functional properties between different nucleic acids or proteins. Accordingly, the compositions and methods of the present application further comprise homologs of the nucleotide sequences and polypeptides of the present application. As used herein, "ortholog" and "orthologous" refer to homologous nucleotide sequences and / or amino acid sequences in different species that arose from a common ancestral gene during speciation. Homologs or orthologs of the nucleotide sequences of the present application have substantial sequence identity (e.g., at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%) to the nucleotide sequences of the present application.
[0068] As used herein, "sequence identity" refers to the extent to which two optimally aligned polynucleotide or polypeptide sequences are invariant throughout a window of alignment of components (e.g., nucleotides or amino acids). "Identity" can be readily calculated by known methods, including, but not limited to, those described in Computational Molecular Biology (Lesk, A. M., ed.) Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, D. W., ed.) Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, A. M., and Griffin, H. G., eds.) Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G., ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.) Stockton Press, New York (1991).
[0069] As used herein, the term "percent sequence identity" or "percent identity" refers to the percentage of identical linear polynucleotide sequences in a reference ("query") polynucleotide molecule (or its complementary strand) as compared to a test ("subject") polynucleotide molecule (or its complementary strand) when the two sequences are optimally aligned. In some embodiments, "percent identity" can refer to the percentage of identical amino acids in an amino acid sequence as compared to a reference polypeptide.
[0070] As used herein, in the context of two nucleic acid molecules, nucleotide sequences, or protein sequences, the phrase "substantially identical" or "substantial identity," when compared and aligned for maximum correspondence, as determined using one of the following sequence comparison algorithms or by visual inspection, means that two or more sequences or subsequences have at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% nucleotide or amino acid residue identity. In some embodiments of the application, substantial identity exists within a region of contiguous nucleotides of a nucleotide sequence of the application, the region being of about 10 nucleotides to about 20 nucleotides, about 10 nucleotides to about 25 nucleotides, about 10 nucleotides to about 30 nucleotides, about 15 nucleotides to about 25 nucleotides, about 30 nucleotides to about 40 nucleotides, about 50 nucleotides to about 60 nucleotides, about 70 nucleotides to about 80 nucleotides, about 90 nucleotides to about 100 nucleotides or more in length, and any range therein, up to the full length of the sequence. In some embodiments, the nucleotide sequence can be substantially identical over at least about 20 nucleotides (e.g., about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 nucleotides). In some embodiments, a nucleotide or protein sequence that is substantially identical performs substantially the same function as the nucleotide (or encoded protein sequence) that it is substantially identical to.
[0071] For sequence comparison, typically one sequence acts as the reference sequence to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are input into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. The sequence comparison algorithm then calculates the percent sequence identity for the test sequence(s) relative to the reference sequence, based on the program parameters.
[0072] Optimal alignment of sequences for comparison can be conducted using tools that are well known to those skilled in the art, and can be performed, for example, by the local homology algorithm of Smith and Waterman, by the homology alignment algorithm of Needleman and Wunsch, by the search for similarity method of Pearson and Lipman, by any sa tool, and optionally by computerized implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI), or by visual inspection. Wisconsin Wisconsin GAP, BESTFIT, FASTA, and TFASTA, Clustal Omega, EMBOSS Needle, EMBOSS Stretcher, EMBOSS Water, LALIGN, GGSEARCH 2SEQ, EMBOS Cons, Kalign, MAFFT, MUSCLE, and / or T-Coffee, available as part of the DNASTAR® software suite (DNASTAR Inc., Madison, WI) and web-based alignment programs such as Clustal Omega, EMBOSS Needle, EMBOSS Stretcher, EMBOSS Water, LALIGN, GGSEARCH 2SEQ, EMBOS Cons, Kalign, MAFFT, MUSCLE, and T-Coffee. In some embodiments, the "best alignment" of two sequences (e.g., two polypeptide sequences) is the alignment that scores the highest, optionally from an alignment by tools such as the local homology algorithm of Smith and Waterman, the homology alignment algorithm of Needleman and Wunsch, the similarity search method of Pearson and Lipman, available as part of the DNASTAR® software suite (DNASTAR Inc., Madison, WI) and web-based alignment programs such as Clustal Omega, EMBOSS Needle, EMBOSS Stretcher, EMBOSS Water, LALIGN, GGSEARCH 2SEQ, EMBOS Cons, Kalign, MAFFT, MUSCLE, and T-Coffee. Wisconsin Wisconsin GAP, BESTFIT, FASTA, and TFASTA, Clustal Omega, EMBOSS Needle, EMBOSS Stretcher, EMBOSS Water, LALIGN, GGSEARCH 2SEQ, EMBOS Cons, Kalign, MAFFT, MUSCLE, and / or T-Coffee, available as part of the DNASTAR® software suite (DNASTAR Inc., Madison, WI) and web-based alignment programs such as Clustal Omega, EMBOSS Needle, EMBOSS Stretcher, EMBOSS Water, LALIGN, GGSEARCH 2SEQ, EMBOS Cons, Kalign, MAFFT, MUSCLE, and T-Coffee. In some embodiments, the "best alignment" of two sequences (e.g., two polypeptide sequences) is the alignment that scores the highest, optionally from an alignment by tools such as the local homology algorithm of Smith and Waterman, the homology alignment algorithm of Needleman and Wunsch, the similarity search method of Pearson and Lipman, available as part of the DNASTAR® software suite (DNASTAR Inc., Madison, WI) and web-based alignment programs such as Clustal Omega, EMBOSS Needle, EMBOSS Stretcher, EMBOSS Water, LALIGN, GGSEARCH 2SEQ, EMBOS Cons, Kalign, MAFFT, MUSCLE, and T-Coffee.
[0073] Two nucleotide sequences can also be considered substantially complementary when they hybridize to each other under stringent conditions. In some representative embodiments, two nucleotide sequences that are considered substantially complementary hybridize to each other under highly stringent conditions.
[0074] In the context of nucleic acid hybridization experiments, such as Southern and Northern hybridizations, "stringent hybridization conditions" and "stringent hybridization wash conditions" are sequence dependent, and vary under different ambient parameters. A broad guide to nucleic acid hybridization is found in Tijssen, Laboratory Techniques in Biochemistry and Molecular Biology - Hybridization with Nucleic Acid Probes, Part I, Chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assays", Elsevier, New York (1993). Typically, highly stringent hybridization and wash conditions are chosen to be about 5°C lower than the thermal melting point (Tm) of the particular sequence at defined ionic strength and pH. m
[0075] T m is the temperature at which 50% of the target sequence hybridizes to a perfectly matched probe (at defined ionic strength and pH). Very stringent conditions are chosen to be equal to the Tmof the particular probe. m In Southern or Northern blots, an example of stringent hybridization conditions in which the complementary nucleotide sequences of more than 100 complementary residues hybridize to the filter is overnight, at 42°C, with 50% formamide and 1 mg of heparin. An example of high stringency wash conditions is 0.15 M NaCl at 72°C. An example of medium stringency wash conditions is 0.2x SSC at 65°C (see Sambrook below for a description of SSC buffer). Typically, low stringency washes are performed prior to high stringency washes to remove background probe signal. For example, an example of low stringency wash for duplexes greater than 100 nucleotides is 4-6x SSC at 40°C. For short probes (e.g., about 10 to 50 nucleotides), stringent conditions will typically involve salt concentrations less than about 1.0 M Na ion, usually about 0.01 to 1.0 M Na ion concentration at pH 7.0 to 8.3, and temperatures of at least about 30°C. Addition of destabilizing agents, such as formamide, can also be used to achieve stringent conditions. Typically, a 2-fold (or greater) difference in signal obtained under otherwise identical conditions, but for the presence of the non-relevant probe, indicates detection of specific hybridization. Nucleotide sequences which will not hybridize to each other under stringent conditions are still essentially identical. Denham, et al. (1993) Biochem. J. 291 : 267- 275. This occurs, e.g., when a copy of a nucleotide sequence is created using the maximum codon degeneracy permitted by the genetic code.
[0076] The polynucleotides and / or recombinant nucleic acid constructs of the present application can be codon optimized for expression. In some embodiments, the polynucleotides, nucleic acid constructs, expression cassettes, and / or vectors of the present application (e.g., which comprise / encode a polypeptide of the present application (e.g., an engineered protein), a nucleic acid binding polypeptide (e.g., a DNA binding polypeptide such as a sequence-specific DNA binding domain from a polynucleotide-guided endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), an Argonaute protein, and / or a CRISPR-Cas effector protein), a guide nucleic acid, a cytosine deaminase, and / or an adenine deaminase) can be codon optimized for expression in an organism (e.g., an animal such as a human, a plant, a fungus, an archaeon, or a bacterium). In some embodiments, a codon-optimized nucleic acid construct, polynucleotide, expression cassette, and / or vector of the present application has about 70% to about 99.9% (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100%) or more identity to a reference nucleic acid construct, polynucleotide, expression cassette, and / or vector that has not been codon optimized.
[0077] In any of the embodiments described herein, the polynucleotides or nucleic acid constructs of the present application can be operably associated with a variety of promoters and / or other regulatory elements for expression in an organism or a cell thereof (e.g., a mammal and / or a mammalian cell, a plant and / or a plant cell, etc.). Thus, in some embodiments, the polynucleotides or nucleic acid constructs of the present application can further comprise one or more promoters, introns, enhancers, and / or terminators operably linked to one or more nucleotide sequences. In some embodiments, a promoter can be operably associated with an intron (e.g., the Ubil promoter and intron). In some embodiments, a promoter associated with an intron can be referred to as a “promoter region” (e.g., the Ubil promoter and intron).
[0078] As used herein, “operably linked” or “operably associated” in reference to polynucleotides means that the indicated elements are functionally related to each other, and are typically also physically related. Thus, as used herein, the term “operably linked” or “operably associated” refers to nucleotide sequences that are functionally associated on a single nucleic acid molecule. Thus, a first nucleotide sequence that is operably linked to a second nucleotide sequence refers to the situation where the first nucleotide sequence is in a functional relationship with the second nucleotide sequence. For example, a promoter is operably associated with a nucleotide sequence if the promoter affects the transcription or expression of the nucleotide sequence. Those skilled in the art will appreciate that a control sequence (e.g., a promoter) need not be contiguous with the nucleotide sequence with which it is operably associated, so long as the control sequence functions to direct the expression of the nucleotide sequence. Thus, for example, there can be intervening untranslated but transcribed nucleic acid sequences between a promoter and a nucleotide sequence, and the promoter can still be considered to be “operably linked” to the nucleotide sequence.
[0079] As used herein, the term “linked” or “fused” in reference to polypeptides refers to the linkage of one polypeptide to another polypeptide. A polypeptide can be linked or fused to another polypeptide (at the N- or C-terminus) either directly (e.g., by a peptide bond) or through a linker (e.g., a peptide linker).
[0080] The term “linker” in reference to polypeptides is art-recognized and refers to a chemical group or molecule that links two molecules or moieties, such as two polypeptides or domains of a fusion protein, such as a CRISPR-Cas effector protein and a peptide tag and / or a polypeptide of interest. A linker can be composed of a single linking molecule (e.g., a single amino acid), or can comprise more than one linking molecule. In some embodiments, a linker can be an organic molecule, group, polymer, or chemical moiety, such as a bivalent organic moiety. In some embodiments, a linker can be an amino acid, or can be a peptide. In some embodiments, a linker is a peptide (e.g., a peptide linker).
[0081] In some embodiments, the length of the peptide linker useful in the present application can be from about 2 to about 100 or more amino acids, for example, from about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids in length (e.g., from about 2 to about 40, about 2 to about 50, about 2 to about 60, about 4 to about 40, about 4 to about 50, about 4 to about 60, about 5 to about 40, about 5 to about 50, about 5 to about 60, about 9 to about 40, about 9 to about 50, about 9 to about 60, about 10 to about 40, about 10 to about 50, about 10 to about 60 amino acids, or from about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids) (e.g.,about 105, 110, 115, 120, 130, 140, 150, or more amino acids). In some embodiments, the peptide linker can be a GS linker. In some embodiments, the peptide linker is a GS linker having 1, 2, 3, or 4 amino acid residues, optionally having 2 or 4 amino acid residues. In some embodiments, the peptide linker has one of the amino acid sequences of SEQ ID NOs: 18-47 or 176-179. In some embodiments, the peptide linker can comprise the following amino acid sequence: (GGS), n , GS, SG, GSSG (SEQ ID NO: 175), GSSGSS (SEQ ID NO: 176), GSSGSSGS (SEQ ID NO: 177), (GSS) n (SEQ ID NO: 178), (GSS) n GS (SEQ ID NO: 179), S(GGS) n (SEQ ID NO: 42), SGGS (SEQ ID NO: 43), or (GGGGS)n(SEQ ID NO: 44), where n is an integer from 1-20 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20). In some embodiments, the peptide linker can comprise the following amino acid sequence: SGGSGGSGGS (SEQ ID NO: 45). In some embodiments, the peptide linker can comprise the following amino acid sequence: SGSETPGTSESATPES (SEQ ID NO: 46), also referred to as an XTEN linker. In some embodiments, the peptide linker can comprise the following amino acid sequence: SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 47), also referred to as a GS-XTEN-GS linker.
[0082] As used herein, the term "linked" or "fused" in reference to polynucleotides refers to the joining of one polynucleotide to another polynucleotide. In some embodiments, two or more polynucleotide molecules can be linked by a linker, which can be an organic molecule, group, polymer, or chemical moiety, such as a divalent organic moiety. A polynucleotide can be linked or fused to another polynucleotide (at the 5' end or 3' end) by covalent or non-covalent linkage or association, including, for example, Watson-Crick base pairing or by one or more linking nucleotides. In some embodiments, a polynucleotide motif of a certain structure can be inserted into another polynucleotide sequence (e.g., an extension of a hairpin structure in a guide RNA). In some embodiments, the linking nucleotide can be a naturally occurring nucleotide. In some embodiments, the linking nucleotide can be a non-naturally occurring nucleotide.
[0083] A "promoter" is a nucleotide sequence that controls or regulates transcription of a nucleotide sequence (e.g., a coding sequence) with which the promoter is operably associated. A coding sequence controlled or regulated by a promoter can encode a polypeptide and / or a functional RNA. Typically, a "promoter" refers to a nucleotide sequence that contains a binding site for RNA polymerase II and directs the initiation of transcription. Typically, a promoter is located 5' or upstream relative to the start of the coding region of a corresponding coding sequence. A promoter can include other elements that act as regulators of gene expression, such as a promoter region. These include a TATA box consensus sequence, and often include a CAAT box consensus sequence (Breathnach and Chambon, (1981) Annu. Rev. Biochem. 50:349). In plants, the CAAT box can be replaced by an AGGA box (Messing et al., (1983) Genetic Engineering of Plants, T. Kosuge, C. Meredith and A. Hollaender (eds.), Plenum Press, pp. 211-227). In some embodiments, a promoter region can include at least one intron (e.g., SEQ ID NO: 48 or SEQ ID NO: 49).
[0084] Promoters useful in the present application can include, for example, constitutive, inducible, temporally regulated, developmentally regulated, chemically regulated, tissue-preferred, and / or tissue-specific promoters, for making recombinant nucleic acid molecules, e.g., "synthetic nucleic acid constructs" or "protein-RNA complexes." These various types of promoters are known in the art.
[0085] The choice of promoter can vary depending on the temporal and spatial requirements for expression, as well as the host cell to be transformed. Promoters for many different organisms are well known in the art. Based on the extensive knowledge that exists in the art, an appropriate promoter can be selected for a particular host organism of interest. Thus, for example, promoters upstream of genes that are highly constitutively expressed in model organisms are well known and such knowledge is readily available and implemented in other systems as appropriate.
[0086] In some embodiments, promoters that are functional in plants can be used with the constructs of the application. Non-limiting examples of promoters that can be used to drive expression in plants include the promoter of the RubisCo small subunit gene 1 (PrbcS1), the promoter of the actin gene (Pactin), the promoter of the nitrate reductase gene (Pnr), and the promoter of the repetitive carbonic anhydrase gene 1 (Pdca1) (see Walker et al. Plant Cell Rep. 23:727-735 (2005); Li et al. Gene 403:132-142 (2007); Li et al. Mol Biol. Rep. 37:1143-1154 (2010)). PrbcS1 and Pactin are constitutive promoters, and Pnr and Pdca1 are inducible promoters. Pnr is induced by nitrate and repressed by ammonium (Li et al. Gene 403:132-142 (2007)), and Pdca1 is induced by salt (Li et al. Mol Biol. Rep. 37:1143-1154 (2010)). In some embodiments, a promoter that can be used in the application is an RNA polymerase II (Pol II) promoter. In some embodiments, a U6 promoter or a 7SL promoter from Zea mays can be used in the constructs of the application. In some embodiments, a U6c promoter and / or a 7SL promoter from Zea mays can be used to drive expression of a guide nucleic acid. In some embodiments, a U6c promoter, a U6i promoter, and / or a 7SL promoter from Glycine max can be used in the constructs of the application. In some embodiments, a U6c promoter, a U6i promoter, and / or a 7SL promoter from Glycine max can be used to drive expression of a guide nucleic acid.
[0087] Examples of constitutive promoters useful for plants include, but are not limited to, the cestrum virus promoter (cmp) (U.S. Patent No. 7,166,770), the rice actin 1 promoter (Wang et al. (1992) Mol. Cell. Biol. 12:3399-3406; and U.S. Patent No. 5,641,876), the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812), the CaMV 19S promoter (Lawton et al. (1987) Plant Mol. Biol. 9:315-324), the nos promoter (Ebert et al. (1987) Proc. Natl. Acad. Sci USA 84:5745-5749), the Adh promoter (Walker et al. (1987) Proc. Natl. Acad. Sci USA 84:6624-6629), the sucrose synthase promoter (Yang and Russell (1990) Proc. Natl. Acad. Sci USA 87:4144-4148), and the ubiquitin promoter. Constitutive promoters derived from ubiquitin accumulate in many cell types. Ubiquitin promoters have been cloned from several plant species for use in transgenic plants, for example sunflower (Binet et al., 1991. Plant Science 79:87-94), maize (Christensen et al., 1989. Plant Mol. Biol. 12:619-632), and arabidopsis (Norris et al. 1993. Plant Mol. Biol. 21 :895-906). The maize ubiquitin promoter (UbiP) has been developed in a transgenic monocot system and its sequence and vectors constructed for monocot transformation are disclosed in European Patent Publication EP 0 342 926. The ubiquitin promoter is suitable for expression of the nucleotide sequences of the application in transgenic plants, especially monocots. In addition, the promoter expression cassette described by McElroy et al. (Mol. Gen. Genet. 231 :150-160 (1991)) can be readily modified for expression of the nucleotide sequences of the application and is particularly suitable for use in monocot hosts.
[0088] In some embodiments, tissue-specific / tissue-preferred promoters can be used to express a heterologous polynucleotide in a plant cell. Tissue-specific or -preferred expression patterns include, but are not limited to, green tissue-specific or -preferred, root-specific or -preferred, stem-specific or -preferred, flower-specific or -preferred, or pollen-specific or -preferred. Promoters suitable for expression in green tissues include many of the promoters that regulate genes involved in photosynthesis, and many of these are cloned from both monocots and dicots. In one embodiment, a promoter useful in the present application is the maize PEPC promoter from the phosphoenolpyruvate carboxylase gene (Hudspeth and Grula, Plant Mol. Biol. 12:579-589 (1989)). Non-limiting examples of tissue-specific promoters include tissue-specific promoters associated with genes encoding seed storage proteins (such as beta-conglycinin, cruciferin, napin, and legumin), zein or oil body proteins (such as oleosin), or proteins involved in fatty acid biosynthesis (including acyl carrier protein, stearoyl-ACP desaturase, and fatty acid desaturase (fad 2-1)), as well as other nucleic acids expressed during embryo development (such as Bce4, see, e.g., Kridl et al. (1991) Seed Sci. Res. 1 :209-219; and EP Patent No. 255378). Tissue-specific or tissue-preferred promoters useful for expressing a nucleotide sequence of the present application in a plant, particularly maize, include, but are not limited to, promoters that direct expression in roots, pith, leaves, or pollen. Such promoters are disclosed, for example, in WO 93 / 07278, which is incorporated herein by reference for its disclosure of promoters.Other non-limiting examples of tissue-specific or tissue-preferred promoters that can be used in the present application are the cotton rubisco promoter disclosed in U.S. Patent 6,040,504; the rice sucrose synthase promoter disclosed in U.S. Patent 5,604,121; the root-specific promoters described by de Framond (FEBS 290: 103-106 (1991); European Patent EP 0452269 to Ciba-Geigy); the stem-specific promoters described in U.S. Patent 5,625,136 (to Ciba-Geigy) and which drive expression of the com trpA gene; the jatropha yellowing curly leaf virus promoter disclosed in WO 01 / 73087; and pollen-specific or -preferred promoters including, but not limited to, ProOsLPS10 and ProOsLPS11 from rice (Nguyen et al. Plant Biotechnol. Reports 9(5): 297-306 (2015)), ZmSTK2_USP from com (Wang et al. Genome 60(6): 485-495 (2017)), LAT52 and LAT59 from tomato (Twell et al. Development 109(3): 705-713 (1990)), Zm13 (U.S. Patent No. 10,421,972), PLA2-delta promoter from Arabidopsis (U.S. Patent No. 7,141,424), and / or ZmC5 promoter from com (International PCT Publication No. WO 1999 / 042587).
[0089] Further examples of plant tissue-specific / tissue-preferred promoters include, but are not limited to, root-hair-specific cis-elements (RHE) (K IMJeong et al. (2006) The Plant Cell 18:2958-2970), root-specific promoters RCc3 (Jeong et al. (2010) Plant Physiol. 153:185-197) and RB7 (U.S. Patent No. 5,459,252), lectin promoters (Lindstrom et al. (1990) Der. Genet. 11 : 160-167; and Vodkin (1983) Prog. Clin. Biol. Res. 138:87-98), maize alcohol dehydrogenase 1 promoter (Dennis et al. (1984) Nucleic Acids Res. 12:3983-4000), S-adenosyl-L-methionine synthetase (SAMS) (Vander Mijnsbrugge et al. (1996) Plant and Cell Physiology, 37(8): 1108-1115), maize light harvesting complex promoter (Bansal et al. (1992) Proc. Natl. Acad. Sci. USA 89:3654-3658), maize heat shock protein promoter (O'Dell et al. (1985) EMBO J. 5:451-458; and Rochester et al. (1986) EMBO J. 5:451-458), pea small subunit RuBP carboxylase promoter (Cashmore, “Nuclear genes encoding the small subunit of ribulose-1,5-bisphosphate carboxylase” pp. 29-39 in Genetic Engineering of Plants (Hollaender, ed. Plenum Press 1983; and Poulsen et al. (1986) Mol. Gen. Genomics 205:193-200), Ti plasmid mannopine synthase promoter (Langridge et al. (1989) Proc. Natl. Acad. Sci. USA 86:3219-3223), Ti plasmid nopaline synthase promoter (Langridge et al. (1989), supra), petunia chalcone isomerase promoter (van Tunen et al. (1988) EMBO J. 7:1257-1263), legume glycine-rich protein 1 promoter (Keller et al. (1989) Genes Dev.The maize ubiquitin-1 promoter (Christou et al. (1992) Proc. Natl. Acad. Sci. USA 89: 6224-6228), the maize heat shock protein promoter (Schilperoot et al. (1988) Plant Physiol. 88: 945-949), the maize sucrose synthase promoter (Vilardel et al. (1996) Transgenic Res. 5: 269-277), the maize 19 kD alpha tubulin promoter (Sullivan et al. (1989) Mol. Gen. Genet. 215: 431-440), the maize 5-adenylate deaminase promoter (Hudspeth and Grula (1989) Plant Mol. Biol. 12: 579-589), the maize R gene complex-related promoter (Chandler et al. (1989) Plant Cell 1 : 1175-1183), and the chalcone synthase promoter (Franken et al. (1991) EMBO J. 10: 2605-2612).
[0090] Useful for seed-specific expression is the pea legumin promoter (Czako et al. (1992) Mol. Gen. Genet. 235: 33-40); and seed-specific promoters disclosed in U.S. Patent No. 5,625,136. Useful for expression in mature leaves is a promoter that switches at the onset of senescence, such as the SAG promoter from Arabidopsis (Gan et al. (1995) Science 270: 1986-1988).
[0091] In addition, promoters that are functional in chloroplasts can be used. Non-limiting examples of such promoters include the bacteriophage T3 gene 9 5' UTR and other promoters disclosed in U.S. Patent No. 7,579,516. Other promoters useful in the present application include, but are not limited to, the S-E9 small subunit RuBP carboxylase promoter and the Kunitz trypsin inhibitor gene promoter (Kti3).
[0092] Additional regulatory elements that can be used in the present application include, but are not limited to, introns, enhancers, termination sequences, and / or 5' and 3' untranslated regions.
[0093] Introns that can be used in the present application can be introns identified in and isolated from plants and then inserted into an expression cassette for plant transformation. As understood by one of skill in the art, introns can contain sequences required for self excision and are incorporated in frame into a nucleic acid construct / expression cassette. Introns can be used as spacers to separate multiple protein coding sequences in one nucleic acid construct, or can be used within one protein coding sequence, for example, to stabilize mRNA. If used within a protein coding sequence, they are inserted "in frame" and include a cleavage site. Introns can also be associated with promoters to improve or alter expression. As an example, a promoter / intron combination that can be used in the present application includes, but is not limited to, the maize Ubil promoter and intron promoter / intron combination.
[0094] Non-limiting examples of introns that can be used in the present application include introns from ADHI genes (e.g., Adh1-S introns 1, 2, and 6), ubiquitin genes (Ubil), RuBisCO small subunit (rbcS) genes, RuBisCO large subunit (rbcL) genes, actin genes (e.g., actin-1 intron), pyruvate dehydrogenase kinase genes (pdk), nitrate reductase genes (nr), repetitive carbonic anhydrase gene 1 (Tdcal), psbA genes, atpA genes, or any combination thereof.
[0095] As used herein, an“editing system” refers to any site-specific (e.g., sequence-specific) nucleic acid editing system now known or hereafter developed that can introduce modifications (e.g., mutations) in a nucleic acid in a target-specific manner. For example, an editing system (e.g., a site-specific and / or sequence-specific editing system) can include, but is not limited to, a CRISPR-Cas editing system, a meganuclease editing system, a zinc finger nuclease (ZFN) editing system, a transcription activator-like effector nuclease (TALEN) editing system, a base editing system, and / or a prime editing system, each of which can comprise one or more polypeptides and / or one or more polynucleotides that can modify (e.g., mutate) a target nucleic acid in a sequence-specific manner when present and / or expressed together in a composition and / or cell (e.g., as a system). In some embodiments, an editing system (e.g., a site-specific and / or sequence-specific editing system) can comprise one or more polynucleotides and / or one or more polypeptides, including but not limited to a nucleic acid binding polypeptide (e.g., a DNA binding domain), a nuclease, another polypeptide, and / or a polynucleotide. In some embodiments, a CRISPR-Cas editing system comprising an engineered protein of the present application is provided and / or used.
[0096] In some embodiments, an editing system comprises one or more sequence-specific nucleic acid binding polypeptides (e.g., DNA binding domains), which can be from, for example, a polynucleotide-guided endonuclease, a CRISPR-Cas endonuclease (e.g., a CRISPR-Cas effector protein), a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein. In some embodiments, an editing system comprises one or more cleaving polypeptides (e.g., nucleases), including but not limited to an endonuclease (e.g., Fok1), a polynucleotide-guided endonuclease, a CRISPR-Cas endonuclease (e.g., a CRISPR-Cas effector protein), a zinc finger nuclease, and / or a transcription activator-like effector nuclease (TALEN).
[0097] As used herein, a "nucleic acid binding polypeptide" refers to a polypeptide or domain that binds and / or is capable of binding a nucleic acid (e.g., a target nucleic acid). A DNA binding domain is an exemplary nucleic acid binding polypeptide and can be a site-specific and / or sequence-specific nucleic acid binding domain. In some embodiments, a nucleic acid binding polypeptide can be a sequence-specific nucleic acid binding polypeptide, such as, but not limited to, a sequence-specific binding domain from, e.g., a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein (e.g., a CRISPR-Cas endonuclease), a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein. In some embodiments, a nucleic acid binding polypeptide comprises a cleavage domain (e.g., a nuclease domain), such as, but not limited to, an endonuclease (e.g., Fokl), a polynucleotide-guided endonuclease, a CRISPR-Cas endonuclease, a zinc finger nuclease, and / or a transcription activator-like effector nuclease (TALEN). In some embodiments, a nucleic acid binding polypeptide is associated with and / or is capable of associating with (e.g., forming a complex with) one or more nucleic acid molecules (e.g., forming a complex with a guide nucleic acid as described herein), which can direct and / or guide the nucleic acid binding polypeptide to a particular target nucleotide sequence (e.g., a locus of a genome) that is complementary to the one or more nucleic acid molecules (or a portion or region thereof), such that the nucleic acid binding polypeptide binds to the nucleotide sequence at the particular target site. In some embodiments, a nucleic acid binding polypeptide is a CRISPR-Cas effector protein as described herein.
[0098] In some embodiments, an editing system comprises or is a ribonucleoprotein, such as a pre-assembled ribonucleoprotein complex (e.g., a ribonucleoprotein comprising a CRISPR-Cas effector protein, a guide nucleic acid, and optionally a deaminase). In some embodiments, the ribonucleoprotein of an editing system can assemble together (e.g., a pre-assembled ribonucleoprotein comprising a CRISPR-Cas effector protein, a guide nucleic acid, and optionally a deaminase), such as when contacted with a target nucleic acid or when introduced into a cell (e.g., a mammalian cell or a plant cell). In some embodiments, the ribonucleoprotein of an editing system can assemble into a complex (e.g., a covalently and / or non-covalently bound complex) while a portion of the ribonucleoprotein contacts a target nucleic acid and / or can assemble after and / or during introduction into a plant cell. In some embodiments, an editing system can assemble (e.g., assemble into a covalently and / or non-covalently bound complex) upon introduction into a plant cell. In some embodiments, a ribonucleoprotein can comprise an engineered protein of the disclosure, a guide nucleic acid, and optionally a deaminase.
[0099] As used herein, the term "transgene" or "transgenic" refers to at least one nucleic acid sequence that is taken or synthetically produced from the genome of one organism and then introduced into a host cell (e.g., a plant cell) or organism or tissue of interest and subsequently integrated into the genome of the host by a "stable" transformation or transfection method. In contrast, the term "transient" transformation or transfection or introduction refers to a manner of introducing a molecular tool that includes at least one nucleic acid (DNA, RNA, single or double stranded or mixtures thereof) and / or at least one amino acid sequence, optionally comprising a suitable chemical or biological agent, to effect transfer into at least one compartment of interest of the cell, including but not limited to the cytoplasm, organelles (including nucleus, mitochondria, vacuole, chloroplast), or membrane, resulting in transcription and / or translation and / or association and / or activity of the introduced at least one molecule without effecting stable integration or incorporation into the genome, and thus the corresponding at least one molecule introduced into the genome of the cell is not inherited. The term "non-transgenic" refers to the absence or non-discovery of a transgene in the genome of a host cell or tissue or organism of interest.
[0100] In some embodiments, the polynucleotides and / or nucleic acid constructs of the present application can be, or can be comprised within, an "expression cassette." As used herein, an "expression cassette" means a recombinant nucleic acid molecule that comprises, for example, a nucleic acid construct of the present application (e.g., a polynucleotide encoding an engineered protein of the present application, a polynucleotide encoding a nuclease, a polynucleotide encoding a cytosine deaminase, a polynucleotide encoding an adenine deaminase, a polynucleotide encoding a deaminase fusion protein, a polynucleotide encoding a peptide tag, a polynucleotide encoding an affinity polypeptide, a polynucleotide encoding a glycosylase, and / or a polynucleotide comprising a guide nucleic acid), wherein the nucleic acid construct is operably associated with at least a control sequence (e.g., a promoter). Thus, some embodiments of the present application provide expression cassettes designed to express, for example, a nucleic acid construct of the present application. When an expression cassette comprises more than one polynucleotide, the polynucleotides can be operably linked to a single promoter that drives expression of all of the polynucleotides, or the polynucleotides can be operably linked to one or more separate promoters (e.g., three polynucleotides can be driven by one, two, or three promoters in any combination). Thus, for example, a polynucleotide encoding an engineered protein, a polynucleotide encoding a deaminase (e.g., an adenine deaminase), and a polynucleotide comprising a guide nucleic acid included in an expression cassette can each be operably associated with a single promoter, or one or more of the polynucleotides can be operably associated with separate promoters (e.g., two or three promoters) in any combination, which can be the same as or different from each other.
[0101] In some embodiments, the expression cassettes comprising the polynucleotides / nucleic acid constructs of the application can be optimized for expression in an organism (e.g., an animal, a plant, a bacterium, etc.).
[0102] The expression cassettes comprising the nucleic acid constructs of the application can be chimeric, meaning that at least one component is heterologous with respect to at least one other component thereof (e.g., a promoter from a host organism operably linked to a polynucleotide of interest to be expressed in the host organism, where the polynucleotide of interest is from a different organism than the host, or is not normally found in association with the promoter). The expression cassettes can also be naturally occurring, but have been obtained in a recombinant form that is useful for heterologous expression.
[0103] The expression cassettes can optionally include a transcriptional and / or translational termination region (i.e., termination region) and / or an enhancer region that is functional in the selected host cell. A variety of transcriptional terminators and enhancers are known in the art and can be used in the expression cassettes. The transcriptional terminator is responsible for termination of transcription and correction of mRNA polyadenylation. The termination region and / or enhancer region can be native to the transcriptional initiation region, can be native to the gene encoding the CRISPR-Cas effector protein or the gene encoding the deaminase, can be native to the gene encoding the polypeptide of the application, can be native to the host cell, or can be native to another source (e.g., foreign or heterologous to the promoter, the gene encoding the CRISPR-Cas effector protein or the gene encoding the deaminase, the host cell, or any combination thereof).
[0104] The expression cassettes of the application can also include a polynucleotide encoding a selectable marker, which can be used to select for transformed host cells. As used herein, a “selectable marker” means a polynucleotide sequence that, when expressed, imparts a unique phenotype to the host cell expressing the marker, and thereby allows such transformed cells to be distinguished from cells that do not have the marker. Such polynucleotide sequences can encode selectable or screenable markers, depending on whether the marker confers a trait that can be selected for by chemical means (such as by the use of a selective agent (e.g., an antibiotic, etc.)), or whether the marker is simply a trait that can be identified by observation or testing (such as by screening (e.g., fluorescence)). Numerous examples of suitable selectable markers are known in the art and can be used in the expression cassettes described herein.
[0105] The expression cassettes, nucleic acid molecules / constructs, and polynucleotide sequences described herein can be used in conjunction with a vector. The term "vector" refers to a composition used to transfer, deliver, or introduce a nucleic acid (or nucleic acids) into a cell. A vector can comprise a nucleic acid construct comprising one or more nucleotide sequences to be transferred, delivered, or introduced into a cell. Vectors for transforming host organisms are well known in the art. Non-limiting examples of general classes of vectors include viral vectors (e.g., adeno-associated virus (AAV) vectors), plasmid vectors, bacteriophage vectors, phagemid vectors, fosmid vectors, phages, artificial chromosomes, minicircles, or Agrobacterium binary vectors in double- or single-stranded linear or circular form, which can or can not be self-transmissible or mobilizable. In some embodiments, viral vectors can include, but are not limited to, retroviral, lentiviral, adenoviral, adeno-associated viral, or herpes simplex viral vectors. Vectors as defined herein can transform prokaryotic or eukaryotic hosts by integration into the cellular genome or exist extrachromosomally (e.g., autonomously replicating plasmids with an origin of replication). Additionally, also included are shuttle vectors, which mean DNA vectors that are capable of replicating naturally or intentionally in two different host organisms, which can be selected from actinomycetes and related species, bacteria, and eukaryotes (e.g., higher plants, mammalian, yeast, or fungal cells). In some embodiments, the nucleic acid in the vector is under the control of, and operably linked to, an appropriate promoter or other regulatory element for transcription in the host cell. The vector can be a bifunctional expression vector that functions in a variety of hosts. In the case of genomic DNA, this can contain its own promoter and / or other regulatory elements, while in the case of cDNA, this can be under the control of an appropriate promoter and / or other regulatory elements for expression in the host cell. Thus, the nucleic acid constructs of the present application and / or expression cassettes comprising the same can be comprised in vectors described herein and known in the art.
[0106] As used herein, "contact," "contacting," "contacted," and grammatical variants thereof, refer to bringing components of a desired reaction together under conditions suitable for the desired reaction to occur (e.g., conversion, transcriptional control, genome editing, nicking, and / or cleavage). Thus, for example, a target nucleic acid can be contacted with a nucleic acid construct of the application encoding, for example, a nucleic acid binding polypeptide (e.g., a DNA binding domain such as a sequence-specific DNA binding protein (e.g., a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein (e.g., a CRISPR-Cas endonuclease), a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein)), a guide nucleic acid, a polynucleotide encoding a polypeptide of the application, and, optionally, a cytosine deaminase and / or an adenine deaminase, under conditions in which the nucleic acid binding polypeptide (e.g., a CRISPR-Cas effector protein) is expressed and the nucleic acid binding polypeptide forms a complex with the guide nucleic acid, which hybridizes to the target nucleic acid, and, optionally, a polypeptide of the application, a cytosine deaminase and / or an adenine deaminase is recruited to the nucleic acid binding polypeptide (and thereby to the target nucleic acid) or is fused to the nucleic acid binding polypeptide, thereby modifying the target nucleic acid. In some embodiments, the polypeptide of the application, a cytosine deaminase and / or an adenine deaminase, and the nucleic acid binding polypeptide are positioned at the target nucleic acid, optionally by covalent and / or non-covalent interactions.
[0107] In some embodiments, a target nucleic acid can be contacted with a nucleic acid construct of the application encoding an engineered protein of the application, a guide nucleic acid, and, optionally, a cytosine deaminase and / or an adenine deaminase, under conditions in which the engineered protein is expressed, or a target nucleic acid can be contacted with an engineered protein of the application, a guide nucleic acid, and, optionally, a cytosine deaminase and / or an adenine deaminase. The engineered protein can form a complex with the guide nucleic acid, and the complex can hybridize to the target nucleic acid, and, optionally, a cytosine deaminase and / or an adenine deaminase is recruited to the engineered protein (and thereby to the target nucleic acid) or is fused to the engineered protein, thereby modifying the target nucleic acid. The cytosine deaminase and / or adenine deaminase and the engineered protein can be positioned at the target nucleic acid, optionally by covalent and / or non-covalent interactions.
[0108] As used herein, “modifying” or “modification” of a target nucleic acid includes editing (e.g., mutation), covalent modification, exchange / substitution of nucleic acid / nucleotide bases, deletion, cleavage, and / or nicking of the target nucleic acid to provide a modified nucleic acid and / or alteration of transcriptional control of the target nucleic acid to provide a modified nucleic acid. In some embodiments, modification may include insertions and / or deletions of any size and / or any type of single-base alteration (SNP). In some embodiments, modification includes an SNP. In some embodiments, modification includes the exchange and / or substitution of one or more (e.g., 1, 2, 3, 4, 5, or more) nucleotides. In some embodiments, the length of the insertion or deletion can be from about 1 base to about 30,000 bases or more (e.g., lengths of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38). 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 8 6, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360 370, 380, 390, 400, 410, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, 1100, 1200, 1300 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000, 11 500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, 20,000, 20,500, 21,000, 2 1,500, 22,000, 22,500, 23,000, 23,500, 24,000, 24,500, 25,000, 25,500, 26,000, 26,500, 27,000, 27,500, 28,000, 28,500, 29,000, 29,500, 30,000 bases or more, or any value or range thereof. Therefore, in some embodiments, the length of the insertion or deletion can be approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 4... 3, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 8889, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300 to about 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000 bases or any range or value therein; about 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300 bases to about 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650,660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 bases or more or any value or range therein; about 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 bases to about 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, or 10,000 bases or more or any value or range therein; or about 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, or 700 bases to about 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000 bases or more or any value or range therein.1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2500, 3000, 3500, 4000, 4500, or 5000 bases or more or any value or range therein. In some embodiments, the length of the insertion or deletion can be about 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, or 10,000 bases to about 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, 20,000, 20,500, 21,000, 21,500, 22,000, 22,500, 23,000, 23,500, 24,000, 24,500, 25,000, 25,500, 26,000, 26,500, 27,000, 27,500, 28,000, 28,500, 29,000, 29,500, or 30,000 bases or more or any value or range therein.
[0109] As used herein, “recruit,” “recruiting,” or “recruitment” refers to the use of protein-protein interactions, nucleic acid protein interactions (e.g., RNA-protein interactions), and / or chemical interactions to draw one or more polypeptides or polynucleotides to another polypeptide or polynucleotide (e.g., a particular location in the genome). Protein-protein interactions can include, but are not limited to, peptide tags (epitopes, multimerization epitopes) and corresponding affinity polypeptides, RNA recruiting motifs and corresponding affinity polypeptides, and / or chemical interactions. Example chemical interactions that can be used with polypeptides and polynucleotides for recruitment purposes can include, but are not limited to, rapamycin-induced FRB-FKBP dimerization; biotin-streptavidin interactions; SNAP tags (Hussain et al. Curr Pharm Des. 19(30):5437-42 (2013)); Halo tags (Los et al. ACS Chem Biol. 3(6):373-82 (2008)); CLIP tags (Gautier et al. Chemistry & Biology 15:128-136 (2008)); DmrA-DmrC heterodimer induced by a compound (Tak et al. Nat Methods 14(12):1163-1166 (2017)); bifunctional ligand approaches (fusing two protein-binding chemicals together) (Voβ et al. Curr Opin Chemical Biology 28:194-201 (2015)) (e.g., dihydrofolate reductase (DHFR) (Kopyteck et al. Cell Chem Biol 7(5):313-321 (2000)).
[0110] In the context of a polynucleotide or editing system of interest, "introducing," "introduce," "introduced" (and grammatical variations thereof) means presenting to a host organism or a cell of the organism (e.g., a host cell; e.g., a plant cell) the nucleotide sequence (e.g., a polynucleotide, a nucleic acid construct, and / or a guide nucleic acid) and / or editing system (e.g., a polynucleotide, a polypeptide, and / or a ribonucleoprotein) of interest in a manner such that the nucleotide sequence and / or editing system gains access to the interior of a cell. Thus, for example, a nucleic acid construct of the application encoding an engineered protein, a guide nucleic acid, and / or a cytosine deaminase and / or an adenine deaminase of the application can be introduced into a cell of an organism, thereby using the engineered protein, the guide nucleic acid, and the cytosine deaminase and / or adenine deaminase to transform the cell. In some embodiments, an engineered protein and / or a guide nucleic acid can be introduced into a cell of an organism, optionally wherein the engineered protein and guide nucleic acid can be comprised in a complex (e.g., a ribonucleoprotein). In some embodiments, the organism is a eukaryote (e.g., a mammal, such as a human).
[0111] As used herein, the term "transforming" refers to introducing a nucleic acid, a polypeptide, and / or a ribonucleoprotein (e.g., a heterologous nucleic acid, a polypeptide, and / or a ribonucleoprotein) into a cell. The transformation of a cell can be stable or transient. Thus, in some embodiments, a host cell or host organism can be stably transformed with a polynucleotide / nucleic acid molecule of the application. In some embodiments, a host cell or host organism can be transiently transformed with a nucleic acid construct, a polypeptide, and / or a ribonucleoprotein of the application.
[0112] In the context of a polynucleotide, a polypeptide, and / or a ribonucleoprotein, "transiently transforming" means that the polynucleotide, polypeptide, and / or ribonucleoprotein is introduced into a cell and does not integrate into the genome of the cell.
[0113] In the context of a polynucleotide being introduced into a cell, "stably introduced" or "stably introduced" means that the introduced polynucleotide is stably incorporated into the genome of the cell, and thus the cell is stably transformed with the polynucleotide.
[0114] As used herein, "stably transforming" or "stably transformed" means that a nucleic acid molecule is introduced into a cell and integrated into the genome of the cell. Thus, the integrated nucleic acid molecule can be inherited by its progeny, more particularly, by successive generations of progeny. As used herein, "genome" includes nuclear genome and plastid genome, and thus includes integration of a nucleic acid into, for example, a chloroplast or mitochondrial genome. As used herein, stably transforming can also refer to a transgene that is maintained extrachromosomally, for example, as a minichromosome or a plasmid.
[0115] Transient transformation can be detected, for example, by enzyme-linked immunoadsorbent assay (ELISA) or Western blot, which can detect the presence of a peptide or polypeptide encoded by one or more transgenes introduced into an organism. Stable transformation of a cell can be detected, for example, by Southern blot hybridization of genomic DNA of the cell to a nucleic acid sequence that specifically hybridizes to the nucleotide sequence of a transgene introduced into an organism (e.g., a mammal, a plant, etc.). Stable transformation of a cell can also be detected, for example, by Northern blot hybridization of RNA of the cell to a nucleic acid sequence that specifically hybridizes to the nucleotide sequence of a transgene introduced into a host organism. Stable transformation of a cell can also be detected, for example, by polymerase chain reaction (PCR) or other amplification reactions well known in the art, which employ specific primer sequences that hybridize to a target sequence of a transgene, resulting in amplification of the transgene sequence, which can then be detected according to standard methods. Transformation can also be detected by direct sequencing and / or hybridization protocols well known in the art.
[0116] Accordingly, in some embodiments, the nucleotide sequences, polynucleotides, nucleic acid constructs, and / or expression cassettes of the present application can be transiently expressed and / or they can be stably incorporated into the genome of a host organism. Accordingly, in some embodiments, the nucleic acid constructs of the present application can be transiently introduced into a cell along with a guide nucleic acid, and thus, no DNA is maintained in the cell.
[0117] The nucleic acid constructs, polypeptides, and / or ribonucleoproteins of the present application can be introduced into a cell by any method known to one of skill in the art. In some embodiments, the method of transformation includes, but is not limited to, transformation by bacterial-mediated nucleic acid delivery (e.g., by Agrobacteria), viral-mediated nucleic acid delivery, silicon carbide and / or nucleic acid whisker-mediated nucleic acid delivery, liposome-mediated nucleic acid delivery, microinjection, microprojectile bombardment, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation, sonication, maceration, PEG-mediated nucleic acid uptake, and any other electrical, chemical, physical (mechanical), and / or biological mechanism that results in the introduction of nucleic acid into a cell (e.g., a plant cell or an animal cell), including any combination thereof. In some embodiments of the present application, transformation of a cell comprises nuclear transformation. In some embodiments, transformation of a cell comprises plastid transformation (e.g., chloroplast transformation). In some embodiments, the recombinant nucleic acid constructs of the present application can be introduced into a cell by conventional breeding techniques.
[0118] Procedures for transforming eukaryotic and prokaryotic organisms are well known and routine in the art and are described throughout the literature (see, e.g., Jiang et al., 2013. Nat. Biotechnol. 31 :233-239; Ran et al. Nature Protocols 8:2281-2308 (2013)). General guidelines for various plant transformation methods known in the art include Miki et al. ("Procedures for Introducing Foreign DNA into Plants" in Methods in Plant Molecular Biology and Biotechnology, Glick, B. R. and Thompson, J. E., eds. (CRC Press, Inc., Boca Raton, 1993), pp. 67-88) and Rakowoczy-Trojanowska (Cell. Mol. Biol. Lett. 7:849-858 (2002)).
[0119] Accordingly, nucleotide sequences, polypeptides, and / or ribonucleoproteins can be introduced into a host organism or cell thereof in a variety of ways well known in the art. The methods of the present application do not depend on a particular method for introducing one or more nucleotide sequences, polypeptides, and / or ribonucleoproteins into an organism, but only on their entry into the interior of at least one cell of the organism. Where more than one nucleotide sequence, polypeptide, and / or ribonucleoprotein is to be introduced, these can be assembled as part of a single nucleic acid construct, or as separate nucleic acid constructs, and can be located on the same or different nucleic acid constructs. Thus, nucleotide sequences, polypeptides, and / or ribonucleoproteins can be introduced into a cell of interest in a single transformation event and / or in separate transformation events, or, in relevant cases, nucleotide sequences can be incorporated into a plant, e.g., as part of a breeding program. In some embodiments, the cell is a eukaryotic cell (e.g., a plant cell or a mammalian cell such as a human).
[0120] In some embodiments, a nucleic acid construct of the present application (e.g., a polynucleotide encoding an engineered protein of the present application, a polynucleotide encoding a deaminase, and / or a guide nucleic acid, and / or an expression cassette and / or a vector comprising the same) can be operably linked to at least one regulatory sequence, optionally wherein the at least one regulatory sequence can be codon optimized for expression in a plant. In some embodiments, the at least one regulatory sequence can be, for example, a promoter, an operator, a terminator, or an enhancer. In some embodiments, the at least one regulatory sequence can be a promoter. In some embodiments, the regulatory sequence can be an intron. In some embodiments, the at least one regulatory sequence can be, for example, a promoter operably associated with an intron or a promoter comprising an intron. In some embodiments, the at least one regulatory sequence can be, for example, a ubiquitin promoter and its associated intron (e.g., Medicago truncatula and / or corn and its associated intron). In some embodiments, the at least one regulatory sequence can be a terminator nucleotide sequence and / or an enhancer nucleotide sequence.
[0121] In some embodiments, a nucleic acid construct of the present application can be operably associated with a promoter region, wherein the promoter region comprises an intron, optionally wherein the promoter region can be a ubiquitin promoter and intron (e.g., a Medicago or corn ubiquitin promoter and intron, such as SEQ ID NO: 48 or SEQ ID NO: 49). In some embodiments, a nucleic acid construct of the present application operably associated with a promoter region comprising an intron can be codon optimized for expression in a plant.
[0122] In some embodiments, a nucleic acid construct of the present application can encode one or more (e.g., 1, 2, 3, 4, or more) polypeptides of interest. The one or more polypeptides of interest can be codon optimized for expression in a eukaryote (e.g., a human or a plant). In some embodiments, an engineered protein can comprise one or more (e.g., 1, 2, 3, 4, or more) polypeptides of interest. For example, a heterologous polypeptide of an engineered protein can comprise or be a polypeptide of interest.
[0123] Polypeptides of interest that can be used in the present application can include, but are not limited to, polypeptides or protein domains having deaminase activity, nickase activity, recombinase activity, transposase activity, methylase activity, glycosylase (DNA glycosylase) activity, glycosylase inhibitor activity (e.g., uracil-DNA glycosylase inhibitor (UGI)), reverse transcriptase, peptide tag (e.g., GCN4 peptide tag), demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, restriction endonuclease activity (e.g., Fokl), nucleic acid binding activity, methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, polymerase activity, ligase activity, helicase activity, nuclear localization sequence or activity, affinity polypeptide, peptide tag, and / or photolyase activity. In some embodiments, the polypeptide of interest is a Fokl nuclease or a uracil-DNA glycosylase inhibitor. When encoded in a nucleic acid (polynucleotide, expression cassette, and / or vector), the encoded polypeptide or protein domain can be codon optimized for expression in an organism. In some embodiments, the polypeptide of interest can be linked to an engineered protein of the present application or to a CRISPR-Cas effector protein domain to provide a CRISPR-Cas fusion protein. In some embodiments, a CRISPR-Cas fusion protein comprising a CRISPR-Cas effector protein domain linked to a peptide tag can also be linked to a polypeptide of interest (e.g., a CRISPR-Cas effector protein domain can be linked, for example, to both a peptide tag (or affinity polypeptide) and, for example, a polypeptide of interest).
[0124] In some embodiments, the editing system of the application comprises a CRISPR-Cas effector protein. As used herein, a “CRISPR-Cas effector protein” is a protein or polypeptide that cleaves, nick, or nicks a nucleic acid; binds a nucleic acid (e.g., a target nucleic acid and / or a guide nucleic acid); and / or identifies, recognizes, or binds a guide nucleic acid as defined herein. In some embodiments, a CRISPR-Cas effector protein can be and / or can function as an enzyme (e.g., a nuclease, an endonuclease, a nickase, etc.). In some embodiments, a CRISPR-Cas effector protein refers to a CRISPR-Cas nuclease. In some embodiments, a CRISPR-Cas effector protein comprises nuclease activity and / or nickase activity, comprises a nuclease domain whose nuclease activity and / or nickase activity has been reduced or eliminated, comprises single-stranded DNA cleavage activity (ss DNAse activity) or has had its ss DNAse activity reduced or eliminated, and / or comprises self-processing RNAse activity or has had its self-processing RNAse activity reduced or eliminated. A CRISPR-Cas effector protein can bind to a target nucleic acid. A CRISPR-Cas effector protein can be a Type I, Type II, Type III, Type IV, Type V, or Type VI CRISPR-Cas effector protein. In some embodiments, a CRISPR-Cas effector protein can be from a Type I CRISPR-Cas system, a Type II CRISPR-Cas system, a Type III CRISPR-Cas system, a Type IV CRISPR-Cas system, a Type V CRISPR-Cas system, or a Type VI CRISPR-Cas system. In some embodiments, a CRISPR-Cas effector protein of the application can be from a Type II CRISPR-Cas system or a Type V CRISPR-Cas system. In some embodiments, a CRISPR-Cas effector protein can be a Type II CRISPR-Cas effector protein, e.g., a Cas9 effector protein. In some embodiments, a CRISPR-Cas effector protein can be a Type V CRISPR-Cas effector protein, e.g., a Cas12 effector protein. In some embodiments, a CRISPR-Cas effector protein can be a Cas12a, and optionally can have the amino acid sequence of any one of SEQ ID NOs: 50-66 or 180 and / or the nucleotide sequence of any one of SEQ ID NOs: 67-69. In some embodiments, a CRISPR-Cas effector protein can be an active Cas12a, and optionally can have the amino acid sequence of SEQ ID NO: 58 or 180. In some embodiments, a CRISPR-Cas effector protein can be an inactive (i.e., dead) Cas12a, and optionally can have the amino acid sequence of SEQ ID NO: 50.In some embodiments, the CRISPR-Cas effector protein can be Cas12b, and optionally can have the amino acid sequence of SEQ ID NO: 151.
[0125] Exemplary CRISPR-Cas effector proteins include, but are not limited to, Cas9, C2cl, C2c3, Casl2a (also known as Cpf1), Casl2b, Casl2c, Casl2d, Casl2e, Casl3a, Casl3b, Casl3c, Casl3d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csx 12), CaslO, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx 10, Csx 16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4 (dinG), and / or Csf5 nuclease, optionally wherein the CRISPR-Cas effector protein can be a Cas9, Casl2a (Cpf1), Casl2b, Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2g, Casl2h, Casl2i, C2c4, C2c5, C2c8, C2c9, C2c 10, Casl4a, Casl4b, and / or Casl4c effector protein.
[0126] In some embodiments, a CRISPR-Cas effector protein useful in the present application can comprise a mutation in its nuclease active site and / or nuclease domain (e.g., RuvC, HNH, e.g., RuvC site of a Casl2a nuclease domain; e.g., RuvC site and / or HNH site of a Cas9 nuclease domain). A CRISPR-Cas effector protein having a mutation in its nuclease active site and / or nuclease domain and thus no longer comprising nuclease activity is often referred to as “inactivated” or “dead,” e.g., dCas9. In some embodiments, a CRISPR-Cas effector protein having a mutation in its nuclease active site and / or nuclease domain can have impaired activity or reduced activity (e.g., nickase activity) compared to the same CRISPR-Cas effector protein without the mutation.
[0127] A CRISPR Cas9 effector protein or Cas9 that can be used in the application can be any known or later identified Cas9 nuclease. In some embodiments, a Cas9 of the application can be a protein from, for example, Streptococcus spp. (e.g., S. pyogenes, S. thermophilus), Lactobacillus spp., Bifidobacterium spp., Kandleria spp., Leuconostoc spp., Oenococcus spp., Pediococcus spp., Weissella spp., and / or Olsenella spp. In some embodiments, a CRISPR-Cas effector protein can be a Cas9, and optionally can have the nucleotide sequence of any one of SEQ ID NOs: 70-80 or 140-143 and / or the amino acid sequence of any one of SEQ ID NOs: 81-82.
[0128] In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Streptococcus pyogenes and / or can recognize the PAM sequence motif NGG, NAG, NGA (Mali et al., Science 2013; 339(6121): 823-826). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Streptococcus thermophiles and / or can recognize the PAM sequence motif NGGNG and / or NNAGAAW (W = A or T) (see, e.g., Horvath et al., Science, 2010; 327(5962): 167-170, and Deveau et al., J Bacteriol 2008; 190(4): 1390-1400). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Streptococcus mutans and / or can recognize the PAM sequence motif NGG and / or NAAR (R = A or G) (see, e.g., Deveau et al., J Bacteriol 2008; 190(4): 1390-1400). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Streptococcus aureus and / or can recognize the PAM sequence motif NNGRR (R = A or G). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from S. aureus and / or can recognize the PAM sequence motif N GRRT (R = A or G). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from S. aureus and / or can recognize the PAM sequence motif N GRRV (R = A or G). In some embodiments, the CRISPR-Cas effector protein can be Cas9 derived from Neisseria meningitidis and / or can recognize the PAM sequence motif N GATT or NGCTT (R = A or G, V = A, G or C) (see, e.g., Hou et al., PNAS 2013, 1-6). In the foregoing embodiments in this paragraph, the N in the PAM sequence motif can be any nucleotide residue, e.g., any of A, G, C or T. In some embodiments, the CRISPR-Cas effector protein can be Casl3a derived from Leptotrichia shahii and / or can recognize a protospacer flanking sequence (PFS) (or RNAPAM (rPAM)) sequence motif of a single 3' A, U or C, which can be located within the target nucleic acid.
[0129] A Type V CRISPR-Cas effector protein that can be used in embodiments of the present application can be any Type V CRISPR-Cas nuclease. Exemplary Type V CRISPR-Cas effector proteins include, but are not limited to, Casl2a (Cpfl), Casl2b, Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2g, Casl2h, Casl2i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Casl4a, Casl4b, and / or Casl4c nucleases. In some embodiments, the Type V CRISPR-Cas effector protein can be Casl2a. In some embodiments, the Type V CRISPR-Cas effector protein can be a nickase, optionally a Casl2a nickase. In some embodiments, the Type V CRISPR-Cas effector protein can be Casl2b (e.g., SEQ ID NO: 151).
[0130] In some embodiments, the CRISPR-Cas effector protein can be a Type V Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-Cas nuclease. Casl2a differs from the more well-known Type II CRISPR Cas9 nuclease in several respects. For example, Cas9 recognizes a G-rich protospacer adjacent motif (PAM) (3'-NGG) located 3' of its guide RNA (gRNA, sgRNA, crRNA, crDNA, CRISPR array) binding site (protospacer, target nucleic acid, target DNA), whereas Casl2a recognizes a T-rich PAM (5'-TTN, 5'-TTTN) located 5' of the target nucleic acid. In fact, the orientation of Cas9 and Casl2a binding to its guide RNA is nearly reversed with respect to its N and C termini. Further, Casl2a enzymes use a single guide RNA (gRNA, CRISPR array, crRNA), rather than the dual guide RNA (sgRNA (e.g., crRNA and tracrRNA)) found in the native Cas9 system, and Casl2a processes its own gRNA. Additionally, Casl2a nuclease activity produces staggered DNA double-strand breaks, rather than the blunt ends produced by Cas9 nuclease activity, and Casl2a relies on a single RuvC domain to cleave both DNA strands, whereas Cas9 utilizes an HNH domain and a RuvC domain to cleave.
[0131] A CRISPR Cas12a effector protein useful in the application can be any known or later identified Cas12a (formerly known as Cpf1) (see, e.g., U.S. Patent No. 9,790,490, the disclosure of which regarding Cpf1 (Cas12a) sequences is incorporated by reference). The term “Cas12a” refers to an RNA-guided protein that can have nuclease activity, the protein comprising a guide nucleic acid binding domain and an active, inactive, or partially active DNA cleavage domain, whereby the RNA-guided nuclease activity of Cas12a can be active, inactive, or partially active, respectively. In some embodiments, a Cas12a useful in the application can comprise a mutation in the nuclease active site (e.g., the RuvC site of the Cas12a domain). A Cas12a having a mutation in its nuclease domain and / or nuclease active site and thus no longer comprising nuclease activity is often referred to as a dead Cas12a (e.g., dCas12a). In some embodiments, a Cas12a having a mutation in its nuclease domain and / or nuclease active site can have impaired activity, e.g., can have reduced nickase activity.
[0132] In some embodiments, a CRISPR-Cas effector protein can be optimized for expression in an organism, e.g., in an animal (e.g., a mammal, such as a human), a plant, a fungus, an archaea, or a bacterium. In some embodiments, a CRISPR-Cas effector protein (e.g., a Cas12a polypeptide / domain or a Cas9 polypeptide / domain) can be optimized for expression in a plant.
[0133] Any deaminase domain / polypeptide useful for base editing can be used in the application. As used herein, “cytosine deaminase” and “cytidine deaminase” refer to a polypeptide or domain thereof that catalyzes or is capable of catalyzing the deamination of cytosine, as the polypeptide or domain catalyzes or is capable of catalyzing the removal of an amine group from a cytosine base. Thus, a cytosine deaminase can result in the conversion of cytosine to thymidine (via a uracil intermediate), resulting in a C to T conversion or a G to A conversion in the complementary strand in the genome. Thus, in some embodiments, a cytosine deaminase encoded by a polynucleotide of the application results in a C to T conversion in the sense (e.g., “+”; template) strand of a target nucleic acid or a G to A conversion in the anti-sense (e.g., “-”, complementary) strand of a target nucleic acid. In some embodiments, a cytosine deaminase encoded by a polynucleotide of the application results in a C to T, G, or A conversion in the complementary strand in the genome.
[0134] A cytosine deaminase useful in the present application can be any known or later identified cytosine deaminase from any organism (see, e.g., U.S. Patent No. 10,167,457 and Thuronyi et al., Nature Biotechnology 37: 1070-1079 (2019), each of which is incorporated herein by reference for its disclosure of cytosine deaminases). A cytosine deaminase can catalyze the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. Thus, in some embodiments, a deaminase or deaminase domain useful in the present application can be a cytidine deaminase domain that catalyzes the hydrolytic deamination of cytosine to uracil. In some embodiments, a cytosine deaminase can be a variant of a naturally occurring cytosine deaminase, including but not limited to primates (e.g., humans, monkeys, chimpanzees, gorillas), dogs, cows, rats, or mice. Thus, in some embodiments, a cytosine deaminase useful in the present application can be about 70% to about 100% identical to a wild-type cytosine deaminase (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to a naturally occurring cytosine deaminase, as well as any range or value therein).
[0135] In some embodiments, the cytosine deaminase used in this invention may be an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the cytosine deaminase may be APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, APOBEC3H deaminase, APOBEC4 deaminase, human activation-inducible deaminase (hAID), rAPOBEC1, FERNY and / or CDA1, optionally pmCDA1, atCDA1 (e.g., At2g19570), and its evolved versions. Evolved deaminases are disclosed, for example, in U.S. Patent No. 10,113,163, Gaudelli et al., Nature 551(7681):464-471 (2017), and Thuronyi et al., Nature Biotechnology 37:1070-1079 (2019), the disclosure of each of these documents regarding its deaminase and evolved deaminases is incorporated herein by reference. In some embodiments, the cytosine deaminase may be the APOBEC1 deaminase having the amino acid sequence of SEQ ID NO:83. In some embodiments, the cytosine deaminase may be the APOBEC3A deaminase having the amino acid sequence of SEQ ID NO:84. In some embodiments, the cytosine deaminase may be the CDA1 deaminase, optionally the CDA1 having the nucleotide sequence of SEQ ID NO:85. In some embodiments, the cytosine deaminase may be the FERNY deaminase, optionally the FERNY having the amino acid sequence of SEQ ID NO:86. In some embodiments, the cytosine deaminase may be rAPOBEC1 deaminase, optionally having the amino acid sequence of SEQ ID NO:87. In some embodiments, the cytosine deaminase may be hAID deaminase, optionally having the amino acid sequence of SEQ ID NO:88 or SEQ ID NO:89. In some embodiments, the cytosine deaminases that can be used in the present invention may be about 70% to about 100% identical in amino acid sequence to naturally occurring cytosine deaminases (e.g., "evolved deaminases") (see, for example, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92) (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%).In some embodiments, cytosine deaminases useful in the application can be about 70% to about 99.5% identical (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical) to the amino acid sequence of any of SEQ ID NOs: 83, 84, or 86-92, or to the nucleotide sequence of SEQ ID NO: 85 (e.g., at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of any of SEQ ID NOs: 83, 84, or 86-92, or to the nucleotide sequence of SEQ ID NO: 85). In some embodiments, a polynucleotide encoding a cytosine deaminase can be codon optimized for expression in a plant, and the codon optimized polypeptide can be about 70% to 99.5% identical to the reference polynucleotide.
[0136] As used herein, “adenine deaminase” and “adenosine deaminase” refer to a polypeptide or domain thereof that catalyzes or is capable of catalyzing the hydrolytic deamination of adenine or adenosine (e.g., removal of the amine group from adenine). In some embodiments, an adenine deaminase can catalyze the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, an adenosine deaminase can catalyze the hydrolytic deamination of adenine or adenosine in DNA. In some embodiments, an adenine deaminase encoded by a nucleic acid construct of the application can produce an A→G conversion in the sense (e.g., “+”; template) strand of a target nucleic acid or a T→C conversion in the anti-sense (e.g., “-”, complementary) strand of a target nucleic acid. Adenine deaminases useful in the application can be any known or later identified adenine deaminase from any organism (see, e.g., U.S. Patent No. 10,113,163, which is incorporated by reference herein for its disclosure of adenine deaminases).
[0137] In some embodiments, the adenosine deaminase can be a variant of a naturally occurring adenine deaminase. Thus, in some embodiments, the adenosine deaminase can be about 70% to 100% identical to a wild-type adenine deaminase (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to a naturally occurring adenine deaminase, and any range or value therein). In some embodiments, the deaminase or deaminase is not naturally occurring and can be referred to as an engineered, mutated, or evolved adenosine deaminase. Thus, for example, an engineered, mutated, or evolved adenine deaminase polypeptide or adenine deaminase domain can be about 70% to 99.9% identical to a naturally occurring adenine deaminase polypeptide / domain (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical to a naturally occurring adenine deaminase polypeptide or adenine deaminase domain, and any range or value therein). In some embodiments, the adenosine deaminase can be from a bacterium (e.g., E. coli, Staphylococcus aureus, Haemophilus influenzae, Caulobacter crescentus, etc.). In some embodiments, the polynucleotide encoding the adenine deaminase polypeptide / domain can be codon optimized for expression in a plant.
[0138] In some embodiments, the adenine deaminase domain can be a wild-type tRNA- specific adenosine deaminase domain, e.g., a tRNA-specific adenosine deaminase (TadA) and / or a mutated / evolved adenosine deaminase domain, e.g., a mutated / evolved tRNA-specific adenosine deaminase domain (TadA*). In some embodiments, the TadA domain can be from E. coli. In some embodiments, the TadA can be modified, e.g., truncated, such that one or more N-terminal and / or C-terminal amino acids are deleted relative to the full-length TadA (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal and / or C-terminal amino acid residues can be deleted relative to the full-length TadA. In some embodiments, the TadA polypeptide or TadA domain does not comprise an N-terminal methionine. In some embodiments, the wild-type E. coli TadA comprises the amino acid sequence of SEQ ID NO: 93. In some embodiments, the mutated / evolved E. coli TadA* comprises the amino acid sequence of any one of SEQ ID NOs: 94-97. In some embodiments, the polynucleotide encoding the TadA / TadA* can be codon-optimized for expression in plants. In some embodiments, the adenine deaminase can comprise all or a portion of the amino acid sequence of any one of SEQ ID NOs: 98-103. In some embodiments, the adenine deaminase can comprise all or a portion of the amino acid sequence of any one of SEQ ID NOs: 93-103.
[0139] In some embodiments, the nucleic acid construct of the application can further encode a glycosylase inhibitor (e.g., a uracil glycosylase inhibitor (UGI), such as a uracil-DNA glycosylase inhibitor). In some embodiments, the application provides a fusion protein comprising an engineered protein and a UGI and / or one or more polynucleotides encoding the same, optionally wherein the one or more polynucleotides can be codon-optimized for expression in plants.
[0140] A“uracil glycosylase inhibitor” useful in the present application can be any protein or polypeptide capable of inhibiting the uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain comprises a wild-type UGI or a fragment thereof. In some embodiments, a UGI domain useful in the present application can be about 70% to about 100% identical (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% identical, and any range or value therein) to the amino acid sequence of a naturally occurring UGI domain. In some embodiments, a UGI domain can comprise the amino acid sequence of SEQ ID NO: 104 or a polypeptide that is about 70% to about 99.5% identical (e.g., at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical) to the amino acid sequence of SEQ ID NO: 104. For example, in some embodiments, a UGI domain can comprise a fragment of the amino acid sequence of SEQ ID NO: 104 that is 100% identical to a portion of contiguous nucleotides of the amino acid sequence of SEQ ID NO: 104 (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 contiguous nucleotides; e.g., about 10, 15, 20, 25, 30, 35, 40, 45 to about 50, 55, 60, 65, 70, 75, 80 contiguous nucleotides). In some embodiments, a UGI domain can be a variant of a known UGI (e.g., SEQ ID NO: 104) that is about 70% to about 99.5% identical (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% identical, and any range or value therein) to the known UGI. In some embodiments, a polynucleotide encoding a UGI can be codon optimized for expression in a plant (e.g., a plant), and the codon optimized polypeptide can be about 70% to about 99.5% identical to the reference polynucleotide.
[0141] The engineered proteins of the present application can be used in combination with a guide nucleic acid (e.g., a guide RNA (gRNA), a CRISPR array, a CRISPR RNA, a crRNA) designed to work with the engineered protein to modify a target nucleic acid. A guide nucleic acid useful in the present application can comprise at least one spacer sequence and at least one repeat sequence. The guide nucleic acid is capable of forming a complex with the engineered protein (e.g., with the nuclease domain of the engineered protein), and the spacer sequence is capable of hybridizing to a target nucleic acid, thereby directing the complex to the target nucleic acid, wherein the target nucleic acid can be modified (e.g., cleaved or edited) and / or regulated (e.g., transcription is regulated) by a deaminase (e.g., a cytosine deaminase and / or an adenine deaminase optionally present in the complex and / or recruited to the complex).
[0142] In some embodiments, an engineered protein comprising a Cas9 domain (or a nucleic acid construct encoding the same) can be used in combination with a Cas9 guide nucleic acid to modify a target nucleic acid, and a deaminase (e.g., cytosine and / or adenine) can be linked to or form a complex with the engineered protein. The cytosine deaminase deaminates a cytosine base in the target nucleic acid, thereby editing the target nucleic acid. The adenine deaminase deaminates an adenine base in the target nucleic acid, thereby editing the target nucleic acid.
[0143] Likewise, an engineered protein can comprise a Cas12a domain (or other selected CRISPR-Cas nuclease, e.g., C2cl, C2c3, Casl2b, Casl2c, Casl2d, Casl2e, Casl3a, Casl3b, Casl3c, Casl3d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), CaslO, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4 (dinG), and / or Csf5) that can form a complex with or be linked to a cytosine deaminase domain and / or an adenine deaminase domain and can be used in combination with a Casl2a guide nucleic acid (or a guide nucleic acid of other selected CRISPR-Cas nuclease) to modify a target nucleic acid, where the cytosine deaminase domain or the adenine deaminase domain of the fusion protein deaminates a cytosine base or an adenine base, respectively, in the target nucleic acid, thereby editing the target nucleic acid.
[0144] As used herein, a “guide nucleic acid,” “guide RNA,” “gRNA,” “CRISPR RNA / DNA,” “crRNA,” or “crDNA” means a nucleic acid comprising at least one spacer sequence complementary to (and hybridizes with) a target nucleic acid (e.g., a target DNA and / or a protospacer) and at least one repeat sequence (e.g., a repeat sequence of a Type V Cas12a CRISPR-Cas system, or a fragment or portion thereof; a repeat sequence of a Type II Cas9 CRISPR-Cas system, or a fragment thereof; a repeat sequence of a Type V C2cl CRISPR Cas system, or a fragment thereof; e.g., C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12i, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csx12), Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4 (dinG), and / or Csf5 CRISPR-Cas system, or a fragment thereof), where the repeat sequence can be linked to the 5' end and / or 3' end of the spacer sequence. In some embodiments, a guide nucleic acid comprises DNA. In some embodiments, a guide nucleic acid comprises RNA (e.g., is a guide RNA). The design of gRNAs of the present application can be based on a Type I, Type II, Type III, Type IV, Type V, or Type VI CRISPR-Cas system.
[0145] In some embodiments, a Cas12a gRNA can comprise, from 5' to 3', a repeat sequence (full length or a portion thereof (“handle”); e.g., a pseudoknot-like structure) and a spacer sequence.
[0146] In some embodiments, the guide nucleic acid can comprise more than one repeat-spacer sequence (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more repeat-spacer sequences) (e.g., repeat-spacer-repeat, e.g., repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer, etc.). The guide nucleic acids of the present invention are synthetic, man-made, and do not exist in nature. The gRNA can be long and can be used as an aptamer (as in the MS2 recruitment strategy) or other RNA structure that dangles a spacer.
[0147] As used herein, “repeat sequence” refers to any repeat sequence of, for example, a wild-type CRISPR Cas locus (e.g., a Cas9 locus, a Cas12a locus, a C2c1 locus, etc.), or a repeat sequence of a synthetic crRNA that functions with a CRISPR-Cas effector protein encoded by a nucleic acid construct of the present invention. Repeat sequences useful in the present invention can be any known or later identified repeat sequence of a CRISPR-Cas locus (e.g., Type I, Type II, Type III, Type IV, Type V, or Type VI), or it can be a synthetic repeat sequence designed to function in a Type I, Type II, Type III, Type IV, Type V, or Type VI CRISPR-Cas system. The repeat sequence can comprise a hairpin structure and / or a stem loop structure. In some embodiments, the repeat sequence can form a pseudoknot-like structure at its 5’ end (i.e., a “handle”). Thus, in some embodiments, the repeat sequence can be identical or substantially identical to a repeat sequence from a wild-type Type I CRISPR-Cas locus, a Type II CRISPR-Cas locus, a Type III CRISPR-Cas locus, a Type IV CRISPR-Cas locus, a Type V CRISPR-Cas locus, and / or a Type VI CRISPR-Cas locus. Repeat sequences from wild-type CRISPR-Cas loci can be determined by established algorithms, such as using CRISPRfinder provided by CRISPRdb (see Grissa et al. Nucleic Acids Res. 35 (Web Server issue): W52-7). In some embodiments, the repeat sequence, or a portion thereof, is linked at its 3’ end to the 5’ end of a spacer sequence, thereby forming a repeat-spacer sequence (e.g., a guide nucleic acid, a guide RNA / DNA, a crRNA, a crDNA).
[0148] In some embodiments, the repeat sequence comprises, consists essentially of, or consists of at least 10 nucleotides, depending on the particular repeat sequence and whether the guide nucleic acid comprising the repeat sequence is processed or unprocessed (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, to 100 or more nucleotides, or any range or value therein; e.g., about). In some embodiments, the repeat sequence comprises, consists essentially of, or consists of about 10 to about 20, about 10 to about 30, about 10 to about 45, about 10 to about 50, about 15 to about 30, about 15 to about 40, about 15 to about 45, about 15 to about 50, about 20 to about 30, about 20 to about 40, about 20 to about 50, about 30 to about 40, about 40 to about 80, about 50 to about 100 or more nucleotides.
[0149] The repeat sequence connected to the 5’ end of the spacer sequence can comprise a portion of the repeat sequence (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, or more contiguous nucleotides of the wild-type repeat sequence). In some embodiments, the portion of the repeat sequence connected to the 5’ end of the spacer sequence can be about five to about ten contiguous nucleotides in length (e.g., about 5, 6, 7, 8, 9, 10 nucleotides) and have at least 90% sequence identity (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more percent) to the same region (e.g., 5’ end) of the wild-type CRISPR Cas repeat nucleotide sequence. In some embodiments, the portion of the repeat sequence can comprise a pseudoknot-like structure (e.g., a “handle”) at its 5’ end.
[0150] As used herein, a “spacer sequence” is a nucleotide sequence that is complementary to a target nucleic acid (e.g., target DNA) (e.g., a protospacer). A spacer sequence can be fully complementary or substantially complementary (e.g., at least about 70% complementary (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher percent)) to the target nucleic acid. Thus, in some embodiments, a spacer sequence can have one, two, three, four, or five mismatches as compared to the target nucleic acid, which can be contiguous or non-contiguous. In some embodiments, a spacer sequence can have 70% complementarity to a target nucleic acid. In other embodiments, a spacer nucleotide sequence can have 80% complementarity to a target nucleic acid. In still other embodiments, a spacer nucleotide sequence can have 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% complementarity to a target nucleic acid (protospacer), etc. In some embodiments, a spacer sequence is 100% complementary to a target nucleic acid. A spacer sequence can be from about 15 nucleotides to about 30 target nucleotides in length (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides, or any range or value therein). Thus, in some embodiments, a spacer sequence can have full or substantial complementarity over a region of a target nucleic acid (e.g., protospacer) that is at least about 15 nucleotides to about 30 nucleotides in length. In some embodiments, a spacer is about 20 nucleotides in length. In some embodiments, a spacer is about 21, 22, or 23 nucleotides in length.
[0151] In some embodiments, the 5' region of the spacer sequence of the guide nucleic acid can be fully complementary to the target nucleic acid, while the 3' region of the spacer can be substantially complementary to the target nucleic acid (as for spacers of Type V CRISPR-Cas systems), or the 3' region of the spacer sequence of the guide nucleic acid can be fully complementary to the target nucleic acid, while the 5' region of the spacer can be substantially complementary to the target nucleic acid (as for spacers of Type II CRISPR-Cas systems), and thus, the overall complementarity of the spacer sequence to the target nucleic acid can be less than 100%. Thus, for example, in a guide nucleic acid of a Type V CRISPR-Cas system, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides in the 5' region of the spacer sequence (i.e., the seed region), e.g., 20 nucleotides, can be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary (e.g., at least about 70% complementary) to the target nucleic acid. In some embodiments, the first 1 to 8 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8 nucleotides and any range therein) at the 5' end of the spacer sequence can be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary (e.g., at least about 50% complementary (e.g., 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher percentage)) to the target nucleic acid.
[0152] As additional examples, in a guide nucleic acid of a Type II CRISPR-Cas system, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides in the 3' region of a 20-nucleotide spacer sequence (i.e., the seed region) can be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary (e.g., at least about 70% complementary) to the target nucleic acid. In some embodiments, the first 1 to 10 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides, and any range therein) of the 3' end of a spacer sequence can be 100% complementary to the target nucleic acid, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary (e.g., at least about 50% complementary (e.g., at least about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more percent, or any range or value therein)) to the target nucleic acid. The recruiting guide RNA further comprises one or more recruiting motifs as described herein, which can be linked to the 5' end or 3' end of the guide, or which can be inserted into the recruiting guide nucleic acid (e.g., within a hairpin loop).
[0153] In some embodiments, the seed region of the spacer can be about 8 to about 10 nucleotides in length, about 5 to about 6 nucleotides in length, or about 6 nucleotides in length.
[0154] “Target nucleic acid,” “target DNA,” “target nucleotide sequence,” “target region,” and “target region in a genome” are used interchangeably herein and refer to a region in a genome of an organism (e.g., a plant) that comprises a sequence that is perfectly complementary (100% complementary) or substantially complementary (e.g., at least 70% complementary (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more percent)) to a spacer sequence in a guide nucleic acid as defined herein. A target nucleic acid is targeted by an editing system (or components thereof) as described herein. A target region that can be used in a CRISPR-Cas system can be located immediately adjacent to a 3’ (e.g., Type V CRISPR-Cas system) or immediately adjacent to a 5’ (e.g., Type II CRISPR-Cas system) of a PAM sequence in a genome of an organism (e.g., a plant genome or a mammalian (e.g., human) genome). A target region can be selected from any region of at least 15 contiguous nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides, etc.) located immediately adjacent to a PAM sequence.
[0155] As used herein, a “protospacer sequence” or “protospacer” refers to a sequence that is perfectly or substantially complementary to (and can hybridize with) a spacer sequence of a guide nucleic acid. In some embodiments, a protospacer is all or a portion of a target nucleic acid as defined herein that is perfectly or substantially complementary to (and hybridizes with) a spacer sequence of a CRISPR repeat sequence-spacer sequence (e.g., a guide nucleic acid, a CRISPR array, a crRNA).
[0156] In the case of Type V CRISPR-Cas (e.g., Cas12a) systems and Type II CRISPR-Cas (Cas9) systems, a protospacer sequence is flanked (e.g., immediately adjacent to) a protospacer adjacent motif (PAM). For Type IV CRISPR-Cas systems, the PAM is located at the 5’ end of the non-target strand and at the 3’ end of the target strand (see below as an example).
[0157]
[0158] In the case of Type II CRISPR-Cas (e.g., Cas9) systems, the PAM is located immediately 3’ of the target region. The PAM for Type I CRISPR-Cas systems is located 5’ of the target strand. No PAM is known for Type III CRISPR-Cas systems. Makarova et al. describe the nomenclature for all classes, types, and subtypes of CRISPR systems (Nature Reviews Microbiology 13:722-736 (2015)). R. Barrangou describes guide structures and PAMs (Genome Biol. 16:247 (2015)).
[0159] Typical Cas12a PAMs are rich in T. In some embodiments, a typical Cas12a PAM sequence can be 5’-TTN, 5’-TTTN, or 5’-TTTV. In some embodiments, a typical Cas9 (e.g., S. pyogenes) PAM can be 5’-NGG-3’. In some embodiments, atypical PAMs can be used, but efficiency can be lower.
[0160] One of skill in the art can determine additional PAM sequences through established experimental and computational methods. Thus, for example, experimental methods include targeting sequences flanked by all possible nucleotide sequences and identifying members of the sequence that do not undergo targeting, such as by transformation of target plasmid DNA (Esvelt et al. 2013. Nature Methods 10:1116-1121; Jiang et al. 2013. Nature Biotechnology 31:233-239). In some aspects, computational methods can include BLAST searching of natural spacers to identify original target DNA sequences in phage or plasmids, and aligning these sequences to determine conserved sequences adjacent to the target sequence (Briner and Barrangou 2014. Appl. Environ. Microbiol. 80:994-1001; Mojica et al. 2009. Microbiology 155:733-740).
[0161] In some embodiments, the present application provides expression cassettes and / or vectors comprising a nucleic acid construct of the present application (e.g., one or more components of an editing system of the present application). In some embodiments, expression cassettes and / or vectors comprising a nucleic acid construct of the present application and / or one or more guide nucleic acids can be provided. In some embodiments, a nucleic acid construct of the present application encodes an engineered protein and / or a deaminase, and each can be comprised on the same or separate expression cassettes or vectors as an expression cassette or vector comprising one or more guide nucleic acids. When a nucleic acid construct encoding an engineered protein or component of an editing system is comprised on a separate expression cassette or vector from an expression cassette or vector comprising a guide nucleic acid, a target nucleic acid and an expression cassette or vector encoding an engineered protein or component of an editing system can be contacted (e.g., provided together) with each other and a guide nucleic acid in any order, e.g., before, simultaneously with, or after providing an expression cassette comprising a guide nucleic acid (e.g., in contact with a target nucleic acid).
[0162] Methods of recruiting one or more components of an editing system to each other and / or a target nucleic acid are known in the art and can include the use of a peptide tag or an affinity polypeptide that interacts with a peptide tag. In some embodiments, a guide nucleic acid can be linked to an RNA recruitment motif and a deaminase can be linked to an affinity polypeptide that is capable of interacting with the RNA recruitment motif, thereby recruiting the deaminase to the target nucleic acid. Alternatively, a polypeptide (e.g., a deaminase) can be recruited to a target nucleic acid using a chemical interaction.
[0163] As used herein, a “recruitment motif” refers to one half of a binding pair that can be used to recruit a compound to which the recruitment motif binds to another compound comprising the other half of the binding pair (i.e., a “counter motif”). A recruitment motif and a counter motif can bind covalently and / or non-covalently. In some embodiments, a recruitment motif is an RNA recruitment motif (e.g., an RNA recruitment motif that is capable of and / or configured to bind to an affinity polypeptide), an affinity polypeptide (e.g., an affinity polypeptide that is capable of and / or configured to bind to an RNA recruitment motif and / or a peptide tag), or a peptide tag (e.g., a peptide tag that is capable of and / or configured to bind to an affinity polypeptide). For example, when a recruitment motif is an RNA recruitment motif, a counter motif for the RNA recruitment motif can be an affinity polypeptide that binds to the RNA recruitment motif. Further examples are when a recruitment motif is a peptide tag, a counter motif for the peptide tag can be an affinity polypeptide that binds to the peptide tag. Thus, a compound comprising a recruitment motif (e.g., an affinity polypeptide) can be recruited to another compound (e.g., a guide nucleic acid) comprising a counter motif (e.g., an RNA recruitment motif) for the recruitment motif.
[0164] Peptide tags (e.g., epitopes) that can be used in the present application can include, but are not limited to, GCN4 peptide tags (e.g., Sun-Tag), c-Myc affinity tags, HA affinity tags, His affinity tags, S affinity tags, methionine-His affinity tags, RGD-His affinity tags, FLAG octapeptides, strep tags or strep tags II, V5 tags, and / or VSV-G epitopes. Any epitope that can be linked to a polypeptide and for which a corresponding affinity polypeptide can be linked to another polypeptide can be used as a peptide tag in the present application. In some embodiments, a peptide tag can comprise 1 or 2 or more copies of a peptide tag (e.g., repeating units, multimerized epitopes (e.g., tandem repeat sequences)) (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or more repeating units). In some embodiments, the affinity polypeptide that interacts / binds with the peptide tag can be an antibody. In some embodiments, the antibody can be an scFv antibody. In some embodiments, the affinity polypeptide that binds with the peptide tag can be synthetic (e.g., evolved for affinity interactions), including, but not limited to, affibodies, anticalins, monobodies, and / or DARPins (see, e.g., Sha et al., Protein Sci. 26(5):910-924 (2017); Gilbreth (Curr Opin Struc Biol 22(4):413-420 (2013)); U.S. Patent No. 9,982,053), each of which is incorporated by reference in its entirety for its teachings regarding affibodies, anticalins, monobodies, and / or DARPins.
[0165] In some embodiments, a guide nucleic acid can be linked to an RNA recruiting motif and a polypeptide (e.g., a deaminase) to be recruited can be fused to an affinity polypeptide that binds to the RNA recruiting motif, wherein the guide binds to a target nucleic acid and the RNA recruiting motif binds to the affinity polypeptide, thereby recruiting the polypeptide to the guide and bringing the target nucleic acid into contact with the polypeptide (e.g., deaminase). In some embodiments, two or more polypeptides can be recruited to a guide nucleic acid, thereby bringing a target nucleic acid into contact with two or more polypeptides (e.g., deaminases).
[0166] In some embodiments of the application, the guide RNA can be linked to one or two or more RNA recruiting motifs (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more motifs; e.g., at least 10 to about 25 motifs), optionally wherein the two or more RNA recruiting motifs can be the same RNA recruiting motif or different RNA recruiting motifs. In some embodiments, the RNA recruiting motif and the corresponding affinity polypeptide can include, but are not limited to, a telomerase Ku binding motif (e.g., a Ku binding hairpin) and the corresponding affinity polypeptide Ku (e.g., a Ku heterodimer), a telomerase Sm7 binding motif and the corresponding affinity polypeptide Sm7, an MS2 bacteriophage operon stem loop and the corresponding affinity polypeptide MS2 coat protein (MCP), a PP7 bacteriophage operon stem loop and the corresponding affinity polypeptide PP7 coat protein (PCP), an SfMu bacteriophage Com stem loop and the corresponding affinity polypeptide Com RNA binding protein, a PUF binding site (PBS) and the affinity polypeptide Pumilio / fem-3 mRNA binding factor (PUF), and / or a synthetic RNA aptamer and aptamer ligand as the corresponding affinity polypeptide. In some embodiments, the RNA recruiting motif and the corresponding affinity polypeptide can be an MS2 bacteriophage operon stem loop and the affinity polypeptide MS2 coat protein (MCP). In some embodiments, the RNA recruiting motif and the corresponding affinity polypeptide can be a PUF binding site (PBS) and the affinity polypeptide Pumilio / fem-3 mRNA binding factor (PUF). Exemplary RNA recruiting motifs and corresponding affinity polypeptides useful in the application can include, but are not limited to, SEQ ID NOs: 108-118.
[0167] In some embodiments, the components for recruiting polypeptides and nucleic acids can include components that act through chemical interactions, which can include, but are not limited to, rapamycin-induced FRB-FKBP dimerization; biotin-streptavidin; SNAP tag; Halo tag; CLIP tag; compound-induced DmrA-DmrC heterodimer; bifunctional ligand (e.g., two protein binding chemicals fused together; e.g., dihydrofolate reductase (DHFR)).
[0168] In some embodiments, a nucleic acid construct, expression cassette, or vector of the application optimized for expression in plants can be about 70% to 100% identical (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%) to a nucleic acid construct, expression cassette, or vector comprising the same polynucleotide but which has not been codon-optimized for expression in plants.
[0169] As described herein, a “peptide tag” can be used to recruit one or more polypeptides. A peptide tag can be any polypeptide that is capable of being bound by a corresponding motif, such as an affinity polypeptide. A peptide tag can also be referred to as an “epitope,” and when provided in multiple copies, as a “multimerized epitope.” Example peptide tags can include, but are not limited to, a GCN4 peptide tag (e.g., Sun-Tag), a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or strep tag II, a V5 tag, and / or a VSV-G epitope. In some embodiments, a peptide tag can also include a phosphorylated tyrosine in a specific sequence context recognized by SH2 domains, a characteristic consensus sequence containing a phosphoserine recognized by 14-3-3 proteins, a proline-rich peptide motif recognized by SH3 domains, PDZ protein interaction domains, or PDZ signal sequences, and an AGO hook motif from plants. Peptide tags are disclosed in WO 2018 / 136783 and U.S. Patent Application Publication No. 2017 / 0219596, which are incorporated by reference for their disclosure of peptide tags. Peptide tags useful in the application can include, but are not limited to, SEQ ID NO: 119 and SEQ ID NO: 120. Affinity polypeptides useful with peptide tags include, but are not limited to, SEQ ID NO: 121.
[0170] A peptide tag can comprise or be present in one copy or 2 or more copies of a peptide tag (e.g., a multimerized peptide tag or a multimerized epitope) (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 9, 20, 21, 22, 23, 24, or 25 or more peptide tags). When multimerized, the peptide tags can be directly fused to one another, or they can be linked to one another by one or more amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20), optionally about 3 to about 10, about 4 to about 10, about 5 to about 10, about 5 to about 15, or about 5 to about 20 amino acids, etc., and any value or range therein. Thus, in some embodiments, a CRISPR-Cas effector protein of the application can comprise a CRISPR-Cas effector protein domain fused to one peptide tag or two or more peptide tags, optionally wherein the two or more peptide tags are fused to one another by one or more amino acid residues. In some embodiments, a peptide tag useful in the application can be a single copy of a GCN4 peptide tag or epitope, or can be a multimerized GCN4 epitope comprising about 2 to about 25 or more copies of a peptide tag (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more copies of a GCN4 epitope, or any range therein).
[0171] In some embodiments, a peptide tag can be fused to a CRISPR-Cas polypeptide or domain. In some embodiments, a peptide tag can be fused or linked to the C-terminus of a CRISPR-Cas effector protein to form a CRISPR-Cas fusion protein. In some embodiments, a peptide tag can be fused or linked to the N-terminus of a CRISPR-Cas effector protein to form a CRISPR-Cas fusion protein. In some embodiments, a peptide tag can be fused within a CRISPR-Cas effector protein (e.g., a peptide tag can be in a loop region of a CRISPR-Cas effector protein). In some embodiments, a peptide tag can be fused to a cytosine deaminase and / or adenine deaminase.
[0172] An“affinity polypeptide” (e.g.,“recruiting polypeptide”) refers to any polypeptide that is capable of binding to its corresponding peptide tag, peptide tag, or RNA recruiting motif. The affinity polypeptide of a peptide tag can be, for example, an antibody and / or a single-chain antibody that specifically binds to the peptide tag, respectively. In some embodiments, the antibody of a peptide tag can be, but is not limited to, an scFv antibody. In some embodiments, the affinity polypeptide can be fused or linked to the N-terminus of a deaminase (e.g., a cytosine deaminase or an adenine deaminase). In some embodiments, the affinity polypeptide is stable under reducing conditions of a cell or cell extract.
[0173] The nucleic acid construct and / or guide nucleic acid of the present application can be comprised in one or more expression cassettes as described herein. In some embodiments, the nucleic acid construct of the present application can be comprised in the same or separate expression cassettes or vectors as the expression cassettes or vectors comprising the guide nucleic acid and / or recruiting guide nucleic acid.
[0174] In some embodiments, the nucleic acid construct, expression cassette, or vector of the present application that is optimized for expression in an organism (e.g., a human or a plant) can be about 70% to 100% identical (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%) to a nucleic acid construct, expression cassette, or vector comprising the same polynucleotide but that has not been codon-optimized for expression in the organism.
[0175] When used in combination with a guide nucleic acid, the nucleic acid construct of the present application (and expression cassettes and / or vectors comprising the same) can be used to modify a target nucleic acid and / or alter its expression. The target nucleic acid can be contacted with the nucleic acid construct of the present application and / or expression cassettes and / or vectors comprising the same before, simultaneously with, or after contacting the target nucleic acid with the guide nucleic acid / recruiting guide nucleic acid and / or expression cassettes and vectors comprising the same.
[0176] According to embodiments of the present application, provided herein are engineered proteins. As used herein, an “engineered protein” refers to a polypeptide or protein that does not naturally occur in nature. In some embodiments, an engineered protein refers to a polypeptide comprising a first polypeptide from a first protein and a second polypeptide from a second protein, wherein the first protein and the second protein are different from each other. In some embodiments, an engineered protein of the present application comprises a polypeptide from a CRISPR-Cas effector protein (i.e., a CRISPR-Cas effector polypeptide) and a polypeptide that is heterologous to the CRISPR-Cas effector polypeptide (i.e., a heterologous polypeptide). A polypeptide from a CRISPR-Cas effector protein is referred to herein as a “CRISPR-Cas effector polypeptide,” and a “CRISPR-Cas effector polypeptide” is all or a portion of a CRISPR-Cas effector protein. In some embodiments, a “CRISPR-Cas effector polypeptide” does not include the entire CRISPR-Cas effector protein, and thus has a reduced number of amino acids compared to the number of amino acids of a CRISPR-Cas effector protein. In some embodiments, a CRISPR-Cas effector polypeptide is a Type V CRISPR-Cas effector polypeptide (i.e., all or a portion of a Type V CRISPR-Cas effector protein). In some embodiments, a Type V CRISPR-Cas effector polypeptide is a portion of a Type V CRISPR-Cas effector protein. In some embodiments, an engineered protein of the present application is a target strand nickase that has about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to one or more wild-type CRISPR-Cas effector proteins (e.g., Type V CRISPR-Cas effector proteins) and / or to the amino acid sequence of one or more of SEQ ID NOs: 50-66, 151, and / or 180. In some embodiments, an engineered protein of the present application is a target strand nickase that has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to one or more wild-type CRISPR-Cas effector proteins (e.g., Type V CRISPR-Cas effector proteins) and / or to the amino acid sequence of one or more of SEQ ID NOs: 50-66, 151, and / or 180.
[0177] The engineered proteins of the present application can include non-native PAM recognition sites and / or sequences. For example, the engineered proteins of the present application can be nucleases that comprise non-native PAM recognition specificity (e.g., altered binding affinity) in addition to or in place of the native PAM recognition specificity (e.g., wild-type binding affinity) of Cas12a. As used herein in reference to modified proteins, engineered proteins, and / or nucleases, “altered PAM specificity” means that the PAM specificity of the modified protein, engineered protein, and / or nuclease is different than the PAM specificity of the wild-type nuclease (e.g., recognizes non-native PAM sequences in addition to and / or in place of native PAM sequences). For example, a modified Cas12a nuclease (e.g., modified protein) is altered in its PAM specificity if it recognizes PAM sequences that are different than and / or in addition to the native Cas12a PAM sequence TTTV, where V is A, C, or G. In some embodiments, the engineered proteins of the present application comprise a polypeptide that is part of Cas12a, and the engineered protein has an altered PAM specificity, i.e., it recognizes PAM sequences that are different than and / or in addition to the native Cas12a PAM sequence TTTV, where V is A, C, or G. In some embodiments, the modified proteins of the present application can be as described in U.S. Patent Application Publication No. 2021 / 0115421, and / or the nucleases or portions thereof (e.g., polypeptides thereof) can be modified and / or altered as described in U.S. Patent Application Publication No. 2021 / 0115421, which is incorporated by reference herein in its entirety. In some embodiments, the engineered proteins of the present application recognize native PAM sequences (e.g., the PAM sequence of TTTV, where V is A, C, or G) and / or recognize non-native PAM sequences (e.g., the PAM sequences of CCCC, TCCA, TCCC, TCCG, TTCA, TTCC, TATA, TATC, and TATG, and / or TTCG).
[0178] In some embodiments, the engineered proteins and / or modified proteins of the application can comprise altered protospacer adjacent motif (PAM) specificity as compared to wild-type LbCas12a (e.g., SEQ ID NO: 180), LbCas12a comprising mutations S542R / K607R (e.g., LbCas12a-RR; e.g., SEQ ID NO: 182), and / or LbCas12a comprising mutations S542R / K548V / N552R (e.g., LbCas12a-RVR; e.g., SEQ ID NO: 183). The engineered proteins and / or modified proteins of the application can have altered PAM specificity, wherein the altered PAM specificity includes, but is not limited to, NNNG, NNNT, NNNA, NNNC, NNG, NNT, NN C, NNA, NG, NT, NC, NA, NN, NNN, NNNN, wherein each N in each sequence is independently selected from any of T, C, G, or A. In some embodiments, the altered PAM specificity can include, but is not limited to, TTTA, TTTC, TTTG, TTTT, TTCA, TTCC, TTCG, TATA, TATC, TCCG, TCCC, TCCA, and / or TATG. In some embodiments, the altered PAM specificity can be NNNN, wherein each N in each sequence is independently selected from any of T, C, G, or A.
[0179] In addition to having altered PAM recognition specificity, the engineered proteins and / or modified proteins of the application can further comprise mutations in the nuclease active site (e.g., RuvC domain) (e.g., dead LbCas12a, dLbCas12a). Such modifications can result in the engineered proteins and / or modified proteins having reduced nuclease activity (e.g., nickase activity) or no nuclease activity.
[0180] In some embodiments, the CRISPR-Cas effector polypeptide (e.g., a Type V CRISPR-Cas effector polypeptide) of the engineered protein of the application lacks a nuclease domain (e.g., lacks a RuvC domain). A polypeptide that is heterologous to a CRISPR-Cas effector polypeptide can be referred to herein as a heterologous polypeptide. The heterologous polypeptide can be a polypeptide of interest described herein. In some embodiments, the engineered protein comprises all or a portion of a deaminase domain (e.g., a cytosine deaminase and / or an adenine deaminase), which can be linked to any portion of the engineered protein. For example, in some embodiments, all or a portion of a deaminase domain is linked to the N- or C-terminus of a CRISPR-Cas effector polypeptide and / or the N- or C-terminus of the engineered protein. In some embodiments, all or a portion of a deaminase domain is positioned between two portions of the engineered protein. In some embodiments, the engineered protein comprises all or a portion of a polypeptide of interest, which can be linked to any portion of the engineered protein.
[0181] The engineered protein can cleave, nick, or incise a nucleic acid; bind a nucleic acid (e.g., a target nucleic acid and / or a guide nucleic acid); and / or identify, recognize, or bind a guide nucleic acid as defined herein. In some embodiments, the engineered protein or a portion thereof can be and / or can function as an enzyme (e.g., a nuclease, an endonuclease, a nickase, etc.). In some embodiments, the engineered protein of the application is an RNA-guided DNA binding protein. In some embodiments, the engineered protein is present in and / or forms a complex with a guide nucleic acid, which is a single guide nucleic acid (e.g., a gRNA, a CRISPR array, and / or a crRNA), optionally wherein the guide nucleic acid can be a single crRNA. In some embodiments, the complex comprises the engineered protein and the guide nucleic acid, and the guide nucleic acid and / or the complex consists of a single guide nucleic acid (e.g., a single crRNA). In some embodiments, the engineered protein binds a single guide nucleic acid (e.g., a single crRNA), recognizes and / or binds a target nucleic acid, and has nuclease activity, optionally wherein the engineered protein cleaves a target strand of the target nucleic acid.
[0182] In some embodiments, the engineered protein comprises a first CRISPR-Cas effector polypeptide and a heterologous polypeptide. The first CRISPR-Cas effector polypeptide can be referred to herein as a first polypeptide, and can be a Type V CRISPR-Cas effector polypeptide. The heterologous polypeptide can be referred to herein as a second polypeptide. The first polypeptide can lack a nuclease domain, optionally lacking a RuvC domain. The second polypeptide can be linked to the N-terminus or C-terminus of the first polypeptide, optionally with or without a linker (e.g., a peptide linker). In some embodiments, the second polypeptide is linked to the N-terminus or C-terminus of the first polypeptide by a peptide linker, optionally wherein the peptide linker is a GS linker. In some embodiments, the peptide linker is a GS linker having 1, 2, 3, or 4 amino acid residues, optionally having 2 or 4 amino acid residues. In some embodiments, the peptide linker has one of the amino acid sequences of SEQ ID NOs: 18-47 or 176-179. In some embodiments, the peptide linker can comprise the following amino acid sequences: (GGS) n , GS, SG, GSSG (SEQ ID NO: 175), GSSGSS (SEQ ID NO: 176), GSSGSSGS (SEQ ID NO: 177), (GSS) n (SEQ ID NO: 178), (GSS) n GS (SEQ ID NO: 179), S(GGS) n (SEQ ID NO: 42), SGGS (SEQ ID NO: 43 or (GGGGS)n(SEQ ID NO: 44), wherein n is an integer from 1-20 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20).
[0183] In some embodiments, the first polypeptide is part of a first CRISPR-Cas effector protein (e.g., part of a Type V CRISPR-Cas effector protein, such as part of Casl2a). In some embodiments, the heterologous polypeptide (e.g., the second polypeptide) comprises a nuclease domain, optionally an HNH domain (e.g., the HNH domain comprises a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the amino acid sequence of one or more of SEQ ID NOs: 1 or 169-174). In some embodiments, the second polypeptide comprises an HNH domain from a CRISPR-Cas effector protein and / or comprises a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the amino acid sequence of any of SEQ ID NOs: 1 or 172. In some embodiments, the second polypeptide comprises an HNH domain that is not from a CRISPR-Cas effector protein and / or comprises a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the amino acid sequence of any of SEQ ID NOs: 169-171 or 173-174. In some embodiments, the second polypeptide is a polypeptide in a CRISPR-Cas effector protein, optionally wherein the second polypeptide is from a CRISPR-Cas effector protein (e.g., a Type II CRISPR-Cas effector protein) that is different from the type of the first CRISPR-Cas effector protein (e.g., a Type V CRISPR-Cas effector protein) of which the first CRISPR-Cas effector polypeptide is a part.
[0184] In some embodiments, the engineered protein comprises a first polypeptide, a second polypeptide, and a third polypeptide, which can be linked together in any order. In some embodiments, the first polypeptide can lack a RuvC domain. The second polypeptide can be linked to the N-terminus or C-terminus of the first polypeptide, optionally with or without a linker (e.g., a peptide linker), and / or the second polypeptide can be linked to the N-terminus or C-terminus of the third polypeptide, optionally with or without a linker (e.g., a peptide linker). In some embodiments, the second polypeptide is between the first polypeptide and the third polypeptide. In some embodiments, the first polypeptide is part of a first CRISPR-Cas effector protein (e.g., part of a Type V CRISPR-Cas effector protein, such as part of Cas12a), and the third polypeptide is part of a second CRISPR-Cas effector protein (e.g., part of a Type V CRISPR-Cas effector protein, such as part of Cas12a), where the first CRISPR-Cas effector protein and the second CRISPR-Cas effector protein can be the same protein or different proteins. In some embodiments, the first CRISPR-Cas effector protein and the second CRISPR-Cas effector protein are the same, so the first polypeptide and the third polypeptide are parts from the same protein, but can be different parts of the CRISPR-Cas effector protein. The first polypeptide and the third polypeptide can have different sequences. In some embodiments, the first polypeptide and the third polypeptide can comprise the same sequence. In some embodiments, the first polypeptide and the third polypeptide together provide the complete sequence of the CRISPR-Cas effector protein. In some embodiments, the first polypeptide and the third polypeptide together do not constitute the complete sequence of the CRISPR-Cas effector protein (i.e., part of the sequence of the CRISPR-Cas effector protein is not present in both the sequence of the first polypeptide and the third polypeptide); for example, 1 or 5 to 10, 15, 20, 25, 30, or more amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, or more amino acids) of the CRISPR-Cas effector protein can not be present in the sequence of the first effector polypeptide and the third effector polypeptide. In some embodiments, the second polypeptide comprises a nuclease domain, optionally an HNH domain (e.g., an HNH domain from a Type II CRISPR-Cas effector protein). In some embodiments, the second polypeptide comprises an HNH domain that is not from a CRISPR-Cas effector protein.In some embodiments, the second polypeptide is a polypeptide from a CRISPR-Cas effector protein, optionally wherein the second polypeptide is from a different type of CRISPR-Cas effector protein (e.g., a Type II CRISPR-Cas effector protein) than the type of the first CRISPR-Cas effector protein (e.g., a Type V CRISPR-Cas effector protein) from which the portion of the first polypeptide is derived, and / or different than the type of the second CRISPR-Cas effector protein (e.g., a Type V CRISPR-Cas effector protein) from which the portion of the third polypeptide is derived. In some embodiments, the second polypeptide is from a Type II CRISPR-Cas effector protein (e.g., is a portion of a Type II CRISPR-Cas effector protein (e.g., an HNH domain or a portion thereof)), the first polypeptide is a portion of a Type V CRISPR-Cas effector protein, and the third polypeptide is a portion of a Type V CRISPR-Cas effector protein, wherein the first and second CRISPR-Cas effector polypeptides are different. In some embodiments, the second polypeptide is heterologous to one of the first polypeptide and the third polypeptide. In some embodiments, the second polypeptide is heterologous to both the first polypeptide and the third polypeptide.
[0185] As used herein, a "heterologous polypeptide" refers to a polypeptide that is not naturally occurring as compared to the CRISPR-Cas effector polypeptides present in the engineered protein. Thus, in nature, the heterologous polypeptide of the engineered protein is not present in at least one CRISPR-Cas effector polypeptide of the engineered protein, and thus the heterologous polypeptide is non-naturally occurring with respect to the at least one CRISPR-Cas effector polypeptide. For example, an engineered protein of the application can include a CRISPR-Cas effector polypeptide as part of a CRISPR-Cas effector protein and a heterologous polypeptide (which is optionally from a CRISPR-Cas effector protein), and the heterologous polypeptide is non-naturally occurring as compared to the CRISPR-Cas effector polypeptide without the heterologous polypeptide (e.g., the CRISPR-Cas effector polypeptide does not have or precedes inclusion (e.g., insertion or fusion) of the heterologous polypeptide and the CRISPR-Cas effector polypeptide); in some embodiments, the heterologous polypeptide is heterologous to the CRISPR-Cas effector protein from which the portion of the CRISPR-Cas effector polypeptide is derived. In some embodiments, the engineered protein includes a heterologous polypeptide, a first CRISPR-Cas effector polypeptide as part of a first CRISPR-Cas effector protein, and a second CRISPR-Cas effector polypeptide as part of a second CRISPR-Cas effector protein, and the heterologous polypeptide is non-naturally occurring (i.e., heterologous) in the first CRISPR-Cas effector polypeptide, the first CRISPR-Cas effector protein, the second CRISPR-Cas effector, and the second CRISPR-Cas effector protein. Similarly, a nucleotide sequence encoding the heterologous polypeptide is heterologous (i.e., non-naturally occurring) to a nucleotide sequence encoding a CRISPR-Cas effector polypeptide of the engineered protein.
[0186] In some embodiments, the heterologous polypeptide comprises a polypeptide or domain from a different type of protein than the CRISPR-Cas effector polypeptide (e.g., a Type V CRISPR-Cas effector polypeptide) of the engineered protein. For example, the heterologous polypeptide of an engineered protein of the application can be part of a first CRISPR-Cas effector protein (e.g., a Type II CRISPR-Cas effector polypeptide), and the heterologous polypeptide is heterologous to another CRISPR-Cas effector polypeptide present in the engineered protein (e.g., a Type V CRISPR-Cas effector polypeptide). In some embodiments, the engineered protein comprises one or more (e.g., 1, 2, 3, or more) portions (i.e., one or more CRISPR-Cas effector polypeptides) from a Type V CRISPR-Cas effector protein (e.g., Cas 12a), and one or more (e.g., 1, 2, 3, or more) polypeptides from a different type of CRISPR-Cas effector protein, such as a Type II CRISPR-Cas effector protein. When two or more portions or polypeptides are from the same protein and each is present in the engineered protein, the two or more portions or polypeptides can be separated from each other in the engineered protein by a linker and / or a heterologous polypeptide (i.e., the two or more portions or polypeptides can not be directly connected), or can be in a different order than the order of the protein from which they are derived (e.g., a wild-type protein and / or a CRISPR-Cas effector protein). In some embodiments, the engineered protein comprises one or more (e.g., 1, 2, 3, or more) portions (i.e., one or more CRISPR-Cas effector polypeptides) from a Type V CRISPR-Cas effector protein (e.g., Cas 12a), and at least one polypeptide from and / or is a portion of a Type II CRISPR-Cas effector protein (e.g., Cas 9). In some embodiments, the engineered protein comprises a first CRISPR-Cas effector polypeptide that is a portion of a Type V CRISPR-Cas effector protein (e.g., Cas 12a), and a heterologous polypeptide from a Type II CRISPR-Cas effector protein (e.g., Cas 9). In some embodiments, the engineered protein comprises a first CRISPR-Cas effector polypeptide that is a portion of a Type V CRISPR-Cas effector protein (e.g., Cas 12a), a heterologous polypeptide from a Type II CRISPR-Cas effector protein (e.g., Cas 9), and a second CRISPR-Cas effector polypeptide that is a portion of a Type V CRISPR-Cas effector protein (e.g., Cas 12a), optionally wherein the first and second CRISPR-Cas effector polypeptides are different portions from the same Type V CRISPR-Cas effector protein (e.g., Cas 12a).In some embodiments, the engineered protein comprises a first CRISPR-Cas effector polypeptide that is part of a Type V CRISPR-Cas effector protein (e.g., Cas 12a), a heterologous polypeptide comprising an HNH domain or portion thereof, and a second CRISPR-Cas effector polypeptide that is part of a Type V CRISPR-Cas effector protein (e.g., Cas 12a), optionally wherein the first and second CRISPR-Cas effector polypeptides are different parts from the same Type V CRISPR-Cas effector protein (e.g., Cas 12a).
[0187] In some embodiments, the engineered protein comprises one or more (e.g., 1, 2, 3, or more) domains or portions thereof from a Type V CRISPR-Cas effector protein (e.g., Cas 12a) and one or more (e.g., 1, 2, 3, or more) domains or portions thereof from a different type of CRISPR-Cas effector protein (e.g., a Type II CRISPR-Cas effector protein). In some embodiments, the engineered protein comprises one or more (e.g., 1, 2, 3, or more) domains or portions thereof from a Type V CRISPR-Cas effector protein (e.g., Cas 12a) and at least one domain or portion thereof from a Type II CRISPR-Cas effector protein (e.g., Cas9). In some embodiments, the heterologous polypeptide of the engineered protein does not interfere with or adversely affect the activity of the CRISPR-Cas effector polypeptide and / or one or more domains (e.g., RuvC domain) of the CRISPR-Cas effector polypeptide.
[0188] In some embodiments, the engineered protein of the application comprises a first polypeptide as a first portion of a modified protein, wherein the modified protein comprises an amino acid sequence that is at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the amino acid sequence of SEQ ID NO: 180 (LbCas12a), and the modified protein has a mutation at one or more positions selected from the group consisting of N100, K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, N211, N263, T296, E330, K387, A404, D405, D423, E484, L498, N527, Q529, G532, D535, K538, E539, D541, Y542, Y553, Y554, D572, L585, K591, M592, K595, V596, S599, K600, K601, Y616, Y646, W649, and any combination thereof, with reference to the position numbering of SEQ ID NO: 180; and a second polypeptide that is heterologous to the first polypeptide and is not a Type V CRISPR-Cas effector polypeptide, wherein the first polypeptide is different from the second polypeptide, and wherein the engineered protein comprises the mutation. Any portion of the engineered protein can include one or more (e.g., 1, 2, 3, 4, 5, or more) mutations. In some embodiments, one or more (e.g., 1, 2, 3, 4, 5, or more) portions of the engineered protein are portions of the modified protein, and the one or more portions can be the same or different from each other. In some embodiments, the engineered protein comprises at least two different portions of the modified protein, and the first portion and / or the second portion can comprise one or more mutations of the modified protein described herein. In some embodiments, the first polypeptide comprises a mutation and / or the second polypeptide comprises a mutation. In some embodiments, the Type V CRISPR-Cas effector polypeptide of the engineered protein comprises the one or more mutations.In some embodiments, the first polypeptide comprises a mutation relative to the amino acid sequence of SEQ ID NO: 180, and the mutation is at one or more positions selected from the group consisting of N100, K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, N211, N263, T296, E330, K387, A404, D405, D423, E484, L498, N527, Q529, G532, D535, K538, E539, D541, Y542, Y553, Y554, D572, L585, K591, M592, K595, V596, S599, K600, K601, Y616, Y646, W649, and any combination thereof, in reference to the position numbering of SEQ ID NO: 180. For example, the first polypeptide, optionally when optimally aligned with the amino acid sequence of SEQ ID NO: 180, has at least one mutation at a position selected from the group consisting of N100, K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, N211, N263, T296, E330, K387, A404, D405, D423, E484, L498, N527, Q529, G532, D535, K538, E539, D541, Y542, Y553, Y554, D572, L585, K591, M592, K595, V596, S599, K600, K601, Y616, Y646, W649, and any combination thereof, in reference to the position numbering of SEQ ID NO: 180. In some embodiments, the engineered protein comprises a third polypeptide as part of the modified protein. The third polypeptide, optionally when optimally aligned with the amino acid sequence of SEQ ID NO: 180, can have at least one mutation at a position selected from the group consisting of N100, K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, N211, N263, T296, E330, K387, A404, D405, D423, E484, L498, N527, Q529, G532, D535, K538, E539, D541, Y542, Y553, Y554, D572, L585, K591, M592, K595, V596, S599, K600, K601, Y616, Y646, W649, and any combination thereof, in reference to the position numbering of SEQ ID NO: 180.
[0189] The modified protein, optionally when optimally aligned with the amino acid sequence of SEQ ID NO: 180, can comprise one or more amino acid mutations selected from the group consisting of N100S, K116D, K116R, K116N, K120R, K120H, K120N, K120T, K120Y, K120Q, K121S, K121T, K121H, K121R, K121G, K121D, K121Q, D122R, D122K, D122H, D122E, D122N, E125G, E125R, E125K, E125Q, E125Y, T148H, T148S, T148A, T148C, T149A, T149C, T149S, T149G, T149H, T149P, T149F, T149N, T149D, T149V, T152R, T152K, T152W, T152Y, T152H, T152Q, T152E, T152L, T152F, D156R, D156K, D156Y, D156W, D156Q, D156H, D156I, D156V, D156L, D156E, E159K, E159R, E159H, E159Y, E159Q, N211S, N263I, T296I, E330V, K387E, A404V, D405G, D423V, E484D, L498M, N527S, Q529N, Q529T, Q529H, Q529A, Q529F, Q529G, Q529S, Q529P, Q529W, Q529D, G532D, G532N, G532S, G532H, G532F, G532K, G532R, G532Q, G532A, G532L, G532C, D535N, D535H, D535V, D535T, D535S, D535A, D535W, D535K, K538R, K538V, K538Q, K538W, K538Y, K538F, K538H, K538L, K538M, K538C, K538G, K538A, K538P, E539V, D541N, D541H, D541R, D541K, D541Y, D541I, D541A, D541S, D541E, Y542R, Y542K, Y542H, Y542Q, Y542F, Y542L, Y542M, Y542P, Y542V, Y542N, Y542T, Y553H, Y554N, D572G, L585Q, L585G, L585H, L585F, K591W, K591F, K591Y, K591H, K591R, K591S,K591A, K591G, K591P, M592R, M592K, M592Q, M592E, M592A, K595R, K595Q, K595Y, K595L, K595W, K595H, K595E, K595S, K595D, K595M, V596T, V596H, V596G, V596A, S599G, S599H, S599N, S599D, K600R, K600H, K600G, K601R, K601H, K601Q, K601T, Y616K, Y616R, Y616E, Y616F, Y616H, Y646R, Y646E, Y646K, Y646H, Y646Q, Y646W, Y646N, W649H, W649K, W649Y, W649R, W649E, W649S, W649V, W649T, and any combination thereof. In some embodiments, the modified protein, optionally when optimally aligned with the amino acid sequence of SEQ ID NO: 180, comprises one or more amino acid mutations selected from the group consisting of N100S, K116D, E125G, T152R, D156E, N211S, N263I, T296I, E330V, K387E, A404V, D405G, D423V, E484D, L498M, N527S, G532R, K538V, E539V, Y542R, Y553H, Y554N, D572G, L585Q, K595R, K595Y, and any combination thereof, with reference to the numbering of the positions of SEQ ID NO: 180.
[0190] In some embodiments, the engineered protein of the application comprises an amino acid sequence that is at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the amino acid sequence of SEQ ID NO: 131 (SYN3298); and the engineered protein comprises a mutation relative to the amino acid sequence of SEQ ID NO: 131 (SYN3298) (e.g., in the optimal alignment with SEQ ID NO: 131) at one or more positions selected from the group of N100, K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, N211, N263, T444, E478, K535, A552, D553, D571, E632, L646, N675, Q677, G680, D683, K686, E687, D689, Y690, Y701, Y702, D720, L733, K739, M740, K743, V744, S747, K748, K749, Y764, Y794, W797, and any combination thereof, in reference to the position numbering of SEQ ID NO: 131. For example, the engineered protein of the application, optionally in the optimal alignment with the amino acid sequence of SEQ ID NO: 131, has at least one mutation at the following positions: N100, K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, N211, N263, T444, E478, K535, A552, D553, D571, E632, L646, N675, Q677, G680, D683, K686, E687, D689, Y690, Y701, Y702, D720, L733, K739, M740, K743, V744, S747, K748, K749, Y764, Y794, W797, and any combination thereof, in reference to the position numbering of SEQ ID NO: 131. In some embodiments, the engineered protein comprises a first polypeptide that is a Type V CRISPR-Cas effector polypeptide and lacks a nuclease domain; and a second polypeptide that is heterologous to the first polypeptide and is not a Type V CRISPR-Cas effector polypeptide. The engineered protein can further comprise a third polypeptide that is a Type V CRISPR-Cas effector polypeptide. Any portion of the engineered protein can include one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, 383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483,5 or more) mutations. In some embodiments, one or more (e.g., 1, 2, 3, 4, 5, or more) portions of the engineered protein are portions of the amino acid sequence of SEQ ID NO: 131 (SYN3298), and the one or more portions can be identical to or different from one another. In some embodiments, the engineered protein comprises at least two different portions of the amino acid sequence of SEQ ID NO: 131 (SYN3298), and the first portion and / or the second portion can comprise one or more mutations described herein. In some embodiments, the engineered protein comprises at least three different portions of the amino acid sequence of SEQ ID NO: 131 (SYN3298), and the first portion and / or the second portion can comprise one or more mutations described herein, and the third portion (e.g., a heterologous polypeptide) can not comprise the one or more mutations. In some embodiments, the V-typed CRISPR-Cas effector polypeptide of the engineered protein comprises the one or more mutations. In some embodiments, the engineered protein, optionally when optimally aligned with the amino acid sequence of SEQ ID NO: 131, comprises one or more amino acid mutations selected from the group consisting of: N100S, K116D, K116R, K116N, K120R, K120H, K120N, K120T, K120Y, K120Q, K121S, K121T, K121H, K121R, K121G, K121D, K121Q, D122R, D122K, D122H, D122E, D122N, E125G, E125R, E125K, E125Q, E125Y, T148H, T148S, T148A, T148C, T149A, T149C, T149S, T149G, T149H, T149P, T149F, T149N, T149D, T149V, T152R, T152K, T152W, T152Y, T152H, T152Q, T152E, T152L, T152F, D156R, D156K, D156Y, D156W, D156Q, D156H, D156I, D156V, D156L, D156E, E159K, E159R, E159H, E159Y, E159Q, N211S, N263I, T444I, E478V, K535E, A552V, D553G, D571V, E632D, L646M, N675S, Q677N, Q677T, Q677H, Q677A, Q677F, Q677G, Q677S, Q677P, Q677W, Q677D, G680D, G680N, G680S, G680H, G680F,G680K, G680R, G680Q, G680A, G680L, G680C, D683N, D683H, D683V, D683T, D683S, D683A, D683W, D683K, K686R, K686V, K686Q, K686W, K686Y, K686F, K686H, K686L, K686M, K686C, K686G, K686A, K686P, E687V, D689N, D689H, D689R, D689K, D689Y, D689I, D689A, D689S, D689E, Y690R, Y690K, Y690H, Y690Q, Y690F, Y690L, Y690M, Y690P, Y690V, Y690N, Y690T, Y701H, Y702N, D720G, L733Q, L733G, L733H, L733F, K739W, K739F, K739Y, K739H, K739R, K739S, K739A, K739G, K739P, M740R, M740K, M740Q, M740E, M740A, K743R, K743Q, K743Y, K743L, K743W, K743H, K743E, K743S, K743D, K743M, V744T, V744H, V744G, V744A, S747G, S747H, S747N, S747D, K748R, K748H, K748G, K749R, K749H, K749Q, K749T, Y764K, Y764R, Y764E, Y764F, Y764H, Y794R, Y794E, Y794K, Y794H, Y794Q, Y794W, Y794N, W797H, W797K, W797Y, W797R, W797E, W797S, W797V, W797T, and any combination thereof. In some embodiments, the engineered protein, optionally when optimally aligned with the amino acid sequence of SEQ ID NO: 131, comprises one or more amino acid mutations selected from the group consisting of N100S, K116D, E125G, T152R, D156E, N211S, N263I, T444I, E478V, K535E, A552V, D553G, D571V, E632D, L646M, N675S, G680R, K686V, E687V, Y690R, Y701H, Y702N, D720G, L733Q, K743R, K743Y, and any combination thereof, with reference to the numbering of the positions of SEQ ID NO: 131.
[0191] In some embodiments, the engineered protein of the application comprises an amino acid sequence that is at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the amino acid sequence of SEQ ID NO: 125 (SYN3287), and the engineered protein comprises a mutation at one or more positions, with reference to the position numbering of SEQ ID NO: 125, selected from the group consisting of: D156, G676, K682, Y686, K739, and any combination thereof. In some embodiments, the engineered protein comprises a first polypeptide that is a Type V CRISPR-Cas effector polypeptide and lacks a nuclease domain; and a second polypeptide that is heterologous to the first polypeptide and is not a Type V CRISPR-Cas effector polypeptide. The engineered protein can further comprise a third polypeptide that is a Type V CRISPR-Cas effector polypeptide. Any portion of the engineered protein can comprise one or more (e.g., 1, 2, 3, 4, 5, or more) mutations. In some embodiments, one or more (e.g., 1, 2, 3, 4, 5, or more) portions of the engineered protein are portions of the amino acid sequence of SEQ ID NO: 125 (SYN3287), and the one or more portions can be the same as or different from one another. In some embodiments, the engineered protein comprises at least two different portions of the amino acid sequence of SEQ ID NO: 125 (SYN3287), and the first portion and / or the second portion can comprise one or more mutations described herein. In some embodiments, the engineered protein comprises at least three different portions of the amino acid sequence of SEQ ID NO: 125 (SYN3287), and the first portion and / or the second portion can comprise one or more mutations described herein, and the third portion (e.g., a heterologous polypeptide) can not comprise the one or more mutations. In some embodiments, the Type V CRISPR-Cas effector polypeptide of the engineered protein comprises the one or more mutations. In some embodiments, the engineered protein, optionally when optimally aligned with the amino acid sequence of SEQ ID NO: 125, comprises one or more amino acid mutations selected from the group consisting of: D156R, G676R, K682V, Y686R, K739R, and any combination thereof, with reference to the position numbering of SEQ ID NO: 125.
[0192] In some embodiments, the engineered protein of the application comprises an amino acid sequence that is at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the amino acid sequence of SEQ ID NO: 126 (SYN3288), and the engineered protein comprises a mutation at one or more positions, with reference to the position numbering of SEQ ID NO: 126, selected from the group consisting of: D156, G678, K684, Y688, K741, and any combination thereof. In some embodiments, the engineered protein comprises a first polypeptide that is a Type V CRISPR-Cas effector polypeptide and lacks a nuclease domain; and a second polypeptide that is heterologous to the first polypeptide and is not a Type V CRISPR-Cas effector polypeptide. The engineered protein can further comprise a third polypeptide that is a Type V CRISPR-Cas effector polypeptide. Any portion of the engineered protein can comprise one or more (e.g., 1, 2, 3, 4, 5, or more) mutations. In some embodiments, one or more (e.g., 1, 2, 3, 4, 5, or more) portions of the engineered protein are portions of the amino acid sequence of SEQ ID NO: 126 (SYN3288), and the one or more portions can be the same as or different from one another. In some embodiments, the engineered protein comprises at least two different portions of the amino acid sequence of SEQ ID NO: 126 (SYN3288), and the first portion and / or the second portion can comprise one or more mutations described herein. In some embodiments, the engineered protein comprises at least three different portions of the amino acid sequence of SEQ ID NO: 126 (SYN3288), and the first portion and / or the second portion can comprise one or more mutations described herein, and the third portion (e.g., a heterologous polypeptide) can not comprise the one or more mutations. In some embodiments, the Type V CRISPR-Cas effector polypeptide of the engineered protein comprises the one or more mutations. In some embodiments, the engineered protein, optionally when optimally aligned with the amino acid sequence of SEQ ID NO: 126, comprises one or more amino acid mutations selected from the group consisting of: D156R, G678R, K684V, Y688R, K741R, and any combination thereof, with reference to the position numbering of SEQ ID NO: 126.
[0193] The second polypeptide and / or heterologous polypeptide of the engineered protein of the present application can be from about 10 to about 300 amino acids in length, such as from about 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids to about 110, 125, 150, 175, 200, 225, 250, 275, or 300 amino acids. In some embodiments, the second polypeptide and / or heterologous polypeptide is about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, or 300 amino acids in length. In some embodiments, the second polypeptide and / or heterologous polypeptide is from about 120, 125, 130, 135, or 140 amino acids to about 145, 150, 155, or 160 amino acids in length. In some embodiments, the second polypeptide and / or heterologous polypeptide is 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, or 160 amino acids in length. In some embodiments, the second polypeptide and / or heterologous polypeptide is located between the first CRISPR-Cas effector polypeptide and the third CRISPR-Cas effector polypeptide, and the second polypeptide and / or heterologous polypeptide is heterologous to one or both of the first and third CRISPR-Cas effector polypeptides.
[0194] In some embodiments, the CRISPR-Cas effector polypeptide (e.g., the first and / or third polypeptide) of the engineered protein of the application is about 100, 150, 200, or 250 amino acids to about 300, 350, or 400 amino acids in length. In some embodiments, the CRISPR-Cas effector polypeptide is about 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, or 400 amino acids in length. In some embodiments, the CRISPR-Cas effector polypeptide is about 800, 850, or 900 amino acids to about 950, 1,000, 1,050, or 1,100 amino acids in length. In some embodiments, the CRISPR-Cas effector polypeptide is about 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1,000, 1,010, 1,020, 1,030, 1,040, 1,050, 1,060, 1,070, 1,080, 1,090, or 1,100 amino acids in length. In some embodiments, the engineered protein comprises a first CRISPR-Cas effector polypeptide that is about 100, 150, 200, or 250 amino acids to about 300, 350, or 400 amino acids in length; a heterologous polypeptide that is about 10, 50, 100, or 140 amino acids to about 160, 200, 250, or 300 amino acids in length; and a second CRISPR-Cas effector polypeptide that is about 100, 200, 300, 400, 500, 600, 700, 800, 850, or 900 amino acids to about 950, 1,000, 1,050, 1,100 amino acids in length.
[0195] In some embodiments, the heterologous polypeptide (e.g., the second polypeptide) comprises a nuclease domain or a portion thereof, which can be referred to herein as a “heterologous nuclease domain or a portion thereof” because the nuclease domain or a portion thereof from the heterologous polypeptide is heterologous to the one or more CRISPR-Cas effector polypeptides present in the engineered protein. The heterologous polypeptide can be a DNA nuclease domain or a portion thereof. In some embodiments, the heterologous nuclease domain or a portion thereof is from a CRISPR-Cas effector protein. In some embodiments, the heterologous nuclease domain or a portion thereof is not from a CRISPR-Cas effector protein. In some embodiments, the heterologous nuclease domain or a portion thereof is from a bacterial protein, optionally wherein the heterologous nuclease domain or a portion thereof is from a restriction endonuclease, a homing endonuclease, colicin, pyocin, a reverse transcriptase, a DNase, and / or a standalone HNH domain. In some embodiments, the engineered protein comprises a heterologous polypeptide that includes a nuclease domain or a portion thereof (i.e., a heterologous nuclease domain or a portion thereof), and the engineered protein is a nuclease that optionally cleaves a target strand of a target nucleic acid and / or a non-target strand of a target nucleic acid. In some embodiments, the engineered protein cleaves a target strand of a target nucleic acid and a non-target strand of a target nucleic acid, and provides a blunt-end double-stranded break of a target nucleic acid or a staggered double-stranded break of a target nucleic acid. In some embodiments, the engineered protein cleaves a target strand of a target nucleic acid and a non-target strand of a target nucleic acid, and the distance (e.g., number of nucleotides) between the cleavage sites is 0, 1, 2, 3, 4, or 5 nucleotides to about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides.
[0196] In some embodiments, the heterologous nuclease domain or a portion thereof can be a target strand nickase domain or a portion thereof. As used herein, a “target strand nickase domain or a portion thereof” refers to a polypeptide that has nickase activity on a target strand of a target nucleic acid when the domain or a portion thereof is in its native protein. That is, the target strand nickase domain or a portion thereof can or is capable of making a nick (e.g., cleaving or breaking) a target strand of a target nucleic acid (also referred to as the sense strand (e.g., “+”; template)) when the domain or a portion thereof is in its native protein. For example, the HNH domain of Cas9 makes a nick and / or has nickase activity on a target strand of a target nucleic acid. As used herein, “nickase activity” refers to the presence of a single-strand break in a nucleic acid.
[0197] In some embodiments, the target strand nuclease domain or a portion thereof can have nuclease activity on the target strand of the target nucleic acid when present in the engineered protein. In some embodiments, the target strand nuclease domain or a portion thereof can have nuclease activity on the non-target strand (also referred to as the anti-sense strand (e.g., “-”; complementary)) of the target nucleic acid when present in the engineered protein. When the target strand nuclease domain or a portion thereof in the engineered protein has nuclease activity on both the target strand and the non-target strand, the target strand nuclease domain or a portion thereof can sequentially cleave both strands. In some embodiments, the target strand nuclease domain or a portion thereof in the engineered protein has a higher activity (e.g., enzymatic activity) on the target strand of the target nucleic acid than on the non-target strand. For example, the target strand nuclease domain or a portion thereof can preferentially or more quickly cleave the target strand of the target nucleic acid than the non-target strand of the target nucleic acid when present in the engineered protein.
[0198] As used herein, a “target strand specific nickase domain” refers to a polypeptide that has nuclease activity only on the target strand of a target nucleic acid and not on the non-target strand of the target nucleic acid. As used herein, a “non-target strand specific nickase domain” refers to a polypeptide that has nuclease activity only on the non-target strand of a target nucleic acid and not on the target strand of the target nucleic acid. As used herein, a “target strand and non-target strand nickase domain” refers to a polypeptide that has nuclease activity on both the target strand and the non-target strand of a target nucleic acid. In some embodiments, the engineered protein comprises a target strand nuclease domain or a portion thereof, and the target strand nuclease domain or a portion thereof is a target strand specific nickase domain in the engineered protein. In some embodiments, the engineered protein comprises a target strand nuclease domain or a portion thereof, and the target strand nuclease domain or a portion thereof is a non-target strand specific nickase domain in the engineered protein. In some embodiments, the engineered protein comprises a target strand nuclease domain or a portion thereof, and the target strand nuclease domain or a portion thereof is a target strand and non-target strand nickase domain in the engineered protein.
[0199] The engineered protein can comprise a heterologous polypeptide (e.g., a second polypeptide) comprising a target strand nuclease domain or a portion thereof. Thus, the engineered protein can have nuclease activity on a target strand of a target nucleic acid and / or on a non-target strand of a target nucleic acid. As such, the engineered protein can be a target strand nuclease and / or a non-target strand nuclease. As used herein, in reference to an engineered protein, a “target strand nuclease” means that the engineered protein can or is capable of cleaving a target strand of a target nucleic acid. As used herein, in reference to an engineered protein, a “non-target strand nuclease” means that the engineered protein can or is capable of cleaving a non-target strand of a target nucleic acid. As used herein, in reference to an engineered protein, a “target strand and non-target strand nuclease” means that the engineered protein can or is capable of cleaving both a target strand and a non-target strand of a target nucleic acid, in any order (e.g., sequentially or simultaneously). In some embodiments, the engineered protein is a target strand nuclease and / or has nuclease activity on a target strand of a target nucleic acid. In some embodiments, the engineered protein is a non-target strand nuclease and / or has nuclease activity on a non-target strand of a target nucleic acid. In some embodiments, the engineered protein is a target strand and non-target strand nuclease and / or has nuclease activity on both a target strand and a non-target strand of a target nucleic acid.
[0200] In some embodiments, the heterologous polypeptide of the engineered protein comprises a target strand nuclease domain or a portion thereof, and the target strand nuclease domain or a portion thereof of the engineered protein has nuclease activity on a target strand of the target nucleic acid, whereby the engineered protein is a target strand nuclease. In some embodiments, the heterologous polypeptide of the engineered protein comprises a target strand nuclease domain or a portion thereof, and the target strand nuclease domain or a portion thereof of the engineered protein has nuclease activity on a non-target strand of the target nucleic acid, whereby the engineered protein is a non-target strand nuclease. In some embodiments, the heterologous polypeptide of the engineered protein comprises a target strand nuclease domain or a portion thereof, and the target strand nuclease domain or a portion thereof of the engineered protein has nuclease activity on both a target strand and a non-target strand of the target nucleic acid, whereby the engineered protein is a target strand and non-target strand nuclease. In some embodiments, the heterologous polypeptide of the engineered protein comprises a target strand nuclease domain or a portion thereof, and the target strand nuclease domain or a portion thereof of the engineered protein has at least nuclease activity on a target strand of the target nucleic acid, and the CRISPR-Cas effector polypeptide of the engineered protein comprises a nuclease domain or a portion thereof that has at least nuclease activity on a non-target strand of the target nucleic acid, whereby the engineered protein is a target strand and non-target strand nuclease. In some embodiments, the CRISPR-Cas effector polypeptide of the engineered protein comprises a nuclease domain or a portion thereof that is a target strand and non-target strand nuclease domain or a portion thereof, but the nuclease domain or a portion thereof is inactivated, such that the nuclease activity on the target strand is inactivated, whereby the target strand of the target nucleic acid is not nicked by the nuclease domain or a portion thereof.
[0201] In some embodiments, the engineered protein can comprise one or more (e.g., 1, 2, or more) nuclease domains or portions thereof. In some embodiments, the engineered protein comprises at least two different nuclease domains or portions thereof. In some embodiments, the engineered protein can comprise a native nuclease domain, optionally one or more (e.g., 1, 2, or more) native nuclease domains. As used herein, a “native nuclease domain” refers to a nuclease domain that is naturally found in a CRISPR-Cas effector protein. In some embodiments, the engineered protein comprises a first heterologous nuclease domain (e.g., from and / or present in a heterologous polypeptide) and a second nuclease domain. The second nuclease domain can be from and / or present in a CRISPR-Cas effector protein. In some embodiments, the first nuclease domain can be a native nuclease domain and / or the second nuclease domain can be a native nuclease domain. In some embodiments, the second nuclease domain is a non-target and target strand nickase domain or portion thereof. As used herein, a “non-target and target strand nickase domain or portion thereof” refers to a polypeptide that, when the domain or portion thereof is in its native protein, has nickase activity on both the non-target strand of a target nucleic acid and the target strand of the target nucleic acid, and cuts the non-target strand before the target strand, or is more prone to or cuts the non-target strand faster than the target strand. The non-target and target strand nickase domain or portion thereof can provide a staggered double-stranded break in a target nucleic acid. In some embodiments, the second nuclease domain is active. In some embodiments, the second nuclease domain is inactivated (i.e., dead, inactive, or lacks nickase activity). In some embodiments, the second nuclease domain only cuts the non-target strand of a target nucleic acid and / or comprises a mutation that inactivates the nickase activity on the target strand of the target nucleic acid. The nuclease domain or portion thereof in the engineered protein can be inactivated by a mutation in the nuclease domain or portion thereof that removes or inactivates the nickase activity. In some embodiments, the engineered protein comprises a nuclease domain or portion thereof from a Type V CRISPR-Cas effector protein, such as Casl2a (e.g., from one of SEQ ID NOs: 50-66 or 180) or Casl2b (e.g., from SEQ ID NO: 151). In some embodiments, the nuclease domain is a RuvC domain from a Type V CRISPR-Cas effector protein, such as Casl2a or Casl2b. The engineered protein can comprise one or more nuclease domains that provide a blunt double-stranded break of a target nucleic acid or a staggered double-stranded break of a target nucleic acid.
[0202] In some embodiments, the heterologous polypeptide (e.g., the second polypeptide) of the engineered protein comprises all or a portion of an HNH domain of a CRISPR-Cas effector protein. The heterologous polypeptide and / or the HNH domain can comprise and / or form a zinc finger motif. In some embodiments, the heterologous polypeptide and / or the HNH domain is about 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids in length to about 110, 125, 150, 175, 200, 225, 250, 275, or 300 amino acids in length. The heterologous polypeptide and / or the HNH domain can comprise about 25 or 30 to about 40 or 45 amino acids, and / or can comprise one or at least two histidines and an asparagine, optionally in a nucleic acid binding and cleavage site. In some embodiments, the heterologous polypeptide and / or the HNH domain can comprise about 25 or 30 to about 40 or 45 amino acids, including two histidines and one asparagine, which are present in and / or form a zinc finger motif. The heterologous polypeptide and / or the HNH domain can comprise and / or form two anti-parallel beta strands connected by a loop, and / or can comprise an alpha helix, optionally wherein a histidine is present in at least one of the beta strands, an asparagine is present in the loop, and / or a histidine or asparagine is present in the alpha helix. The heterologous polypeptide can comprise all or a portion of an HNH domain having the structure described in Pediaditakis M, et al. J Bacteriol. 194(22); 6184-6194. In some embodiments, the heterologous polypeptide of the engineered protein comprises all or a portion of an HNH domain of a Type II CRISPR-Cas effector protein (e.g., a Cas9 HNH domain). The heterologous polypeptide of the engineered protein can comprise all or a portion of an inactive HNH domain (e.g., a Cas9 HNH domain). The HNH domain or portion thereof can have an inactivating mutation (e.g., a mutation that removes nickase activity). In some embodiments, the heterologous polypeptide of the engineered protein comprises all or a portion of an HNH domain having an inactivating mutation, and / or the HNH domain is inactive (e.g., does not have nickase activity). In some embodiments, the heterologous polypeptide of the engineered protein comprises all or a portion of an inactivated HNH domain having a H840A mutation. In some embodiments, the heterologous polypeptide of the engineered protein comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the amino acid sequence of one or more of SEQ ID NOs: 1 or 169-174.In some embodiments, the heterologous polypeptide of the engineered protein comprises the amino acid sequence of any one of SEQ ID NOs: 1 or 169-174.
[0203] In some embodiments, the heterologous polypeptide of the engineered protein comprises an amino acid sequence that, when the amino acid sequence of the heterologous polypeptide is optimally aligned with SEQ ID NO: 81, has an amino acid residue that is not a histidine residue at a position corresponding to amino acid residue number 839 of SEQ ID NO: 81. In some embodiments, the heterologous polypeptide of the engineered protein comprises an amino acid sequence that, when the amino acid sequence of the heterologous polypeptide is optimally aligned with SEQ ID NO: 1, has an amino acid residue that is not a histidine residue at a position corresponding to amino acid residue number 75 of SEQ ID NO: 1. In some embodiments, the heterologous polypeptide of the engineered protein comprises an amino acid sequence that, when the amino acid sequence of the heterologous polypeptide is optimally aligned with SEQ ID NO: 81, has an alanine residue at a position corresponding to amino acid residue number 839 of SEQ ID NO: 81. In some embodiments, the heterologous polypeptide of the engineered protein comprises an amino acid sequence that, when the amino acid sequence of the heterologous polypeptide is optimally aligned with SEQ ID NO: 1, has an alanine residue at a position corresponding to amino acid residue number 75 of SEQ ID NO: 1.
[0204] In some embodiments, the heterologous polypeptide of the engineered protein can be located between and / or linked to (e.g., directly or indirectly) two consecutive or non-consecutive amino acids present in a CRISPR-Cas effector protein. In some embodiments, the engineered protein is prepared by inserting the heterologous polypeptide between two consecutive or non-consecutive amino acids of a CRISPR-Cas effector protein or a portion thereof. In some embodiments, the engineered protein can comprise, in an amino-terminal to carboxyl-terminal direction, a first CRISPR-Cas effector polypeptide, a heterologous polypeptide, and a second CRISPR-Cas effector polypeptide, wherein the first and second CRISPR-Cas effector polypeptides are from the same CRISPR-Cas effector protein.
[0205] In some embodiments, the CRISPR-Cas effector polypeptide comprises a portion of a Type V CRISPR-Cas effector protein (such as Casl2a or Casl2b). The CRISPR-Cas effector polypeptide can comprise all or a portion of a nucleic acid binding domain, such as a nucleic acid binding domain from a Type V CRISPR-Cas effector protein (e.g., Casl2a or Casl2b). In some embodiments, the CRISPR-Cas effector polypeptide of the engineered protein comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a portion of the amino acid sequence of one or more of SEQ ID NOs: 50-66, 151, or 180. In some embodiments, the CRISPR-Cas effector polypeptide comprises a portion of the amino acid sequence of any one of SEQ ID NOs: 50-66, 151, or 180. In some embodiments, the engineered protein comprises two or more (e.g., 2, 3, 4, or more) separated portions of the amino acid sequence of any one of SEQ ID NOs: 50-66, 151, or 180.
[0206] In some embodiments, the engineered proteins of the application can lack about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more amino acids that are present in a CRISPR-Cas effector protein, such as a CRISPR-Cas effector protein having the sequence of any one of SEQ ID NOs: 50-66, 151, or 180. In some embodiments, the engineered proteins of the application can lack 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids that are present in a CRISPR-Cas effector protein, such as a CRISPR-Cas effector protein having the sequence of any one of SEQ ID NOs: 50-66, 151, or 180. For example, the engineered protein can lack one or more amino acids from amino acid residue 283 to amino acid residue 293 of SEQ ID NO: 50 or SEQ ID NO: 58 or SEQ ID NO: 180; amino acid residue 331 to amino acid residue 341 of SEQ ID NO: 55; amino acid residue 312 to amino acid residue 322 of SEQ ID NO: 51; or corresponding amino acid residues of a sequence that optimally aligns with SEQ ID NO: 50, 51, 58, 55, or 180 (e.g., amino acid residues corresponding to amino acid residues 283-293 when a sequence (e.g., SEQ ID NO: 52) optimally aligns with SEQ ID NO: 50). In some embodiments, the engineered protein lacks one or more (e.g., 1, 2, 3, 4 or more) trans- domain linker regions (e.g., regions located between two domains, such as two adjacent domains), which are present in a CRISPR-Cas effector protein that is part of the engineered protein (e.g., a CRISPR-Cas effector protein having the sequence of any one of SEQ ID NOs: 50-66, 151, or 180) and present in the CRISPR-Cas effector protein.
[0207] In some embodiments, the heterologous polypeptide of the engineered protein can be located between and / or linked (e.g., directly or indirectly) to two consecutive or non-consecutive amino acids of a CRISPR-Cas effector protein (e.g., a CRISPR-Cas effector protein having the amino acid sequence of any one of SEQ ID NOs: 50-66, 151, or 180). For example, from N-terminus to C-terminus, the engineered protein can comprise a first CRISPR-Cas effector polypeptide, an HNH domain, and a second CRISPR-Cas effector polypeptide, wherein the first and second CRISPR-Cas effector polypeptides are each part of a CRISPR-Cas effector protein, and the last amino acid residue at the C-terminus of the first CRISPR-Cas effector polypeptide and the first amino acid residue at the N-terminus of the second CRISPR-Cas effector polypeptide are two consecutive or non-consecutive amino acid residues of the CRISPR-Cas effector protein. The heterologous polypeptide can be directly linked to one or both of the two consecutive or non-consecutive amino acids of the CRISPR-Cas effector protein (i.e., without using a linker to link one end of the heterologous polypeptide to an end of a CRISPR-Cas effector polypeptide that is part of the CRISPR-Cas effector protein). In some embodiments, the heterologous polypeptide can be indirectly linked (e.g., by a linker, such as a peptide linker) to one or both of the two consecutive or non-consecutive amino acids of the CRISPR-Cas effector protein. In some embodiments, the heterologous polypeptide of the engineered protein can be located between and / or linked (e.g., directly or indirectly) to two consecutive amino acids of a CRISPR-Cas effector protein (e.g., a CRISPR-Cas effector protein having the amino acid sequence of any one of SEQ ID NOs: 50-66, 151, or 180). In some embodiments, the heterologous polypeptide of the engineered protein can be located between and / or linked (e.g., directly or indirectly) to two non-consecutive amino acids of a CRISPR-Cas effector protein (e.g., a CRISPR-Cas effector protein having the amino acid sequence of any one of SEQ ID NOs: 50-66, 151, or 180).
[0208] In some embodiments, the two contiguous or non-contiguous amino acids are two of the amino acid residues in amino acid residues 250, 260, 270, or 280 to amino acid residues 290, 300, 310, 320, 330, 340, or 350, respectively. In some embodiments, a heterologous polypeptide can be located between and / or linked (e.g., directly or indirectly) to two contiguous or non-contiguous amino acids that are two of the following amino acid residues of a CRISPR-Cas effector protein (e.g., a CRISPR-Cas effector protein having the amino acid sequence of any one of SEQ ID NOs: 50-66, 151, or 180): amino acid residues 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, and 350.In some embodiments, the heterologous polypeptide of the engineered protein can be located between and / or linked (e.g., directly or indirectly) to two noncontiguous amino acids of a CRISPR-Cas effector protein (e.g., a CRISPR-Cas effector protein having the amino acid sequence of any one of SEQ ID NOs: 50-66, 151, or 180), wherein one of the two noncontiguous amino acid residues is amino acid residue 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285, and the other is amino acid residue 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, or 305, optionally from SEQ ID NOs: 50-66, 151, or 180. In some embodiments, the heterologous polypeptide is located between and / or linked (e.g., directly or indirectly) to amino acid residues 290 and 291, amino acid residues 291 and 292, amino acid residues 292 and 293, amino acid residues 293 and 294, amino acid residues 320 and 321, amino acid residues 321 and 322, amino acid residues 339 and 340, or amino acid residues 340 and 341 of a CRISPR-Cas effector protein (e.g., a CRISPR-Cas effector protein having the amino acid sequence of any one of SEQ ID NOs: 50-66, 151, or 180).For example, in some embodiments, the heterologous polypeptide can be located between and / or linked (e.g., directly or indirectly) to amino acid residues 290 and 291 of SEQ ID NO: 50; amino acid residues 291 and 292 of SEQ ID NO: 50; amino acid residues 291 and 292 of SEQ ID NO: 58; amino acid residues 292 and 293 of SEQ ID NO: 58; amino acid residues 320 and 321 of SEQ ID NO: 51; amino acid residues 321 and 322 of SEQ ID NO: 51; amino acid residues 322 and 323 of SEQ ID NO: 51; amino acid residues 339 and 340 of SEQ ID NO: 55; amino acid residues 340 and 341 of SEQ ID NO: 55; amino acid residues 290 and 291 of SEQ ID NO: 180; amino acid residues 291 and 292 of SEQ ID NO: 180; amino acid residues 292 and 293 of SEQ ID NO: 180; or corresponding amino acid residues of a sequence optimally aligned with one of SEQ ID NOs: 50, 51, 58, 55, or 180 (e.g., amino acid residues corresponding to amino acid residues 291 and 292 when a sequence (e.g., SEQ ID NO: 52) is optimally aligned with SEQ ID NO: 50). In some embodiments, the heterologous polypeptide of the engineered protein can be located in an interdomain linker region of a CRISPR-Cas effector protein (e.g., a region located between two domains, such as two adjacent domains). In some embodiments, the heterologous polypeptide can be positioned in the engineered protein such that it is adjacent to an exposed portion of a target strand of a target nucleic acid.
[0209] In some embodiments, the engineered protein comprises all or a portion of a wedge domain, a Rec1 domain, a Rec2 domain, a PAM interaction domain, a RuvC domain, a bridge helix, and / or a Nuc domain, each of which can be from a Type V CRISPR-Cas effector protein, such as Cas12a, Cas12b, and / or a protein having the sequence of any one of SEQ ID NOs: 50-66, 151, or 180. In some embodiments, the engineered protein comprises all or a portion of a Cas12a domain, which has a structure as described in Yamano, Takashi, et al. Mol Cell 67:633-645 (2017). In some embodiments, the engineered protein comprises all or a portion of an alpha-helix recognition (REC) lobe, optionally wherein the domains of all or a portion of the REC lobe in the engineered protein can differ in order and / or structure from that in a Type V CRISPR-Cas effector protein. The REC lobe can contain a Rec1 domain and a Rec2 domain. The Rec1 domain can comprise 13 alpha helices, and / or the Rec2 domain can comprise 10 alpha helices and two beta strands that can form a small antiparallel sheet. In some embodiments, the heterologous polypeptide of the engineered protein can be located between a polypeptide of all or a portion of the Rec1 domain and a polypeptide of all or a portion of the Rec2 domain, each of which can be from a Type V CRISPR-Cas effector protein, such as Cas12a, Cas12b, and / or a protein having the sequence of any one of SEQ ID NOs: 50-66, 151, or 180. In some embodiments, all or a portion of the heterologous polypeptide is located at an exposed surface or interface of the engineered protein. In some embodiments, the CRISPR-Cas effector polypeptide of the engineered protein comprises all or a portion of a RuvC domain. As understood by one of skill in the art, some domains (e.g., the wedge domain and the RuvC domain of Cas12a) are not contiguous in sequence and can be divided into two or more (e.g., 2, 3, 4, or more) non-contiguous sequences. For example, a polypeptide of Cas12 can have, from N-terminus to C-terminus, a first portion of a wedge domain (WED-1), a Rec1 domain, a Rec2 domain, a second portion of a wedge domain (WED-2), a PAM interaction domain (PI), a third portion of a wedge domain (WED-3), a first portion of a RuvC domain (RuvC-1), a bridge helix, a second portion of a RuvC domain (RuvC-2), a Nuc domain, and a third portion of a RuvC domain (RuvC-3). In some embodiments, the engineered protein comprises all or a portion of an active RuvC domain.In some embodiments, the engineered protein comprises all or a portion of an inactivated RuvC domain, optionally all or a portion of an inactivated RuvC domain with a D10A mutation. In some embodiments, the engineered protein comprises all or a portion of an inactivated RuvC domain, and the polypeptide comprising all or a portion of an inactivated RuvC domain has an alanine at a position corresponding to amino acid residue 831 of SEQ ID NO: 50 when the polypeptide is optimally aligned with SEQ ID NO: 50, optionally wherein the mutation is referred to as a D10A and / or D832A mutation.
[0210] The CRISPR-Cas effector polypeptide can comprise a nuclease, optionally a RuvC-like nuclease. In some embodiments, the CRISPR-Cas effector polypeptide comprises a RuvC domain or a portion thereof. In some embodiments, the CRISPR-Cas effector polypeptide comprises a nuclease belonging to the RNase H superfamily. In some embodiments, the CRISPR-Cas effector polypeptide comprises an RNase H-like enzyme with a catalytic core, which can include a beta sheet comprising five beta strands ordered 3, 2, 1, 4, 5, optionally wherein beta strand 2 is antiparallel to the other beta strands. Flanking the central beta sheet can be alpha helices, the number of which can vary between related enzymes. In some embodiments, the CRISPR-Cas effector polypeptide comprises an RNase H-like catalytic core, wherein active site residues include one or more of aspartate, glutamate, and histidine. In some embodiments, the CRISPR-Cas effector polypeptide comprising an RNase H-like catalytic core can include a negatively charged side chain in the active site of the RNase H-like polypeptide that directly or through a water molecule participates in coordinating a divalent metal ion. In some embodiments, the CRISPR-Cas effector polypeptide comprises an RNase H-like catalytic core that uses a catalytic mechanism that depends on two ions, optionally wherein the ions are Mg 2+ and / or Mn 2+ . In some embodiments, the CRISPR-Cas effector polypeptide comprises a nuclease and / or an RNase H-like nuclease, as described in Majorek KA et al. Nucleic Acids Res. 2014; 42(7):4160-4179, which is incorporated by reference herein in its entirety.
[0211] In some embodiments, the CRISPR-Cas effector polypeptide comprises one or more (e.g., 1, 2, 3, 4, or more) mutations. The one or more mutations can improve or alter the activity of the heterologous polypeptide and / or the activity of the CRISPR-Cas effector polypeptide. In some embodiments, the CRISPR-Cas effector polypeptide can comprise an inactivating mutation, such as a D10A mutation in the RuvC domain. In some embodiments, the CRISPR-Cas effector polypeptide comprises all or a portion of a Rec 1 domain comprising one or more (e.g., 1, 2, 3, 4, or more) mutations, such as one or more amino acid residues located in amino acid residues 243-253 and / or in the sequence GFVTESGEKIK (SEQ ID NO: 122) of a CRISPR-Cas effector protein (e.g., a CRISPR-Cas effector protein having the amino acid sequence of any one of SEQ ID NOs: 50-66, 151, or 180). In some embodiments, the CRISPR-Cas effector polypeptide comprises a hairpin and / or the sequence GFVTESGEKIK (SEQ ID NO: 122), and one or more of the amino acid residues in the hairpin and / or sequence can be mutated. In some embodiments, the CRISPR-Cas effector polypeptide comprises a hairpin and / or the sequence GFVTESGEKIK (SEQ ID NO: 122), and all or a portion of the hairpin and / or sequence is deleted. In some embodiments, the CRISPR-Cas effector polypeptide comprises a hairpin and / or the sequence GFVTESGEKIK (SEQ ID NO: 122), and 1, 2, 3, 4, 5, or more amino acid residues are added at one or both ends of the hairpin and / or sequence.
[0212] In some embodiments, the engineered protein comprises, from N- to C-terminus, a first CRISPR-Cas effector polypeptide, an HNH domain, and a second CRISPR-Cas effector polypeptide, wherein the first and second CRISPR-Cas effector polypeptides are each a portion of an inactivated LbCasl2a (e.g., LbCasl2a having the sequence of SEQ ID NO: 50), and the last amino acid residue of the C-terminus of the first CRISPR-Cas effector polypeptide and the first amino acid residue of the N-terminus of the second CRISPR-Cas effector polypeptide are two consecutive amino acid residues in the inactivated LbCasl2a. The HNH domain can be from S. pyogenes Cas9 (SpCas9), and / or can have a sequence comprising SEQ ID NO: 1. In some embodiments, the HNH domain can have the sequence of any one of SEQ ID NOs: 1 or 169-174. The HNH domain can be positioned in the engineered protein such that it is adjacent to an exposed portion of a target strand of a target nucleic acid. The engineered protein can be a target strand nuclease. In some embodiments, the engineered protein only nick the target DNA strand. In some embodiments, the engineered protein is a target strand and non-target strand nuclease.
[0213] One or more (e.g., 1, 2, 3, 4, or more) linkers can be present in the engineered protein. For example, a linker can be present between a CRISPR-Cas effector polypeptide and a heterologous polypeptide (e.g., a second polypeptide). In some embodiments, a linker can be present between a first CRISPR-Cas effector polypeptide and a heterologous polypeptide, and a linker can be present between the heterologous polypeptide and a second CRISPR-Cas effector polypeptide. Exemplary linkers include, but are not limited to, the linkers described herein. In some embodiments, a linker comprises 1 to 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids, and / or comprises glycine and / or serine. In some embodiments, a linker comprises 1, 2, 3, or 4 amino acids that are glycine and / or serine (i.e., a GS linker). In some embodiments, the engineered protein lacks a linker between a CRISPR-Cas effector polypeptide and a heterologous polypeptide. In some embodiments, a heterologous polypeptide is indirectly linked to an amino acid residue at the N-terminus of a CRISPR-Cas effector polypeptide by a linker, and / or a heterologous polypeptide is indirectly linked to an amino acid residue at the C-terminus of a CRISPR-Cas effector polypeptide by a linker. Two or more peptide linkers present in an engineered protein of the disclosure can be selected independently of one another and can be the same as or different from one another. In some embodiments, an engineered protein comprises a peptide linker that is a GS linker having 1, 2, 3, or 4 amino acid residues, optionally 2 or 4 amino acid residues. In some embodiments, an engineered protein comprises a peptide linker having the amino acid sequence (GGS)n, wherein n is an integer from 1-20; GS; SG; and / or the amino acid sequence of any one of SEQ ID NOs: 18-47 or 176-179. n n In some embodiments, a heterologous polypeptide is directly (i.e., without a linker) linked to an amino acid residue at the N-terminus of a CRISPR-Cas effector polypeptide, and / or a heterologous polypeptide is directly (i.e., without a linker) linked to an amino acid residue at the C-terminus of a CRISPR-Cas effector polypeptide.
[0214] In some embodiments, a heterologous polypeptide is directly (i.e., without a linker) linked to an amino acid residue at the N-terminus of a CRISPR-Cas effector polypeptide, and / or a heterologous polypeptide is directly (i.e., without a linker) linked to an amino acid residue at the C-terminus of a CRISPR-Cas effector polypeptide.
[0215] The engineered protein can comprise an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to any one of SEQ ID NOs: 2-17, 125-132, or 157-168. In some embodiments, the engineered protein comprises and / or has the amino acid sequence of any one of SEQ ID NOs: 2-17, 125-132, or 157-168. The engineered protein can have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to all or a portion of the amino acid sequence of a wild-type CRISPR-Cas effector protein. In some embodiments, the engineered protein has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to all or a portion of the amino acid sequence of any one of SEQ ID NOs: 50-66, 151, or 180. In some embodiments, the engineered protein has from about 70%, 75%, or 80% to about 85%, 90%, 95%, or 98% sequence identity to all or a portion of the amino acid sequence of any one of SEQ ID NOs: 50-66, 151, or 180.
[0216] The engineered protein has increased efficiency compared to a CRISPR-Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NO: 50-66, 151, or 180, and / or a wild-type CRISPR-Cas effector protein), such as increased efficiency in nicking a target strand and / or a non-target strand of a target nucleic acid. In some embodiments, the engineered protein has increased efficiency compared to a CRISPR-Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NO: 50-66, 151, or 180, and / or a wild-type CRISPR-Cas effector protein) in nicking a target strand of a target nucleic acid. In some embodiments, the engineered protein can provide an increased number of nicks in a target strand of a target nucleic acid compared to a number of nicks in a target strand of a target nucleic acid having a CRISPR-Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NO: 50-66, 151, or 180, and / or a wild-type CRISPR-Cas effector protein). In some embodiments, the engineered protein has increased efficiency compared to a CRISPR-Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NO: 50-66, 151, or 180, and / or a wild-type CRISPR-Cas effector protein) in modifying a target nucleic acid.
[0217] According to embodiments of the present application, compositions, complexes, and systems comprising the engineered protein can be provided. In some embodiments, the compositions, complexes, and / or systems comprising the engineered protein can be base editing compositions, complexes, and / or systems. The compositions, complexes, and / or systems of the present application can include a guide nucleic acid (e.g., a guide RNA) and / or a deaminase (e.g., a cytosine deaminase and / or an adenine deaminase). In some embodiments, the engineered protein, the guide nucleic acid, and optionally the deaminase form a complex, or are comprised in a complex (e.g., a ribonucleoprotein). The engineered protein, the guide nucleic acid, and optionally the deaminase can not naturally coexist, and / or the complex comprising the engineered protein, the guide nucleic acid, and optionally the deaminase can not naturally coexist. In some embodiments, the engineered protein comprises a deaminase, and / or is fused to a deaminase (e.g., an adenine deaminase and / or a cytosine deaminase).
[0218] Also provided herein are nucleic acid molecules encoding the engineered proteins of the present application, as well as expression cassettes and / or vectors comprising the nucleic acid molecules of the present application.
[0219] According to some embodiments, a method is provided that comprises contacting a target nucleic acid with: an engineered protein of the invention, a guide nucleic acid (e.g., a guide RNA), and optionally a deaminase. In some embodiments, the engineered protein, the guide nucleic acid, and / or the deaminase form a complex, or are comprised in a complex. In some embodiments, the method can modify the target nucleic acid, and / or provide one or more single-strand breaks in the target nucleic acid.
[0220] In some embodiments, a composition, system, method, and / or complex comprising an engineered protein has increased efficiency compared to a composition, system, method, and / or complex comprising a CRISPR-Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NO: 50-66, 151, or 180, and / or a wild-type CRISPR-Cas effector protein). In some embodiments, a composition, system, method, and / or complex comprising an engineered protein that is a target strand nickase has increased efficiency compared to a composition, system, method, and / or complex comprising a CRISPR-Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NO: 50-66, 151, or 180, and / or a wild-type CRISPR-Cas effector protein). This can be because nicking the target strand can increase the efficiency of genome editing tools, such as base editors and / or base diversifiers. In some embodiments, a composition, system, method, and / or complex comprising an engineered protein can provide an increased number of target strand breaks in a target nucleic acid compared to a target nucleic acid having a composition, system, method, and / or complex comprising a CRISPR-Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NO: 50-66, 151, or 180, and / or a wild-type CRISPR-Cas effector protein).
[0221] The engineered proteins and / or compositions, systems, methods, and / or complexes comprising the engineered proteins can provide: improved or altered size and / or composition of insertions and / or deletions (e.g., indels), improved or altered size of deletions in a target nucleic acid, improved or altered nicking ability on either strand (i.e., a target strand or a non-target strand of a target nucleic acid), and / or increased nuclease activity as compared to a CRISPR-Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NOs: 50-66, 151, or 180, and / or a wild-type CRISPR-Cas effector protein) and / or a composition, system, method, and / or complex comprising a CRISPR-Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NOs: 50-66, 151, or 180, and / or a wild-type CRISPR-Cas effector protein). In some embodiments, the engineered proteins and / or compositions, systems, methods, and / or complexes comprising the engineered proteins confer nuclease functionality to a catalytically-inactivated CRISPR-Cas effector protein. In some embodiments, the engineered proteins and / or compositions, systems, methods, and / or complexes comprising the engineered proteins provide a different editing profile and / or a different cleavage pattern of a target nucleic acid as compared to an editing profile and / or a cleavage pattern of a target nucleic acid of a Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NOs: 50-66, 151, or 180, and / or a wild-type CRISPR-Cas effector protein) and / or a composition, system, method, and / or complex comprising a CRISPR-Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NOs: 50-66, 151, or 180, and / or a wild-type CRISPR-Cas effector protein).
[0222] In some embodiments, the methods of the application have increased efficiency in modifying a target nucleic acid as compared to the efficiency of a control method (e.g., a method comprising contacting a target nucleic acid with a CRISPR-Cas effector protein (e.g., Cas12a, a CRISPR-Cas effector protein having a sequence of SEQ ID NOs: 50-66, 151, or 180 amino acid sequence, and / or a wild-type CRISPR-Cas effector protein) and / or a method lacking the engineered protein.
[0223] As described herein, the polypeptides, engineered proteins, nucleic acids, expression cassettes, and / or vectors of the application can be codon optimized for expression in an organism. The organism useful in the application can be any organism or cell thereof for which nucleic acid modification is useful. The organism can include, but is not limited to, any animal (e.g., a mammal), any plant, any fungus, any archaea, or any bacteria. In some embodiments, the organism can be a plant or cell thereof. In some embodiments, the organism is an animal, such as a mammal (e.g., a human).
[0224] The target nucleic acid can be a genomic sequence from any organism (e.g., a eukaryote, such as a mammal or a plant). In some embodiments, the target nucleic acid is a genomic sequence from a model organism, such as, but not limited to, Escherichia coli, an immortalized human cell line (e.g., HEK293, HeLa, etc.), Caenorhabditis elegans, and / or Drosophila Melanogaster. In some embodiments, the target nucleic acid is a genomic sequence from a non-model organism. Exemplary non-model organisms include, but are not limited to, a crop plant (e.g., a fruit crop plant, a vegetable crop plant, and / or a field crop plant) and / or an animal, such as a human, a primate, and / or a mouse. In some embodiments, the non-model organism is a crop plant, such as corn, soybean, wheat, or canola. In some embodiments, the non-model organism is an animal used to test and / or use human therapeutics.
[0225] A target nucleic acid of any plant or plant part can be modified using the nucleic acid constructs of the application. Any plant (or group of plants, such as a genus or higher classification), including angiosperms, gymnosperms, monocots, dicots, C3, C4, CAM plants, bryophytes, ferns, and / or fern allies, microalgae, and / or macroalgae, can be modified using the engineered proteins of the application. The plants and / or plant parts useful in the application can be plants and / or plant parts of any plant species / variety / cultivar. As used herein, the term “plant part” includes, but is not limited to, an embryo, pollen, an ovule, a seed, a leaf, a stem, a shoot, a flower, a branch, a fruit, a kernel, an ear, a cob, a husk, a stalk, a root, a root tip, an anther, a plant cell (including a whole plant cell in a plant and / or plant part), a plant protoplast, a plant tissue, a plant cell tissue culture, a plant callus, a plant clump, and the like. As used herein, “shoot” refers to the above ground portion, including leaves and stems. Further, as used herein, “plant cell” refers to a structural and physiological unit of a plant that contains a cell wall and can also refer to a protoplast. The plant cell can be in isolated single cell form, or can be a cultured cell, or can be a higher order tissue unit, such as a part of a plant tissue or plant organ.
[0226] Non-limiting examples of plants that can be used in the present application include turf grasses (e.g., bluegrass, bentgrass, ryegrass, fescue), feather reed grass, pampas grass, miscanthus, giant reed, switchgrass, vegetable crops including artichoke, kohlrabi, arugula, leek, asparagus, lettuce (e.g., iceberg lettuce, leaf lettuce, romaine lettuce), yellow fleshed yam, melon (e.g., muskmelon, watermelon, crenshaw melon, cantaloupe, honeydew melon), brassica crops (e.g., broccoli, cabbage, cauliflower, kale, collard greens, bok choy, Chinese cabbage), prickly chrysanthemum, radish, chard, okra, onion, celery, parsley, chickpea, cow parsnip, chicory, pepper, potato, cucurbitaceae (e.g., zucchini, cucumber, green zucchini, pumpkin, squash, cantaloupe, watermelon, honeydew melon), radish, dry bulb onion, turnip rape, eggplant, ginseng, broad-leaved chicory, shallot, endive, garlic, spinach, green onion, pumpkin, leafy greens, sugar beet (sugar beet and fodder beet), sweet potato, turnip greens, horseradish, tomato, red radish, and spices; fruit crops such as apple, apricot, cherry, nectarine, peach, pear, plum, prune, cherry, sea buckthorn, fig, nuts (e.g., chestnut, pecan, pistachio, hazelnut, macadamia, peanut, walnut, macadamia, almond, etc.), citrus (e.g., clementine, kumquat, orange, grapefruit, tangerine, mandarin, lemon, lime, etc.), blueberry, raspberry, boysenberry, cranberry, currant, gooseberry, loganberry, raspberry, strawberry, blackberry, grape (wine grape and table grape), avocado, banana, kiwifruit, persimmon, pomegranate, pineapple, tropical fruit, pome, mango, papaya, and lychee, field crop plants such as clover, alfalfa, timothy, evening primrose, meadow grass, corn / maize (fodder corn, sweet corn, popcorn), hops, jojoba, buckwheat, safflower, quinoa, wheat, rice, barley, rye, millet, sorghum, oat, triticale, kudzu, tobacco, kapok, legume (bean (e.g., green bean and dry bean), lentil, pea (e.g., field pea, snow pea, snap pea), soybean, chickpea (chickpea)), oil plant (rape, rapeseed, mustard, poppy, olive, sunflower, coconut, castor oil plant, cocoa bean, peanut, oil palm), duckweed, Arabidopsis thaliana, fiber plant (cotton, flax, hemp), Cannabis (e.g., Cannabis sativa, Cannabis indica, and Cannabis ruderalis), Lauraceae (cinnamon, camphor), or plants such as coffee, sugar cane, tea, and natural rubber plants;and / or ornamental plants (e.g., roses, tulips, violets), and trees, such as forest trees (broadleaf and evergreen trees, such as conifers; e.g., elm, ash, oak, maple, fir, spruce, cedar, pine, birch, cypress, eucalyptus, willow) and shrubs and other nursery stock. In some embodiments, the nucleic acid constructs of the application and / or expression cassettes and / or vectors encoding the same can be used to modify corn, soybean, wheat, canola, rice, tomato, pepper, sunflower, raspberry, blackberry, black raspberry, and / or cherry.
[0227] In some embodiments, the application provides a cell (e.g., a plant cell, an animal cell, a bacterial cell, an archaeal cell, etc.) comprising a polypeptide, a polynucleotide, a nucleic acid construct, an expression cassette, or a vector of the application.
[0228] The application further comprises a kit or kits for carrying out the methods of the application. The kits of the application can comprise reagents, buffers, and equipment for mixing, measuring, sorting, labeling, etc., as well as instructions for use in modifying a target nucleic acid, etc.
[0229] In some embodiments, the application provides a kit comprising one or more nucleic acid constructs of the application as described herein, and / or an expression cassette and / or vector and / or cell comprising the same, and optionally instructions for use thereof. In some embodiments, the kit can further comprise a CRISPR-Cas guide nucleic acid (corresponding to an engineered protein, which can be encoded by a polynucleotide of the application) and / or an expression cassette and / or vector and / or cell comprising the same. In some embodiments, the guide nucleic acid can be provided on the same expression cassette and / or vector as one or more nucleic acid constructs of the application. In some embodiments, the guide nucleic acid can be provided on a separate expression cassette or vector from an expression cassette or vector comprising one or more nucleic acid constructs of the application.
[0230] Thus, in some embodiments, a kit is provided comprising a nucleic acid construct comprising (a) a polynucleotide as provided herein, and (b) a promoter driving expression of the polynucleotide of (a). In some embodiments, the kit can further comprise a nucleic acid construct encoding a guide nucleic acid, wherein the construct comprises a cloning site for cloning a nucleic acid sequence identical or complementary to a target nucleic acid sequence into the backbone of the guide nucleic acid.
[0231] In some embodiments, the nucleic acid construct of the application can be an mRNA, which can encode one or more introns within the encoded polynucleotide. In some embodiments, the nucleic acid construct of the application and / or expression cassette and / or vector comprising the same can further encode one or more selectable markers useful for identifying transformants (e.g., nucleic acids encoding antibiotic resistance genes, herbicide resistance genes, etc.).
[0232] The polypeptides, polynucleotides, nucleic acid constructs, expression cassettes, vectors, compositions, kits, systems, and / or cells of the application can comprise all or a portion of the sequence of one or more of SEQ ID NOs: 1-192. In some embodiments, the polypeptides, polynucleotides, nucleic acid constructs, expression cassettes, vectors, compositions, kits, systems, and / or cells of the application can comprise at least about 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more contiguous amino acids of the sequence of one or more of SEQ ID NOs: 1-192.
[0233] The application will now be described with reference to the following examples. It is to be understood that these examples do not intend to limit the scope of the claims of the application, but rather are intended as examples of certain embodiments. Any variations of the example methods that occur to the skilled person are intended to fall within the scope of the application.
[0234] Example
[0235] Example 1:
[0236] Using existing domain annotations and visual inspection of the SpCas9 crystal structure (PDB ID 4UN3) in PyMOL (PyMOL Molecular Graphics System, Version 2.0, Schrodinger, LLC), the complete HNH domain from SpCas9 was first identified ( Figure 3 ), and its residue boundaries were determined. The domain is mostly resolved in the crystal structure, but several residues connecting the N-terminus of the HNH domain to the Rec1 domain are not resolved in the crystal structure. The position of the Cas9 target DNA strand cleavage site relative to the HNH domain was also recorded. This relative orientation was mimicked in the subsequent rational positioning of the HNH domain relative to the Cas12a target DNA strand.
[0237] Next, the crystal structure of the LbCas12a ternary complex (PDB ID 5XUS) was examined to locate the accessible region of the target DNA strand. Although the side of the target DNA / crRNA duplex closest to the RuvC domain was largely obscured by other Cas12a domains, the interface between the Rec1 and Rec2 domains on the opposite side of the protein ( Figure 4 A portion of the target DNA is exposed at this location (by...) Figure 4 (Indicated by the left arrow in the diagram). The connector connecting these two structural domains (by... Figure 4 The linker (indicated by the right arrow) is located near this exposure site and has little interaction with other residues in LbCas12a; this linker was selected as a candidate site for domain insertion.
[0238] To determine the precise placement of the SpCas9 HNH domain relative to LbCas12a, the exposed DNA bases in the groove between the Rec1 and Rec2 domains of LbCas12a were identified. Then, the HNH domain and its target DNA (two bases on each side of the cleavage site, for a total of four bases) were treated as a unit, and PyMOL was used to test the alignment of the HNH target DNA with the exposed target strand of LbCas12a in a sliding window until an alignment was identified that placed the HNH domain near the insertion loop and minimized conflicts with other domains of LbCas12a. The position of the HNH domain was then manually adjusted using PyMOL to minimize conflicts between HNH and Cas12a.
[0239] The final choice of the HNH domain location is as follows: Figure 5 As shown. Although the C-terminus of the HNH domain is very close to the C-terminus of the insertion loop, the N-terminus of the HNH domain is relatively far from the insertion loop; however, this structure does not include unstructured residues connecting SpCas9 Rec1 and the HNH domain. A highly conserved hairpin in this region that interacts with the target DNA / crRNA duplex was further identified as a potential site for subsequent design.
[0240] To prepare the Cas12a-HNH fusion structure for computational linker modeling, the N-terminus of the HNH domain was initially extended using PyMOL by appending residues from SpCas9 (which connect the Rec1 and HNH domains and are not resolved in the SpCas9 crystal structure). The resulting structure was derived and prepared for linker modeling using a custom Python script that inserts HNH domain residues into possible insertion sites throughout the insertion loop, as shown in Table 1.
[0241] Table 1: Preliminary calculation and screening results of possible insertion sites in Cas12a.
[0242]
[0243] A quick computational screen was then performed to test the ability of the HNH domain end to connect to the ends of the linker cleavage site using the Rosetta Relax protocol included in the Rosetta macromolecular modeling software package (Huang P.S. et al. (2011) PLoS ONE 6(8):e24109. https: / / doi.org / 10.1371 / journal.pone.002410). Ten iterations of loop closure were performed for each insertion point (without sequence design or insertion). The number of times the linker was able to successfully connect across the ten iterations was tallied and compared (Table 1). Based on a combination of their loop closure success rate and manual inspection, two of these insertion sites were selected for more thorough linker modeling, including variations in linker length (shown in bold in Table 1).
[0244] For these two selected insertion sites, a refinement test was then performed using small-scale (2- to 4-residue) glycine-serine insertions or deletions in the N- and C-terminal linkers, and a more extensive sampling (100 iterations each) was performed. The residues that could be deleted were selected based on manual inspection of the sequence. Based on the linker modeling results, eight designs (four each for the insertion sites) were selected for experimental testing, including N-terminal linker extensions of 0, 2, or 4 residues, and C-terminal linker extensions of 0 or 2 residues.
[0245] Example 2:
[0246] The DNA coding regions of the 8 LbCas12a-HNH constructs (HNH-3287, HNH-3288, HNH-3289, HNH-3290, HNH-3296, HNH-3297, HNH-3298, and HNH-3299) were synthesized using solid state synthesis with a 6-Histidine tag. The coding regions were cloned into the pET28a plasmid (Novagen) behind an inducible T7 promoter and transfected into BL21(DE3)-Star cells (Invitrogen) and plated on kanamycin. Individual colonies were grown in 30 ml Luria Broth at 37°C until the A600 optical density reached 0.5. 500 mM IPTG was added and the temperature was reduced to 18°C for 18 hours of expression. Cells were pelleted and lysed using BugBuster Master Mix (Millipore) following the manufacturer’s instructions. Cell debris was pelleted and the soluble fraction was imaged on a 4-12% Bis-Tis PAGE gel (Invitrogen) under reducing conditions and visualized using Coomassie staining. All eight HNH constructs showed soluble protein expression at approximately 160 kDa MW (arrow). Figure 6
[0247] Soluble protein expression of all eight constructs containing HNH nucleases in the middle of the Cas12a protein demonstrated the quality of the fusion design. Large domain insertions in the middle of a protein often result in insoluble protein expression or no expression in E. coli. It was observed that all eight proteins were expressed at high levels, indicating that the chimeric proteins were correctly folded and did not cause disruption to the folding of either protein.
[0248] The expression protocol was repeated to generate protein suitable for nuclease assays. After eight constructs were pelleted, E. coli cells were frozen, thawed, and suspended in Buffer A (20 mM HEPES-KOH pH 7.5, 500 mM NaCl, 10% glycerol, 2 mM TCEP, and 10 mM imidazole, pH 7.5). 0.3 mg / ml lysozyme was added and cells were incubated at room temperature for 20 minutes, followed by sonication (QSonica) at 1 / 8 inch probe, power 25%, 10 seconds on, 30 seconds off for 2.25 minutes. Cell debris was pelleted and supernatant was loaded onto Ni-NTA agarose (Bio-Rad), washed with 20 mM imidazole in Buffer A, and eluted with 300 mM imidazole in Buffer A. The approximate concentration of protein was 0.5 mg / ml to 2 mg / ml (estimated by NanoDrop A280 absorbance), with a total eluate of 200 μL.
[0249] Example 3:
[0250] Plasmid-based assays were used to assess the nicking activity of purified HNH-3287, HNH-3288, HNH-3289, HNH-3290, HNH-3296, HNH-3297, and HNH-3298. The working principle of the plasmid nicking assay is that supercoiled plasmids extracted from bacteria run smaller on agarose gels than linearized, double-cut plasmids. Furthermore, if only one strand is nicked, the plasmid runs even larger than linearized plasmids. This assay has been widely used in the CRISPR field to assess whether an enzyme is a double-stranded nuclease or a single-stranded nuclease (Jinek et al., Science. 2012 Aug 17;337(6096):816-21) (Zetsche et al., Cell. 2015 Oct 22;163(3):759-71).
[0251] The sequence synthesized and cloned into the pUC18 plasmid was 5'-TTTAGGAAT CCCTTCTGC AGCACCTGG-3' (SEQ ID NO: 123), with the protospacer adjacent motif (PAM) in bold. The plasmid was expressed in DH5a cells and purified using a plasmid miniprep kit (Qiagen). A CRISPR RNA molecule was synthesized (Synthego) that contained no chemical modifications, with the sequence 5'-AAUUUCUACU AAGUGUAGAU GGAAUCCCUUCUGCAGCACC UGG-3' (SEQ ID NO: 124), with the portion complementary to the plasmid in bold. Thirty microliters of reaction was assembled in an RNA:protein:plasmid ratio of 10:10:1, incubated at 37°C for 15 minutes, heat inactivated at 85°C for 2 minutes, and loaded onto a 1% agarose gel containing 1 / 100 v / v SYBR-Safe dye (Invitrogen).
[0252] The proteins tested were wild-type LbCasl2a (wt LbCasl2a), LbCasl2a-Rl 138A, and various chimeric HNH proteins. Rl 138A is a point mutation in LbCasl2a that corresponds to a known non-template strand nickase mutation in AsCasl2a (Rl 226A) (Yamano T, et al. Cell. 2016 May 5;165(4):949-62). For wt LbCasl2a and LbCasl2a-Rl 138, the concentration tested was 33 nM. For the various HNH constructs, a lower 9 nM was used in order to distinguish the most active nucleases by approaching the expected Kd, rather than generating full cleavage.
[0253] The resulting gel Figure 7 ) shows that HNH-3287, HNH-3288, HNH-3289, HNH-3290, HNH-3296, HNH-3297, and HNH-3298 are all nickases, with a percentage of cleavage from about 25% efficiency to about 75% efficiency at low protein concentrations of 9 nM (with the upper band being the cleaved plasmid compared to the supercoiled plasmid in the lower band). Longer incubations or higher concentrations result in full cleavage, but this does not allow for comparison of relative mutant activity. Chimeric HNH-3298 exhibits the highest percentage of cleavage activity at 9 nM [protein] for 15 minutes at 37°C.
[0254] Example 4:
[0255] Methods
[0256] Protein expression and purification
[0257] In initial expression and activity testing, His-tagged proteins SYN3287 (SEQ ID NO: 125), SYN3288 (SEQ ID NO: 126), SYN3289 (SEQ ID NO: 127), SYN3290 (SEQ ID NO: 128), SYN3296 (SEQ ID NO: 129), SYN3297 (SEQ ID NO: 130), SYN3298 (SEQ ID NO: 131), and SYN3299 (SEQ ID NO: 132) were expressed in BL21 cells in 30 mL cultures. Each of the proteins included an active HNH domain and an inactivated RuvC domain. After cell pellets were frozen overnight, they were lysed by sonication. The proteins were then purified from the lysate using HisPur TM Ni-NTA spin columns.
[0258] For assays of SYN3298 and SYN3289, the proteins were expressed in the same manner with the following changes: the proteins were expressed in 1 L cultures and purified by FPLC using a HisTrap-HF column. Fractions containing the protein of interest were further purified by cation exchange and stored in 50% glycerol.
[0259] Plasmid nicking enzyme assay
[0260] To determine the activity of the purified proteins as nickases or nucleases, 30 μΐ, reactions were prepared containing lx NEBuffer 3.1, 100 femtomoles of DNA substrate, and an equal amount of purified protein and appropriate guide RNA (1 picomole each, unless otherwise noted). The reactions were incubated at 37°C for 30 minutes, stopped by proteinase K digestion at room temperature for 20 minutes, and separated on a 1% agarose gel. The sequence of the target site for the plasmid nickase assay is SEQ ID NO: 133.
[0261] Fluorescent nicking enzyme assay
[0262] DNA substrates were generated by annealing one labeled DNA strand (SEQ ID NO: 134) and one unlabeled DNA strand (SEQ ID NO: 135) to generate a Cy5-labeled substrate with the PAM-containing or non-PAM-containing strand in the respective position. The spacer for this assay comprises the sequence of SEQ ID NO: 150. Nicking reactions were prepared as described for the plasmid nickase assay and incubated at 37°C for 30 min. Reactions were stopped by digesting the samples with proteinase K for 10 min. Then, all samples were mixed with urea loading buffer to lx concentration and heated to 90°C for 5 min to denature the substrates. Samples were separated by running a 6% TBE urea gel at 100 V for 6 h at 4°C.
[0263] HEK293T cell transfection
[0264] Eukaryotic HEK293T (ATCC CRL-3216) cells were cultured in Dulbecco's Modified Eagle's Medium (ThermoFisher) supplemented with 10% (v / v) FBS (FBS) at 37°C with 5% CO2. Protein components were synthesized using gene synthesis and subsequently cloned into plasmids with a CMV promoter. Guide RNAs were cloned using the human U6 promoter. HEK293T cells were seeded on 48-well collagen-coated BioCoat plates (Corning). Cells were transfected at approximately 70% confluency. Each well was transfected with 375 ng of CRISPR plasmid and 125 ng of guide RNA expression plasmid using 1.5 μΐ Lipofectamine 3000 (ThermoFisher Scientific) following the manufacturer's protocol. Genomic DNA of transfected cells was obtained 3 days later and indels were detected and quantified using high-throughput Illumina amplicon sequencing.
[0265] To determine which strand a designed protein preferentially nicks, a pair of guides was designed such that the Cas9 guide and the guide for the designed protein are on the same strand and will be cleaved when in close proximity to each other (within about 10 bp). Each design tested was paired with one of the following: a nuclease-dead SpCas9, SpCas9 D10A target strand nickase, or SpCas9 H840A non-target strand nickase. If the synthetic nickase and its paired Cas9 nickase cleave opposite strands, their editing frequency is expected to be higher than if they cleave the same strand due to the generation of a double-strand break.
[0266] Results
[0267] Designed synthetic nickases with His tags were successfully expressed in BL21 E. coli.
[0268] As described above, following crude purification of the designed nickases, all samples showed a band at the expected size (approximately 160 kDa), as shown in Figure 8 , indicating that these nickases were expressed soluble in E. coli.
[0269] Initial plasmid nicking activity observed from crude purified synthetic nickases.
[0270] Plasmid nickase experiments were performed as described in the Methods section above, using crude purified nickases as shown in Figure 8 . Due to low yield of partially purified material, all designed nickases were tested at low concentrations so that they could be directly compared. As shown in Figure 9 , all but one of the designs showed a band, indicating that the plasmid was nicked, which was not evident in the negative control sample, indicating that these designs were able to nick the DNA substrate.
[0271] RNA dependence of plasmid nicking using crude purified synthetic nickases.
[0272] To ensure that the observed plasmid nicking and cleavage was guide-dependent, rather than random nuclease activity, the plasmid nickase assay was repeated for selected designs in the presence of the target crRNA. Designs SYN3288, SYN3296, and SYN3298 all showed a reduction in the amount of uncleaved plasmid in the presence of the crRNA, as shown in Figure 10 , indicating that their nuclease activity was RNA-dependent.
[0273] Plasmid nicking activity of purified synthetic nickases SYN3298.
[0274] Different amounts of protein + guide were tested relative to the concentration of LbCas12a control used (e.g., 30x means that 30 picomoles of protein and guide were included in the reaction). Nicking of the plasmid was observed at all concentrations of SYN3298 tested, and to a lesser extent cleavage Figure 11 , confirming that this design functions as a DNA nickase.
[0275] Fluorescent nicking enzyme experiments using purified synthetic nickases SYN3298 and SYN3289.
[0276] Substrates with fluorescent Cy5 labels on the target strand Figure 12 or non-target strand Figure 13 were incubated with designed nickases that included an active HNH domain and an inactivated RuvC domain, LbCas12a, or LbCas12a R1138A mutant, a non-target strand nickase, and resolved on denaturing TBE-urea gels. Movement of the labeled band position indicates that this strand was cleaved.Figure 14 A portion of the gel showing the sample incubated with the labeled target strand, Figure 15 A portion of the gel showing the sample incubated with the labeled non-target strand, and Figure 16 The full gel, including controls, sample incubated with the labeled target strand (box lane labeled with “a)”), and sample incubated with the labeled non-target strand (box lane labeled with “b)”). SYN3298 shows a band at the expected location of the substrate cleaved at the target DNA strand, but not on the non-target DNA strand, indicating that it functions as a target strand nickase.
[0277] Sequence-based strand-specific nickase genomic DNA assay in HEK293T cells.
[0278] Synthetic nickases were co-transfected with an adjacent Cas9 nickase (e.g., Cas9(H840A) or Cas9(D10A), which cleave the target strand (e.g., Cas9(D10A)) or non-target strand (e.g., Cas9(H840A)), respectively). Information on the spacers used in the sequence-based strand-specific nickase assay is provided in Table 2. “Upstream guide” refers to a spacer that is predicted to be cleaved closer to the 5’ end of the PAM-containing DNA strand. The estimated distance between cleavage sites was determined based on the predicted cleavage site of each native nuclease domain.
[0279] Table 2: Spacer information for sequence-based strand-specific nickase assay.
[0280]
[0281] For each pair of enzymes, the editing efficiency was normalized relative to the level of indel observed when the Cas9 nickase was paired with a nuclease-dead LbCas12a at the same target site Figure 17 The numbers in parentheses in the Figure 17 indicate the observed editing efficiency before normalization. If the synthetic enzyme (SYN) (i.e., SYN3289, SYN3290, or SYN3298) preferentially cleaves the target strand, then (H480A::SYN) / (D10A::SYN) > 1. If the synthetic enzyme (SYN) (i.e., SYN3289, SYN3290, or SYN3298) preferentially cleaves the non-target strand, then (H480A::SYN) / (D10A::SYN) < 1.
[0282] To determine which strand the designed protein preferentially nicked, a pair of guides were designed such that the Cas9 guide and the guide for the designed protein were on the same strand and would be cut when in close proximity to each other (within about 10 bp). Each design tested was paired with one of the following: nuclease dead SpCas9, SpCas9 D10A target strand nickase, or SpCas9 H840A non-target strand nickase. If the synthetic nickase cut the opposite strand from its paired Cas9 nickase, a higher editing frequency was expected to be observed due to the generation of a double-strand break than if the same strand was cut. When all of the designed nickases were paired with the Cas9 non-target strand nickase, an increase in indel frequency was consistently observed (about a 3-fold increase) compared to when paired with the Cas9 target strand nickase, indicating that the designed nickases preferentially cut the target DNA strand.
[0283] Example 5:
[0284] Cytidine base editing data was obtained for the base editors combining A3A cytosine deaminase (SEQ ID NO: 152) with SYN3289, SYN3290, or SYN3298. Figures 18-21 ) Three configurations were tested for each enzyme: A3A fused to the N-terminus of the synthetic enzyme using a linker (SEQ ID NO: 22) and UGI (SEQ ID NO: 104) fused to the C-terminus of the synthetic enzyme using a linker of SEQ ID NO: 45, providing SEQ ID NOs: 160-162; A3A fused to the N-terminus of the synthetic enzyme using a previously published linker (SEQ ID NO: 153; Li et al. Nature Biotechnology 36, 324-327 (2018)) and UGI (SEQ ID NO: 104) fused to the C-terminus of the synthetic enzyme using a linker of SEQ ID NO: 154, providing SEQ ID NOs: 163-165; or Suntag-based UGI recruitment (SEQ ID NO: 104) fused to the C-terminus of A3A (SEQ ID NO: 152), providing SEQ ID NO: 156, recruited to one of SEQ ID NOs: 157-159 with a peptide tag. Shown in Figures 18-21 All percentages represent the average of three data points. All enzymes tested exhibited cytidine base editing in all three configurations tested. Figure 18 The spacer of SYN3289 is SEQ ID NO: 144, Figure 19 The spacer of SYN3290 is SEQ ID NO: 145, Figure 20 The spacer of SYN3291 is SEQ ID NO: 146, and Figure 21 The spacer of SYN3292 is SEQ ID NO: 147.
[0285] Example 6:
[0286] Synthetic enzymes SYN3289, SYN3290, and SYN3298 were obtained as N-terminal fusions to TadA8e adenine deaminase for adenine base editing data Figures 22-23 ). The synthetic enzymes were fused to TadA8e (SEQ ID NO: 155) using a linker (SEQ ID NO: 47) to provide SEQ ID NOs: 166-168. Shown in Figures 22-23 All percentages represent the average of three data points. All three designs tested exhibited adenine base editing activity when fused to TadA8e. Figure 22 The spacer of SYN3298 is SEQ ID NO: 148, and Figure 23 The spacer of SYN3299 is SEQ ID NO: 149.
[0287] Example 7:
[0288] PAM specificity of the engineered proteins described herein will be tested using the PAM-SCANR assay (Leenay, R. T. et al. Molecular Cell 62, 137-147 (2016)). Briefly, an inactivated version of the synthetic enzyme (e.g., SYN3298 with a H370A mutation (i.e., SYN3298 with an alanine at amino acid residue number 370, providing an inactivation mutation in the HNH domain)) will be expressed in a bacterial library containing all 256 possible 4-base 5’ PAMs located in the lacI gene promoter, and a GFP gene controlled by the lacZ promoter. In cells where the synthetic enzyme binds to a PAM, the lacI protein will not be expressed, allowing GFP expression, and the cells will fluoresce. Fluorescent cells will then be isolated by cell sorting, and sequenced to determine which PAMs the synthetic enzyme is able to bind.
[0289] Example 8:
[0290] The PAM specificity of synthetic enzymes SYN3287, SYN3288, and SYN3298, and mutants thereof, was tested using natural TTTV PAM sequences (where V is A, G, or C) and / or non-natural PAM sequences (e.g., CCCC, TCCA, TCCC, TCCG, TTCA, TTCC, TTCG, TATA, TATC, or TATG). Specifically, SYN3287, SYN3288, or SYN3298 with the following were tested: G532R and K595R (RR) mutations, reference to position numbering of SEQ ID NO: 180 (LbCasl2a); G532R, K538V, and Y542R (RVR) mutations, reference to position numbering of SEQ ID NO: 180 (LbCasl2a); or a D156R mutation, and indel generation frequencies were determined. RR-LbCasl2a has G532R and K595R mutations (SEQ ID NO: 181); 3298RR has G680R and K743R mutations (SEQ ID NO: 184); 3287RR has G676R and K739R mutations (SEQ ID NO: 187); 3288RR has G678R and K741R mutations (SEQ ID NO: 190); RVR-LbCasl2a has G532R, K538V, and Y542R mutations (SEQ ID NO: 182); 3298RVR has G680R, K686V, and Y690R mutations (SEQ ID NO: 185); 3287RVR has G676R, K682V, and Y686R mutations (SEQ ID NO: 188); 3288RVR has G678R, K684V, and Y688R mutations (SEQ ID NO: 191); LbCasl2a D156R has a D156R mutation (SEQ ID NO: 183); 3298D156R has a D156R mutation (SEQ ID NO: 186); 3287D156R has a D156R mutation (SEQ ID NO: 189); and 3288D156R has a D156R mutation (SEQ ID NO: 192), each reference to position numbering of the sequence itself.
[0291] HEK293-T cell testing
[0292] HEK293T cells were seeded into 48-well collagen-coated plates (Corning) using DMEM media without antibiotics. When cells reached 70-80% confluency, transfections were performed using 1.5 pL Lipofectamine 3000 (Thermo Fisher Scientific), 750 ng editor plasmid, and 250 ng guide RNA plasmid according to the manufacturer’s protocol. Three days later, cells were lysed and DNA was extracted using MagMax DNA extraction kit (Applied Biosystems). Samples were sequenced by NGS and editing efficiency was calculated as the average percentage of reads containing the desired edit across all available technical replicates.
[0293] Plant testing
[0294] For each PAM sequence to be tested, five spacers were designed in corn and soybean, each targeting a unique site. Plasmids expressing the editor and the spacers to be tested were transformed into corn or soybean dry desiccated excised embryos (DEE) by standard Agrobacterium transformation protocol. Ten weeks later, destructive sampling was performed on the small plants and all editing targets were sequenced using NGS. For each target, editing efficiency was calculated as the percentage of analyzed samples containing the desired edit, including at least 1% or 10% of the sequenced reads.
[0295] Results
[0296] Results are provided in Figures 24-33 and Tables 3-10. Except for SYN3288, all tested proteins appeared to be as efficient as the LbCas12a and D156R control against all eight tested TTV spacers. The D156R mutant did not significantly alter PAM specificity and did not negatively impact recognition of TTV PAM sequences Figures 28-29 .
[0297] In soybean and corn plants, RR and RVR mutations of SYN3298 (SYN3298-RR and SYN3298-RVR) showed binding affinity to natural TTV PAM sequences and non-natural PAM sequences (Tables 3-10). SYN3298 (SEQ ID NO: 131), SYN3298-RR (SEQ ID NO: 184), and SYN3298-RVR (SEQ ID NO: 185).
[0298] Table 3. Total counts of desired edits for five different spacers of SYN3298, SYN3298-RR, and SYN3298-RVR tested in soybean.
[0299]
[0300] Table 4. Percentage of samples with >1% edits in each of the five different spacers used to test SYN3298, SYN3298-RR, and SYN3298-RVR in soybean. “-” indicates no data returned for a given spacer.
[0301]
[0302]
[0303] Table 5. Percentage of samples with >10% edits in each of the five different spacers used to test SYN3298, SYN3298-RR, and SYN3298-RVR in soybean. “-” indicates no data returned for a given spacer.
[0304]
[0305] Table 6. Overall PAM specificity of SYN3298, SYN3298-RR, and SYN3298-RVR in soybean.
[0306]
[0307]
[0308] Table 7. Expected edit total counts for five different spacers tested in corn for SYN3298, SYN3298-RR, and SYN3298-RVR.
[0309]
[0310] Table 8. Percentage of samples with >1% edits in each of the five different spacers used to test SYN3298, SYN3298-RR, and SYN3298-RVR in corn. “-” indicates no data returned for a given spacer.
[0311]
[0312]
[0313] Table 9. Percentage of samples with >10% edits in each of the five different spacers used to test SYN3298, SYN3298-RR, and SYN3298-RVR in corn. “-” indicates no data returned for a given spacer.
[0314]
[0315] Table 10. PAM specificity of SYN3298, SYN3298-RR, and SYN3298-RVR in corn.
[0316]
[0317]
[0318] The foregoing is illustrative of the present application, and is not to be construed as limiting thereof. The application is defined by the following claims, with equivalents of the claims to be included therein.
Claims
1. An engineered protein comprising: The first polypeptide, which is a first portion of a modified protein, wherein the modified protein comprises an amino acid sequence having at least 80%, 85%, 90%, or 95% identity with the amino acid sequence of SEQ ID NO:180 (LbCas12a), and, relative to the amino acid sequence of SEQ ID NO:180, the modified protein comprises mutations at one or more positions selected from the group consisting of: [Refer to SEQ ID NO:180] The position numbers N100, K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, N211, N263, T296, E330, K387, A404, D405, D423, E484, L498, N527, Q529, G532, D535, K538, E539, D541, Y542, Y553, Y554, D572, L585, K591, M592, K595, V596, S599, K600, K601, Y616, Y646, W649 and any combination thereof; and The second polypeptide is heterologous to the first polypeptide and is not a type V CRISPR-Cas effector polypeptide. The first polypeptide and the second polypeptide are different from each other, and The engineered protein described therein contains the mutation.
2. The engineered protein of claim 1, wherein the modified protein comprises one or more amino acid mutations selected from the group consisting of: (Ref. SEQ ID) The positions numbered NO:180 are N100S, K116D, K116R, K116N, K120R, K120H, K120N, K120T, K120Y, K120Q, K121S, K121T, K121H, K121R, K121G, K121D, K121Q, D122R, D122K, D122H, D122E, D122N, E125G, E125R, E125K, E125Q, E125Y, T148H, T148S, T148A, T148C, T149A, T149C, T149S, T149G, T149H, and T149P. T149F, T149N, T149D, T149V, T152R, T152K, T152W, T152Y, T152H, T152 Q, T152E, T152L, T152F, D156R, D156K, D156Y, D156W, D156Q, D156H, D15 6I, D156V, D156L, D156E, E159K, E159R, E159H, E159Y, E159Q, N211S, N 263I, T296I, E330V, K387E, A404V, D405G, D423V, E484D, L498M, N527S, Q529N, Q529T, Q529H, Q529A, Q529F, Q529G, Q529S, Q529P, Q529W, Q529 D. G532D, G532N, G532S, G532H, G532F, G532K, G532R, G532Q, G532A, G53 2L, G532C, D535N, D535H, D535V, D535T, D535S, D535A, D535W, D535K, K 538R, K538V, K538Q, K538W, K538Y, K538F, K538H, K538L, K538M, K538C, K538G, K538A, K538P, E539V, D541N, D541H, D541R, D541K, D541Y, D541 I, D541A, D541S, D541E, Y542R, Y542K, Y542H, Y542Q, Y542F, Y542L, Y54 2M, Y542P, Y542V, Y542N, Y542T, Y553H, Y554N, D572G, L585Q, L585G, L 585H, L585F, K591W, K591F, K591Y, K591H, K591R, K591S, K591A, K591G,K591P, M592R, M592K, M592Q, M592E, M592A, K595R, K595Q, K595Y, K595L, K595W, K595H, K595E, K595S, K595 D. K595M, V596T, V596H, V596G, V596A, S599G, S599H, S599N, S599D, K600R, K600H, K600G, K601R, K601H, K60 1Q, K601T, Y616K, Y616R, Y616E, Y616F, Y616H, Y646R, Y646E, Y646K, Y646H, Y646Q, Y646W, Y646N, W649H, W649K, W649Y, W649R, W649E, W649S, W649V, W649T, and any combination thereof, optionally wherein said modified protein comprises one or more amino acid mutations selected from the group consisting of: (Refer to SEQ) Location numbers with ID NO:180 include N100S, K116D, E125G, T152R, D156E, N211S, N263I, T296I, E330V, K387E, A404V, D405G, D423V, E484D, L498M, N527S, G532R, K538V, E539V, Y542R, Y553H, Y554N, D572G, L585Q, K595R, K595Y, and any combinations thereof.
3. The engineered protein according to claim 1 or 2, wherein the engineered protein has altered PAM specificity compared to PAM (protospacer adjacent motif) specificity for LbCas12a.
4. The engineered protein according to any one of claims 1 to 3, further comprising a mutation (e.g., dead LbCas12a, dLbCas12a) located at the nuclease active site (e.g., RuvC domain), optionally wherein the modified protein comprises an amino acid sequence having at least 80%, 85%, 90% or 95% identity with the amino acid sequence of SEQ ID NO:
50.
5. The engineered protein according to any one of claims 1 to 4, wherein the first polypeptide comprises amino acid residue 1 to about amino acid residue 250, 300 or 350 with reference to the position number of the modified protein, optionally wherein the first polypeptide comprises the mutation, and / or the first polypeptide comprises amino acid residue 1 to about amino acid residue 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292 or 293 with reference to the position number of the modified protein.
6. The engineered protein according to any one of claims 1 to 5, wherein the length of the second polypeptide is about 10, 50, 100 or 150 to about 200, 250 or 300 amino acids, and / or the length of the first polypeptide is about 100, 200 or 250 to about 300, 350 or 400 amino acids, optionally wherein the length of the second polypeptide is about 140 or 150 to about 160, 170, 180, 190, 200, 210, 220, 230, 240 or 250 amino acids, and / or the length of the first polypeptide is about 250 or 275 to about 300 or 350 amino acids.
7. The engineered protein according to any one of claims 1 to 6, wherein the second polypeptide comprises a first nuclease domain or a portion thereof.
8. The engineered protein according to any one of claims 1 to 7, wherein the second polypeptide comprises a target chain cleavage enzyme domain or a portion thereof, optionally wherein the second polypeptide comprises a target chain-specific cleavage enzyme domain, a non-target chain-specific cleavage enzyme domain, or a target chain and non-target chain cleavage enzyme domain.
9. The engineered protein according to any one of claims 1 to 8, further comprising a third polypeptide, the third polypeptide comprising a second nuclease domain or a portion thereof, optionally wherein the third polypeptide comprises the mutation, and / or wherein the first polypeptide and the third polypeptide are discontinuous (i.e., separated from each other (optionally separated by at least 10, 50, 100 or more amino acids) and not directly connected to each other).
10. The engineered protein of claim 9, wherein the second polypeptide and the third polypeptide are heterologous.
11. The engineered protein of claim 9 or 10, wherein the third polypeptide is a type V CRISPR-Cas effector polypeptide, optionally wherein the third polypeptide is a second part of the modified protein and is different from the first polypeptide.
12. The engineered protein according to any one of claims 9 to 11, wherein the second nuclease domain or a portion thereof is a non-target strand and a target strand cleavage enzyme domain or a portion thereof.
13. The engineered protein according to any one of claims 9 to 12, wherein the second nuclease domain is active.
14. The engineered protein according to any one of claims 9 to 12, wherein the second nuclease domain is inactive.
15. The engineered protein according to any one of claims 1 to 14, wherein the second polypeptide comprises an HNH domain, optionally wherein the HNH domain comprises a mutation that modifies the activity of the HNH domain (e.g., an H840A mutation).
16. The engineered protein according to any one of claims 9 to 15, wherein the first polypeptide and the third polypeptide are each part of the modified protein, and the second polypeptide is located between and / or linked to (e.g., directly or indirectly) two amino acids of the modified protein, the two amino acids being two consecutive or discontinuous amino acids.
17. The engineered protein of claim 16, wherein the second polypeptide is located in the engineered protein at a position corresponding to the interdomain linker region of the modified protein.
18. The engineered protein according to any one of claims 1 to 17, wherein the second polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NO:1 or 169-174, optionally wherein the second polypeptide comprises an amino acid sequence of any one of SEQ ID NO:1 or 169-174.
19. The engineered protein according to any one of claims 9 to 18, wherein the engineered protein comprises the first polypeptide, the second polypeptide, and the third polypeptide in the direction from the amino terminus to the carboxyl terminus, optionally wherein the length of the third polypeptide is about 800 or 850 to about 900, 1,000, or 1,100 amino acids (e.g., about 900 to about 950 or 1,000 amino acids).
20. The engineered protein according to any one of claims 1 to 19, further comprising all or a portion of the wedge domain, Rec1 domain, Rec2 domain, PAM interaction domain, RuvC domain, bridging helix and / or Nuc domain of the modified protein, optionally wherein the engineered protein comprises all or a portion of the wedge domain, Rec1 domain, Rec2 domain, PAM interaction domain, RuvC domain, bridging helix and / or Nuc domain of LbCas12a.
21. The engineered protein of claim 20, wherein the engineered protein comprises the Rec1 domain and the Rec2 domain, and the second polypeptide is located between the Rec1 domain and the Rec2 domain.
22. The engineered protein according to any one of claims 1 to 21, wherein the engineered protein lacks at least a portion of the modified protein, optionally wherein the engineered protein lacks at least a portion of LbCas12a.
23. The engineered protein according to any one of claims 9 to 22, further comprising a first linker between the first polypeptide and the second polypeptide and / or a second linker between the second polypeptide and the third polypeptide, optionally wherein the first linker and / or the second linker comprises (GGS). n The amino acid sequence, wherein n is an integer from 1 to 20; GS; SG; and / or the amino acid sequence of one of SEQ ID NO: 18-47 or 176-179.
24. The engineered protein of claim 23, wherein the first linker and / or the second linker comprises 1 to 10 amino acids, optionally wherein the first linker and / or the second linker comprises 1, 2, 3 or 4 amino acids.
25. The engineered protein of claim 23 or 24, wherein the first linker and / or the second linker comprises glycine and / or serine.
26. The engineered protein according to any one of claims 1 to 25, wherein the engineered protein comprises an amino acid sequence having about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with the amino acid sequence of a wild-type CRISPR-Cas effector protein, optionally wherein the engineered protein comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NO: 50-66, 151 or 180.
27. The engineered protein according to any one of claims 1 to 26, wherein, when the modified protein is optimally aligned with SEQ ID NO: 50, 58, or 180, the second polypeptide is located between one set of the following amino acid residue groups present in the engineered protein: Between amino acid residues 290 and 291 of SEQ ID NO:50 Between amino acid residues 291 and 292 of SEQ ID NO:50 Between amino acid residues 291 and 292 of SEQ ID NO:58, or Between amino acid residues 292 and 293 of SEQ ID NO:58 The amino acid residues between 290 and 291 of SEQ ID NO:180 Between amino acid residues 291 and 292 of SEQ ID NO:180, or Between amino acid residues 292 and 293 of SEQ ID NO:
180.
28. The engineered protein according to any one of claims 1 to 27, wherein the engineered protein is a nuclease, optionally wherein the engineered protein is a target strand cleavage enzyme, a non-target strand cleavage enzyme, or a target strand and non-target strand cleavage enzyme.
29. The engineered protein according to any one of claims 1 to 28, wherein the engineered protein is more efficient in cleaving the target strand and / or non-target strand of the target nucleic acid compared to a CRISPR-Cas effector protein (e.g., a wild-type CRISPR-Cas effector protein and / or a protein having the sequence of one of SEQ ID NO: 50-66 or 151 or 180).
30. The engineered protein according to any one of claims 1 to 29, wherein the engineered protein recognizes a non-natural protospacer adjacent motif (PAM) site and / or sequence (e.g., the engineered protein contains altered PAM specificity compared to wild-type LbCas12a).
31. An engineered protein comprising: Amino acid sequences having at least 80%, 85%, 90%, or 95% identity with the amino acid sequence of SEQ ID NO:131 (SYN3298); and The mutation relative to the amino acid sequence of SEQ ID NO:131 (SYN3298) is located at one or more positions selected from the group consisting of: See SEQ ID The position numbers of NO:131 are N100, K116, K120, K121, D122, E125, T148, T149, T152, D156, E159, N211, N263, T444, E478, K535, A552, D553, D571, E632, L646, N675, Q677, G680, D683, K686, E687, D689, Y690, Y701, Y702, D720, L733, K739, M740, K743, V744, S747, K748, K749, Y764, Y794, W797, and any combination thereof.
32. The engineered protein of claim 31, wherein the mutation is one or more amino acid mutations selected from the group consisting of: (Ref. SEQ ID) The positions numbered NO:131 are N100S, K116D, K116R, K116N, K120R, K120H, K120N, K120T, K120Y, K120Q, K121S, K121T, K121H, K121R, K121G, K121D, K121Q, D122R, D122K, D122H, D122E, D122N, E125G, E125R, E125K, E125Q, E125Y, T148H, T148S, T148A, T148C, T149A, T149C, T149S, T149G, T149H, and T149P. T149F, T149N, T149D, T149V, T152R, T152K, T152W, T152Y, T152H, T152 Q, T152E, T152L, T152F, D156R, D156K, D156Y, D156W, D156Q, D156H, D15 6I, D156V, D156L, D156E, E159K, E159R, E159H, E159Y, E159Q, N211S, N 263I, T444I, E478V, K535E, A552V, D553G, D571V, E632D, L646M, N675S, Q677N, Q677T, Q677H, Q677A, Q677F, Q677G, Q677S, Q677P, Q677W, Q677 D. G680D, G680N, G680S, G680H, G680F, G680K, G680R, G680Q, G680A, G68 0L, G680C, D683N, D683H, D683V, D683T, D683S, D683A, D683W, D683K, K 686R, K686V, K686Q, K686W, K686Y, K686F, K686H, K686L, K686M, K686C, K686G, K686A, K686P, E687V, D689N, D689H, D689R, D689K, D689Y, D689 I, D689A, D689S, D689E, Y690R, Y690K, Y690H, Y690Q, Y690F, Y690L, Y69 0M, Y690P, Y690V, Y690N, Y690T, Y701H, Y702N, D720G, L733Q, L733G, L 733H, L733F, K739W, K739F, K739Y, K739H, K739R, K739S, K739A, K739G,K739P, M740R, M740K, M740Q, M740E, M740A, K743R, K743Q, K743Y, K743L, K743W, K743H, K743E, K743S, K74 3D, K743M, V744T, V744H, V744G, V744A, S747G, S747H, S747N, S747D, K748R, K748H, K748G, K749R, K749H, K749Q, K749T, Y764K, Y764R, Y764E, Y764F, Y764H, Y794R, Y794E, Y794K, Y794H, Y794Q, Y794W, Y794N, W797H, W797K, W797Y, W797R, W797E, W797S, W797V, W797T, and any combination thereof, optionally wherein said mutation is one or more amino acid mutations selected from the group consisting of: [Refer to SEQ] Location ID NO:131 includes the following: N100S, K116D, E125G, T152R, D156E, N211S, N263I, T444I, E478V, K535E, A552V, D553G, D571V, E632D, L646M, N675S, G680R, K686V, E687V, Y690R, Y701H, Y702N, D720G, L733Q, K743R, K743Y, and any combination thereof.
33. The engineered protein according to claim 31 or 32, wherein the engineered protein has altered PAM specificity compared to PAM (protospacer adjacent motif) specificity for LbCas12a.
34. The engineered protein according to any one of claims 31 to 33, wherein the mutation is located in the nuclease active site (e.g., RuvC domain) (e.g., dead LbCas12a, dLbCas12a), optionally wherein the engineered protein comprises an amino acid sequence having at least 80%, 85%, 90%, or 95% identity with the amino acid sequence of SEQ ID NO:
50.
35. The engineered protein according to any one of claims 31 to 34, wherein the engineered protein comprises a first polypeptide comprising amino acid residue 1 at position numbered according to the amino acid sequence of SEQ ID NO:180 (LbCas12a) to about amino residues 250, 300 or 350 and optionally the mutations described herein, wherein optionally the first polypeptide comprises amino acid residue 1 at position numbered according to the amino acid sequence of SEQ ID NO:180 (LbCas12a) to about amino residues 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292 or 293 and optionally the mutations described herein.
36. The engineered protein of claim 35, wherein the engineered protein comprises a second polypeptide, the second polypeptide being heterologous to the first polypeptide and not a type V CRISPR-Cas effector polypeptide, and the second polypeptide having a length of about 10, 50, 100, or 150 to about 200, 250, or 300 amino acids, and / or the first polypeptide having a length of about 100, 200, or 250 to about 300, 350, or 400 amino acids, optionally wherein the second polypeptide has a length of about 140 or 150 to about 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 amino acids, and / or the first polypeptide has a length of about 250 or 275 to about 300 or 350 amino acids.
37. The engineered protein of claim 36, wherein the second polypeptide comprises a first nuclease domain or a portion thereof.
38. The engineered protein according to any one of claims 36 to 37, wherein the second polypeptide comprises a target chain cleavage enzyme domain or a portion thereof, optionally wherein the second polypeptide comprises a target chain-specific cleavage enzyme domain, a non-target chain-specific cleavage enzyme domain, or a target chain and non-target chain cleavage enzyme domain.
39. The engineered protein according to any one of claims 35 to 38, further comprising a third polypeptide, the third polypeptide comprising a second nuclease domain or a portion thereof, optionally wherein the third polypeptide comprises the mutation, and / or the first polypeptide and the third polypeptide are discontinuous (i.e., separated from each other (optionally separated by at least 10, 50, 100 or more amino acids) and not directly connected to each other).
40. The engineered protein of claim 39, wherein the second polypeptide and the third polypeptide are heterologous.
41. The engineered protein of claim 39 or 40, wherein the third polypeptide is a type V CRISPR-Cas effector polypeptide, and optionally wherein the third polypeptide comprises the mutation and / or a portion of the amino acid sequence of SEQ ID NO:180 (LbCas12a).
42. The engineered protein according to any one of claims 39 to 41, wherein the second nuclease domain or a portion thereof is a non-target strand and a target strand cleavage enzyme domain or a portion thereof.
43. The engineered protein according to any one of claims 39 to 42, wherein the second nuclease domain is active.
44. The engineered protein according to any one of claims 39 to 42, wherein the second nuclease domain is inactive.
45. The engineered protein according to any one of claims 36 to 44, wherein the second polypeptide comprises an HNH domain, optionally wherein the HNH domain comprises a mutation that alters the activity of the HNH domain (e.g., an H840A mutation).
46. The engineered protein according to any one of claims 39 to 45, wherein the first polypeptide and the third polypeptide are each part of a type V CRISPR-Cas effector protein, and the second polypeptide is located between and / or linked to (e.g., directly or indirectly) two amino acids of the type V CRISPR-Cas effector protein, the two amino acids being two consecutive or discontinuous amino acids.
47. The engineered protein of claim 46, wherein the second polypeptide is located in the engineered protein at a position corresponding to the interdomain linker region of the V-type CRISPR-Cas effector protein.
48. The engineered protein according to any one of claims 36 to 47, wherein the second polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with one or more of SEQ ID NO:1 or 169-174, optionally wherein the second polypeptide comprises an amino acid sequence of any one of SEQ ID NO:1 or 169-174.
49. The engineered protein according to any one of claims 39 to 48, wherein the engineered protein comprises the first polypeptide, the second polypeptide, and the third polypeptide in the direction from the amino terminus to the carboxyl terminus, optionally wherein the length of the third polypeptide is about 800 or 850 to about 900, 1,000, or 1,100 amino acids (e.g., about 900 to about 950 or 1,000 amino acids).
50. The engineered protein according to any one of claims 31 to 49, further comprising all or a portion of the wedge domain, Rec1 domain, Rec2 domain, PAM interaction domain, RuvC domain, bridging helix, and / or Nuc domain of the type V CRISPR-Cas effector protein, optionally wherein the engineered protein comprises all or a portion of the wedge domain, Rec1 domain, Rec2 domain, PAM interaction domain, RuvC domain, bridging helix, and / or Nuc domain of LbCas12a.
51. The engineered protein of claim 50, wherein the engineered protein comprises the Rec1 domain and the Rec2 domain, and the second polypeptide is located between the Rec1 domain and the Rec2 domain.
52. The engineered protein according to any one of claims 31 to 51, wherein the engineered protein lacks at least a portion of the type V CRISPR-Cas effector protein, optionally wherein the engineered protein lacks at least a portion of LbCas12a.
53. The engineered protein according to any one of claims 39 to 52, further comprising a first linker between the first polypeptide and the second polypeptide and / or a second linker between the second polypeptide and the third polypeptide, optionally wherein the first linker and / or the second linker comprises (GGS). n The amino acid sequence, wherein n is an integer from 1 to 20; GS; SG; and / or the amino acid sequence of one of SEQ ID NO: 18-47 or 176-179.
54. The engineered protein of claim 53, wherein the first linker and / or the second linker comprises 1 to 10 amino acids, optionally wherein the first linker and / or the second linker comprises 1, 2, 3 or 4 amino acids.
55. The engineered protein according to claim 53 or 54, wherein the first linker and / or the second linker comprises glycine and / or serine.
56. The engineered protein according to any one of claims 31 to 55, wherein the engineered protein comprises an amino acid sequence having about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with the amino acid sequence of a wild-type CRISPR-Cas effector protein, optionally wherein the engineered protein comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with any one of SEQ ID NO: 50-66, 151 or 180.
57. The engineered protein according to any one of claims 31 to 56, wherein the engineered protein is a nuclease.
58. The engineered protein according to any one of claims 31 to 57, wherein the engineered protein is a target chain cleavage enzyme, a non-target chain cleavage enzyme, or a target chain and non-target chain cleavage enzyme.
59. The engineered protein according to any one of claims 31 to 58, wherein the engineered protein is more efficient in cleaving the target strand and / or non-target strand of the target nucleic acid compared to a CRISPR-Cas effector protein (e.g., a wild-type CRISPR-Cas effector protein and / or a protein having the sequence of one of SEQ ID NO: 50-66, 151 or 180).
60. The engineered protein according to any one of claims 31 to 59, wherein the engineered protein recognizes non-natural PAM sites and / or sequences (e.g., the engineered protein contains altered PAM specificity compared to wild-type LbCas12a).
61. The engineered protein according to any one of claims 1 to 60, wherein the engineered protein comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with any of SEQ ID NO:181-192, optionally wherein the engineered protein comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with any of SEQ ID NO:184-192.
62. A complex comprising: The engineered protein according to any one of claims 1 to 61; Guide nucleic acids (e.g., guide RNA); and Optional deaminase.
63. A nucleic acid molecule encoding an engineered protein according to any one of claims 1 to 61.
64. An expression cassette or vector comprising a nucleic acid molecule according to claim 63 or a nucleotide sequence encoding an engineered protein according to any one of claims 1 to 61.
65. A method for modifying a target nucleic acid, the method comprising: Contact the target nucleic acid with the following: The engineered protein according to any one of claims 1 to 61, and Guide nucleic acids (e.g., guide RNA), Optionally, the engineered protein and the guide nucleic acid form a complex or are contained in the complex, thereby modifying the target nucleic acid.
66. The method of claim 65, wherein the method is more efficient in modifying the target nucleic acid and / or cleaving the target strand and / or non-target strand of the target nucleic acid compared to a control method (e.g., a method comprising contacting the target nucleic acid with a wild-type CRISPR-Cas effector protein).
67. A method for improving the efficiency of modifying target nucleic acids, the method comprising: Contact the target nucleic acid with the following: The engineered protein according to any one of claims 1 to 61, and Guide nucleic acids (e.g., guide RNA), Optionally, the engineered protein and the guide nucleic acid form a complex or are contained in the complex, thereby modifying the target nucleic acid and thus improving the efficiency of modifying the target nucleic acid compared to a control method (e.g., a method comprising contacting the target nucleic acid with a wild-type CRISPR-Cas effector protein and lacking the engineered protein).
68. The method according to claims 65 to 67, wherein the target nucleic acid is present in eukaryotic cells, optionally wherein the target nucleic acid is present in plant cells.
69. The method according to any one of claims 65 to 69, wherein the engineered protein provides a different editing profile of the target nucleic acid compared to the editing profile of the target nucleic acid provided by the wild-type CRISPR-Cas effector protein.
70. The method according to any one of claims 65 to 69, wherein the engineered protein provides a different cleavage pattern of the target nucleic acid compared to the cleavage pattern provided by the wild-type CRISPR-Cas effector protein.
71. The method according to any one of claims 65 to 70, wherein the engineered protein recognizes non-natural PAM sites and / or sequences (e.g., the engineered protein contains altered PAM specificity compared to wild-type LbCas12a).
Citation Information
Patent Citations
Seed specific transcriptional regulation
EP0255378A2
Plant ubiquitin promoter system
EP0342926A2
Tissue-preferential promoters
EP0452269A2
Adenosine nucleobase editors and uses thereof
US10113163B2
Nucleobase editors and uses thereof
US10167457B2