Type v crispr-cas base editors and methods of use thereof
Type V CRISPR-Cas effector proteins combined with deaminases and guide nucleic acids enhance the efficiency and versatility of nucleic acid modifications, addressing the limitations of existing tools by enabling precise editing in various organisms, including plants.
Patent Information
- Application Number
- JP2025107892
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-10-30
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-15
AI Technical Summary
Existing base editing tools are less efficient across a variety of organisms, including plants, and there is a need for more versatile nucleic acid modification methods.
The use of Type V CRISPR-Cas effector proteins in conjunction with deaminases and guide nucleic acids, either through protein-protein or RNA-protein interactions, to modify target nucleic acids, facilitated by fusion proteins and recruitment motifs, enabling co-expression for enhanced editing efficiency.
This approach enhances the versatility and efficiency of base editing across different organisms, including plants, by recruiting deaminases to target nucleic acids, thereby improving the precision and effectiveness of nucleic acid modifications.
Smart Images

Figure 2025157268000001_ABST
Abstract
Description
[Technical Field]
[0001] [Statement regarding electronic filing of sequence listings] An ASCII text sequence listing filed under 37 CFR § 1.821 entitled 1499-10WO_ST25.txt, 351,999 bytes in size, created on October 30, 2020, and submitted via EFS-Web, is provided in lieu of a paper copy. This sequence listing is incorporated herein by reference for its disclosure.
[0002] FIELD OF THE INVENTION The present invention relates to type V CRISPR-Cas effector proteins, deaminases, and fusion and mobilization nucleic acid constructs thereof. The present invention further relates to methods for targeted nucleic acid modification using the same. [Background technology]
[0003] Gene editing is a process that utilizes site-specific nucleases to introduce variations at targeted genomic locations. C to T base editing can be achieved by a base editor that uses Cas9 as a targeting module. For example, Komor et al. (Sci Advances 3(8):eaao4774) (2017)) and Koblan et al. (Nat Biotechnol 36(9):843-846 (2018)) disclose a base editor that uses Cas9 as a fusion protein containing APOBEC1 deaminase, Cas9 nickase (D10A), and two copies of uracil glycosylase inhibitor (UGI). Li et al. (Nat Biotechnol. 36(4):324-327 (2018)) replaced Cas9 with an inactivated Cpf1. Although active in human cells, this Cpf1 construct is less efficient than its Cas9 counterpart. Summary of the Invention [Problem to be solved by the invention]
[0004] To make base editing more useful across a larger number of organisms, including plants, new base editing tools are needed. [Means for solving the problem]
[0005] One aspect of the present invention provides a method of modifying a target nucleic acid, the method comprising contacting the target nucleic acid with (a) a type V CRISPR-Cas effector protein; (b) a deaminase (wherein the target nucleic acid may be contacted with two or more deaminases); and (c) a guide nucleic acid, wherein the deaminase is recruited to the type V CRISPR-Cas effector protein (e.g., via protein-protein interaction, RNA-protein interaction, and / or chemical interaction), thereby modifying the target nucleic acid, wherein the type V CRISPR-Cas effector protein, the deaminase, and the guide nucleic acid may be co-expressed.
[0006] A second aspect of the present invention provides a method for modifying a target nucleic acid, the method comprising contacting the target nucleic acid with (a) a type V CRISPR-Cas fusion protein comprising a type V CRISPR-Cas effector protein fused to a peptide tag (e.g., an epitope or a multimerization epitope); (b) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the peptide tag (the target nucleic acid may be contacted with two or more deaminase fusion proteins); and (c) a guide nucleic acid, wherein the type V CRISPR-Cas fusion protein, deaminase fusion protein and guide nucleic acid may be co-expressed, thereby modifying the target nucleic acid.
[0007] A third aspect of the present invention provides a method for modifying a target nucleic acid, the method comprising contacting the target nucleic acid with (a) a type V CRISPR-Cas fusion protein comprising a type V CRISPR-Cas effector protein fused to an affinity polypeptide that binds to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to a peptide tag (e.g., an epitope or a multimerization epitope) (the target nucleic acid may be contacted with two or more deaminase fusion proteins); and (c) a guide nucleic acid, wherein the type V CRISPR-Cas fusion protein, deaminase fusion protein and guide nucleic acid may be co-expressed, thereby modifying the target nucleic acid.
[0008] A fourth aspect provides a method of modifying a target nucleic acid, the method comprising contacting the target nucleic acid with (a) a Type V CRISPR-Cas effector protein; (b) a recruitment guide nucleic acid comprising a guide nucleic acid linked to an RNA recruitment motif, and (c) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif (the target nucleic acid may be contacted with two or more deaminase fusion proteins); the Type V CRISPR-Cas effector protein, the deaminase fusion protein and the recruitment guide nucleic acid may be co-expressed, thereby modifying the target nucleic acid.
[0009] A fifth aspect of the present invention provides a nucleic acid construct comprising: (a) a Type V CRISPR-Cas fusion protein comprising a Type V CRISPR-Cas effector protein fused to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the peptide tag; and (c) a guide nucleic acid.
[0010] A sixth aspect of the present invention provides a nucleic acid construct comprising: (a) a type V CRISPR-Cas effector protein; (b) a recruitment guide nucleic acid comprising a guide nucleic acid linked to an RNA recruitment motif; and (c) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif.
[0011] A seventh aspect of the present invention provides a V-type CRISPR-Cas fusion protein, comprising: (a) a V-type CRISPR-Cas effector protein fused to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the peptide tag; and (c) a V-type Clustered Regularly Interspaced Short Palindromic Determinant (CDR) fusion protein comprising a guide nucleic acid comprising a spacer sequence and a repeat sequence. The present invention provides a CRISPR-associated Repeats (CRISPR) (Cas) (CRISPR-Cas) system, wherein a guide nucleic acid is capable of forming a complex with a type V CRISPR-Cas effector protein of a type V CRISPR-Cas fusion protein, wherein a spacer sequence is capable of hybridizing to a target nucleic acid, thereby guiding the type V CRISPR-Cas fusion protein to the target nucleic acid, and wherein a deaminase fusion protein is recruited to the type V CRISPR-Cas fusion protein and the target nucleic acid by binding of an affinity polypeptide to a peptide tag fused to the type V CRISPR-Cas fusion protein, thereby enabling the system to modify (e.g., cleave or edit) the target nucleic acid.
[0012] An eighth aspect of the present invention provides a V-type Clustered Regularly Interspaced Short Palindromic RNA (V-CRISPR) fusion protein comprising: (a) a V-type CRISPR-Cas effector protein; (b) a recruitment guide nucleic acid comprising a guide nucleic acid linked to an RNA recruitment motif; and (c) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif. The present invention provides a Crisp Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) system, wherein a recruitment guide nucleic acid comprises a spacer sequence and a repeat sequence, the guide nucleic acid is capable of forming a complex with a V-type CRISPR-Cas effector protein, the recruitment guide nucleic acid is capable of hybridizing to a target nucleic acid, thereby guiding the V-type CRISPR-Cas effector protein to the target nucleic acid, and wherein a deaminase fusion protein is recruited to the V-type CRISPR-Cas effector protein and the target nucleic acid by binding of an affinity polypeptide to an RNA recruitment motif fused to the recruitment guide nucleic acid, thereby enabling the system to modify (e.g., cleave or edit) the target nucleic acid.
[0013] The present invention further provides expression cassettes and / or vectors comprising the nucleic acid constructs of the present invention, and cells comprising the polypeptides, fusion proteins and / or nucleic acid constructs of the present invention. Additionally, the present invention provides kits comprising the nucleic acid constructs of the present invention and expression cassettes, vectors and / or cells comprising same.
[0014] It should be noted that aspects of the invention described with respect to one embodiment may be incorporated into a different embodiment even if not specifically described therein. That is, all embodiments and / or features of any embodiment may be combined in any manner and / or combination. Applicant reserves the right to modify any originally filed claims and / or file any new claims as appropriate, including the right to amend any originally filed claims to depend on and / or incorporate any feature of any other claim(s), even if not originally claimed in that manner. These and other objects and / or aspects of the invention are described in detail in the specification set forth below. Additional features, advantages, and details of the invention will be understood by those skilled in the art from a reading of the following drawings and detailed description of the preferred embodiments, such description being merely illustrative of the invention. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a graph showing C to T editing efficiency for an editing system using pWg120029 as a guide nucleic acid according to some embodiments of the present invention. [Figure 2] 1 is a graph showing C to T editing efficiency for an editing system using pWg120360 as a guide nucleic acid according to some embodiments of the present invention. [Figure 3] 1 is a graph showing C to T editing efficiency for an editing system using pWg120300 as a guide nucleic acid according to some embodiments of the present invention. [Figure 4] 1 is a graph showing C to T editing efficiency for an editing system using pWg120301 as a guide nucleic acid according to some embodiments of the present invention. [Figure 5] 1 is a graph showing A to G editing efficiency for editing systems using guide nucleic acids different from some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] The present invention is described below with reference to the accompanying drawings and examples, in which embodiments of the invention are shown. This description is not intended to be a detailed catalog of all the different ways in which the invention may be practiced or all the features that may be added to the invention. For example, features illustrated with respect to one embodiment may be incorporated into other embodiments, and features illustrated with respect to a particular embodiment may be omitted from that embodiment. Thus, it is contemplated that the invention may exclude or omit, in some embodiments of the invention, any feature or combination of features illustrated herein. Moreover, numerous variations and additions to the various embodiments suggested herein will be apparent to those skilled in the art in light of this disclosure and do not depart from the invention. Therefore, the following description is intended to illustrate some particular embodiments of the invention, but is not intended to exhaustively identify all permutations, combinations, and variations thereof.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. The terminology used in the description of the present invention herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present invention.
[0018] All publications, patent applications, patents, and other references cited herein are incorporated by reference in their entirety for the teachings relevant to the sentence and / or paragraph in which the reference is set forth.
[0019] Unless the context dictates otherwise, it is expressly intended that the various features of the invention described herein can be used in any combination. Moreover, it is also contemplated that in some embodiments of the invention, any feature or combination of features described herein can be excluded or omitted. By way of example, if the specification states that a composition includes components A, B, and C, it is expressly intended that any of A, B, or C, or any combination thereof, alone or in any combination, can be omitted and negated.
[0020] As used in the description of this invention and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise.
[0021] Also, as used herein, "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or").
[0022] As used herein, the term "about," when referring to a measurable value, such as an amount or concentration, means to encompass the specified value as well as a variation of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the specified value. For example, "about X," where X is a measurable value, means to include X and a variation of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of X. Ranges provided herein for measurable values may include any other ranges and / or individual values therein.
[0023] As used herein, phrases such as "between X and Y" and "between about X and Y" should be interpreted to include X and Y. As used herein, phrases such as "between about X and Y" mean "between about X and about Y," and phrases such as "from about X to Y" mean "from about X to about Y."
[0024] The recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each individual value falling within the range, unless otherwise indicated herein, and each individual value is incorporated herein as if it were individually recited herein. For example, if the range 10 to 15 is disclosed, then 11, 12, 13, and 14 are also disclosed.
[0025] As used herein, the terms "comprise", "comprises" and "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0026] As used herein, the transitional phrase "consisting essentially of" means that the claims should be construed to include the specified materials or steps recited in the claims and that do not materially affect the basic and novel feature(s) of the claimed invention. Thus, the term "consisting essentially of," when used in the claims of the present invention, is not intended to be construed as equivalent to "comprising."
[0027] As used herein, the terms "increase," "increasing," "enhance," "enhancing," "improve," and "improving" (and grammatical variations thereof) describe an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 150%, 200%, 300%, 400%, 500% or more compared to a control.
[0028] As used herein, the terms "reduce," "reduced," "reducing," "reduction," "diminish," and "decrease" (and grammatical variations thereof) describe a reduction of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%, for example, as compared to a control. In certain embodiments, the reduction results in no or essentially no detectable activity or amount (i.e., an insignificant amount, e.g., less than about 10% or even 5%).
[0029] A "heterologous" or "recombinant" nucleotide sequence is a nucleotide sequence that is not naturally associated with the host cell into which it is introduced, and includes non-naturally occurring multiple copies of a naturally occurring nucleotide sequence.
[0030] A "native" or "wild-type" nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence refers to a naturally occurring or endogenous nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence. Thus, for example, a "wild-type mRNA" is an mRNA that occurs naturally in a reference organism or is endogenous to the reference organism. A "homologous" nucleic acid sequence is a nucleotide sequence that is naturally associated with the host cell into which it is introduced.
[0031] As used herein, the terms "nucleic acid," "nucleic acid molecule," "nucleotide sequence," and "polynucleotide" refer to linear or branched, single-stranded, or double-stranded RNA or DNA, or a hybrid thereof. The term also encompasses RNA / DNA hybrids. When dsRNA is produced synthetically, less common bases, such as inosine, 5-methylcytosine, 6-methyladenine, hypoxanthine, and others, can also be used for pairing antisense, dsRNA, and ribozymes. For example, polynucleotides containing C-5 propyne analogs of uridine and cytidine have been shown to bind to RNA with high affinity and to be potent antisense inhibitors of gene expression. Other modifications, such as modifications to the phosphodiester backbone of RNA or the 2'-hydroxyl in the ribose sugar group, can also be made.
[0032] As used herein, the term "nucleotide sequence" refers to a heteropolymer of nucleotides or the sequence of these nucleotides from the 5' to 3' end of a nucleic acid molecule, including DNA or RNA molecules, including cDNA, DNA fragments or portions, genomic DNA, synthetic (e.g., chemically synthesized) DNA, plasmid DNA, mRNA, and antisense RNA, any of which may be single-stranded or double-stranded. The terms "nucleotide sequence," "nucleic acid," "nucleic acid molecule," "nucleic acid construct," "recombinant nucleic acid," "oligonucleotide," and "polynucleotide" are also used interchangeably herein to refer to a heteropolymer of nucleotides. Nucleic acid molecules and / or nucleotide sequences provided herein are presented herein in the 5' to 3' direction, from left to right, and are represented using the standard code for representing nucleotide symbols as set forth in the U.S. Sequence Code, 37 CFR §§ 1.821-1.825, and World Intellectual Property Organization (WIPO) Standard ST.25. As used herein, the term "5' region" can refer to the region of a polynucleotide closest to the 5' end of the polynucleotide. Thus, for example, an element in the 5' region of a polynucleotide can be located anywhere from the first nucleotide located at the 5' end of the polynucleotide to a nucleotide located in the middle of the polynucleotide. As used herein, the term "3' region" can refer to the region of the polynucleotide closest to the 3' end of the polynucleotide. Thus, for example, an element in the 3' region of a polynucleotide can be located anywhere from the first nucleotide located at the 3' end of the polynucleotide to a nucleotide located in the middle of the polynucleotide.
[0033] As used herein, the term "gene" refers to a nucleic acid molecule that can be used to produce mRNA, antisense RNA, miRNA, anti-microRNA antisense oligodeoxyribonucleotide (AMO), etc. A gene may or may not be capable of being used to produce a functional protein or gene product. A gene can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and / or 5' and 3' untranslated regions). A gene may be "isolated," thereby meaning a nucleic acid that is substantially or essentially free from components normally found associated with the nucleic acid in its natural state. Such components include other cellular material, culture medium from recombinant production, and / or various chemicals used in the chemical synthesis of nucleic acids.
[0034] The term "mutation" refers to point mutations (e.g., missense, or nonsense, or single base pair insertions or deletions resulting in frameshifts), insertions, deletions, and / or truncations. When a mutation is a substitution of a residue in an amino acid sequence for another residue, or a deletion or insertion of one or more residues in a sequence, the mutation is typically described by identifying the original residue followed by the position of that residue in the sequence and the identity of the newly substituted residues.
[0035] As used herein, the term "complementary" or "complementarity" refers to the natural binding of polynucleotides by base pairing under permissive salt and temperature conditions. For example, the sequence "AGT" (5' to 3') binds to the complementary sequence "TCA" (3' to 5'). Complementarity between two single-stranded molecules can be "partial," where only a small portion of the nucleotides bind, or "complete," where there is overall complementarity between the single-stranded molecules. The degree of complementarity between nucleic acid strands has a significant effect on the efficiency and strength of hybridization between nucleic acid strands.
[0036] As used herein, "complement" can mean 100% complementarity with a comparator nucleotide sequence, or it can mean less than 100% complementarity (e.g., "substantially complementary," e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, etc. complementarity).
[0037] A "portion" or "fragment" of a nucleotide sequence or polypeptide is a nucleotide sequence or polypeptide of reduced length (e.g., a reduction of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more residue(s) (e.g., nucleotide(s) or peptide(s))) compared to a reference nucleotide sequence or polypeptide, respectively, and is not intended to be limiting. "Nucleic acid fragment" is understood to mean a nucleotide sequence or polypeptide comprising, consisting essentially of, and / or consisting of consecutive residues identical or nearly identical (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical) to a nucleic acid sequence or polypeptide. Such a nucleic acid fragment or portion according to the invention may, where appropriate, be comprised within a larger polynucleotide of which it is a component. By way of example, the repeat sequence of a guide nucleic acid of the invention may comprise a portion of a wild-type CRISPR-Cas repeat sequence (e.g., a wild-type V-type CRISPR Cas repeat, such as a repeat from a CRISPR Cas system including, but not limited to, Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c).
[0038] Different nucleic acids or proteins that share homology are referred to herein as "homologs." The term "homolog" includes homologous sequences from the same and other species and orthologous sequences from the same and other species. "Homology" refers to the level of similarity between two or more nucleic acid and / or amino acid sequences in terms of percent positional identity (i.e., sequence similarity or identity). Homology also refers to the concept of similar functional properties between different nucleic acids or proteins. Thus, the compositions and methods of the present invention further include homologs to the nucleotide and polypeptide sequences of the present invention. As used herein, "orthologous" refers to homologous nucleotide and / or amino acid sequences in different species that arose from a common ancestral gene during speciation. Homologs of the nucleotide sequences of the present invention have substantial sequence identity (e.g., at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100%) to the above-described nucleotide sequences of the present invention.
[0039] "Sequence identity," as used herein, refers to the degree to which two optimally aligned polynucleotide or polypeptide sequences are invariant throughout the window of alignment of the components (eg, nucleotides or amino acids). "Identity" can be readily calculated by known methods, including, but not limited to, those described in Computational Molecular Biology (Lesk, AM, ed.) Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, DW, ed.) Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, AM, and Griffin, HG, eds.) Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G., ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.) Stockton Press, New York (1991).
[0040] As used herein, the term "percent sequence identity" or "percent identity" refers to the percentage of identical nucleotides in a linear polynucleotide sequence of a reference ("query") polynucleotide molecule (or its complementary strand) compared to a test ("subject") polynucleotide molecule (or its complementary strand) when the two sequences are optimally aligned. In some embodiments, "percent identity" can refer to the percentage of identical amino acids in an amino acid sequence compared to a reference polypeptide.
[0041] As used herein, the phrase "substantially identical" or "substantial identity" in the context of two nucleic acid molecules, nucleotide sequences, or protein sequences refers to two or more sequences or subsequences that have at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% nucleotide or amino acid residue identity when compared and aligned for maximum correspondence as measured using one of the following sequence comparison algorithms or by visual inspection. In some embodiments of the present invention, substantial identity exists over a region of contiguous nucleotides of the nucleotide sequences of the present invention that is about 10 to about 20 nucleotides, about 10 to about 25 nucleotides, about 10 to about 30 nucleotides, about 15 to about 25 nucleotides, about 30 to about 40 nucleotides, about 50 to about 60 nucleotides, about 70 to about 80 nucleotides, about 90 to about 100 nucleotides, or more nucleotides in length, and any range therein (up to the full length of the sequence). In some embodiments, nucleotide sequences can be substantially identical over at least about 20 nucleotides (e.g., about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 nucleotides). In some embodiments, substantially identical nucleotide or protein sequences perform substantially the same function as the nucleotides (or encoded protein sequences) to which they are substantially identical.
[0042] For sequence comparison, typically one sequence serves as a reference sequence, to which test sequences are compared.When using sequence comparison algorithm, test sequences and reference sequences are input into a computer, subsequence coordinates are designated as necessary, and sequence algorithm program parameters are designated.Then, the sequence comparison algorithm calculates the percent sequence identity for test sequence(s) compared to the reference sequence based on the designated program parameters.
[0043] Optimal sequence alignment for aligning a comparison window is well known to those skilled in the art and may be performed by tools such as the Smith and Waterman local homology algorithm, the Needleman and Wunsch homology alignment algorithm, the Pearson and Lipman similarity search method, or by computerized implementations of algorithms such as GAP, BESTFIT, FASTA, and TFASTA, available as part of the GCG® Wisconsin Package® (Accelrys Inc., San Diego, CA). The "identity fraction" for an aligned segment of a test sequence and a reference sequence is the number of identical elements shared by the two aligned sequences divided by the total number of elements in the reference sequence segment (e.g., the entire reference sequence or a smaller, defined portion of the reference sequence). The percent sequence identity is expressed as the identity fraction multiplied by 100. Comparison of one or more polynucleotide sequences may be to a full-length polynucleotide sequence or a portion thereof, or to a longer polynucleotide sequence. For purposes of the present invention, "percent identity" may be determined using BLASTX version 2.0 for translated nucleotide sequences and BLASTN version 2.0 for polynucleotide sequences.
[0044] Two nucleotide sequences may be considered to be substantially complementary if the two sequences hybridize to each other under stringent conditions. In some exemplary embodiments, two nucleotide sequences considered to be substantially complementary hybridize to each other under very stringent conditions.
[0045] "Stringent hybridization conditions" and "stringent hybridization wash conditions" in the context of nucleic acid hybridization experiments such as Southern and Northern hybridizations are sequence-dependent and will vary under different environmental parameters. A comprehensive guide to nucleic acid hybridization can be found in Tijssen Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, part I, chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assays," Elsevier, New York (1993). In general, highly stringent hybridization and wash conditions are those that achieve the thermal melting point (T) for a specific sequence at a defined ionic strength and pH. m ) is chosen to be approximately 5°C lower than
[0046] T m is the temperature (under defined ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched probe. Very stringent conditions are defined as the T mAn example of stringent hybridization conditions for hybridization of complementary nucleotide sequences with more than 100 complementary residues on a filter in a Southern or Northern blot is 50% formamide containing 1 mg heparin at 42°C, with overnight hybridization. An example of very stringent wash conditions is 0.15 M NaCl at 72°C for approximately 15 minutes. An example of stringent wash conditions is a 0.2×SSC wash at 65°C for 15 minutes (see Sambrook, infra, for a description of SSC buffers). Often, a low stringency wash is performed before a high stringency wash to remove background probe signal. For example, an example of a medium stringency wash for a duplex of more than 100 nucleotides is 1×SSC at 45°C for 15 minutes. An example of a low stringency wash, for example for duplexes of more than 100 nucleotides, is 4-6x SSC at 40°C for 15 minutes. For short probes (e.g., about 10-50 nucleotides), stringent conditions typically include a salt concentration of less than about 1.0 M Na ion, typically about 0.01-1.0 M Na ion (or other salt) (pH 7.0-8.3), and a temperature typically of at least about 30°C. Stringent conditions can also be achieved by the addition of destabilizing agents such as formamide. Generally, a signal-to-noise ratio of 2x (or higher) than that observed for an unrelated probe in a particular hybridization assay indicates detection of specific hybridization. Nucleotide sequences that do not hybridize to each other under stringent conditions are substantially identical if the proteins they encode are substantially identical. This can occur, for example, when copies of nucleotide sequences are made using the maximum codon degeneracy permitted by the genetic code.
[0047] Polynucleotides and / or recombinant nucleic acid constructs of the invention may be codon-optimized for expression. In some embodiments, polynucleotides, nucleic acid constructs, expression cassettes, and / or vectors of the invention (comprising / encoding a Type V CRISPR-Cas effector protein, a deaminase fusion protein, and a recruitment guide nucleic acid, or a Type V CRISPR-Cas fusion protein, a deaminase fusion protein, and a guide nucleic acid) may be codon-optimized for expression in an organism (e.g., an animal, a plant, a fungus, an archaea, or a bacterium). In some embodiments, a codon-optimized nucleic acid construct, polynucleotide, expression cassette, and / or vector of the invention has about 70% to about 99.9% (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9% or 100%) identity or more to a reference nucleic acid construct, polynucleotide, expression cassette, and / or vector that is not codon-optimized.
[0048] In any of the embodiments described herein, the polynucleotides or nucleic acid constructs of the invention may be operably associated with various promoters and / or other regulatory elements for expression in an organism or cells thereof (e.g., plants and / or cells of plants). Thus, in some embodiments, the polynucleotides or nucleic acid constructs of the invention may further comprise one or more promoters, introns, enhancers, and / or terminators operably linked to one or more nucleotide sequences. In some embodiments, a promoter may be operably associated with an intron (e.g., the Ubi1 promoter and an intron). In some embodiments, a promoter associated with an intron may be referred to as a "promoter region" (e.g., the Ubi1 promoter and an intron).
[0049] As used herein, "operably linked" or "operably associated" in reference to a polynucleotide means that the indicated elements are functionally related to each other, and generally also physically related. Thus, as used herein, the terms "operably linked" or "operably associated" refer to nucleotide sequences on a single nucleic acid molecule that are functionally related. Thus, a first nucleotide sequence operably linked to a second nucleotide sequence refers to a situation in which the first nucleotide sequence is placed in a functional relationship with the second nucleotide sequence. For example, a promoter is operably associated with a nucleotide sequence if it acts to transcribe or express the nucleotide sequence. Those skilled in the art will understand that control sequences (e.g., promoters) need not be contiguous with the nucleotide sequence with which they are operably associated, so long as the control sequence functions to direct expression. Thus, for example, an intervening transcribed but untranslated nucleic acid sequence may be present between the promoter and the nucleotide sequence, and the promoter would still be considered "operably linked" to the nucleotide sequence.
[0050] The terms "linked" or "fused," as used herein in reference to polypeptides, refer to the attachment of one polypeptide to another. A polypeptide may be linked or fused to another polypeptide (at either the N- or C-terminus) directly (e.g., via a peptide bond) or via a linker (e.g., a peptide linker).
[0051] The term "linker," in reference to a polypeptide, is art-recognized and refers to a chemical group or molecule that links two molecules or moieties, e.g., two polypeptides (e.g., domains) of a fusion protein, e.g., a Type V CRISPR-Cas effector protein and a peptide tag and / or a polypeptide of interest. A linker may be composed of a single linking molecule (e.g., a single amino acid) or may include more than one linking molecule. In some embodiments, a linker can be an organic molecule, group, polymer, or chemical moiety, e.g., a bivalent organic moiety. In some embodiments, a linker can be an amino acid or a peptide. In some embodiments, a linker is a peptide.
[0052] In some embodiments, peptide linkers useful in the invention are from about 2 to about 100 or more amino acids in length, e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids in length (e.g., about 2 to about 40, about 2 to about 50, about 2 to about 60, about 4 to about 40, about 4 to about 50, about 4 to about 60, About 5 to about 40, about 5 to about 50, about 5 to about 60, about 9 to about 40, about 9 to about 50, about 9 to about 60, about 10 to about 40, about 10 to about 50, about 10 to about 60, or about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47 , 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids in length (e.g., about 105, 110, 115, 120, 130, 140, 150 or more amino acids in length). In some embodiments, the peptide linker may be a GS linker.
[0053] In some embodiments, two or more polynucleotide molecules may be linked by a linker, which may be an organic molecule, group, polymer, or chemical moiety, e.g., a divalent organic moiety. A polynucleotide may be linked or fused to another polynucleotide (at the 5' or 3' end) via a covalent or non-covalent bond, or via a bond involving, for example, Watson-Crick base pairing, or by one or more linking nucleotides. In some embodiments, a polynucleotide motif of a specific structure may be inserted into another polynucleotide sequence (e.g., an extension of a hairpin structure in a guide RNA). In some embodiments, the linking nucleotide may be a naturally occurring nucleotide. In some embodiments, the linking nucleotide may be a non-naturally occurring nucleotide.
[0054] A "promoter" is a nucleotide sequence that controls or regulates the transcription of a nucleotide sequence (e.g., a coding sequence) operably associated with the promoter. The coding sequence controlled or regulated by a promoter can encode a polypeptide and / or functional RNA. Typically, a "promoter" refers to a nucleotide sequence that contains a binding site for RNA polymerase II and directs the initiation of transcription. Generally, promoters are found 5', i.e., upstream, to the start of the coding region of a corresponding coding sequence. A promoter may contain other elements that act as regulators of gene expression; for example, a promoter region. These include a TATA box consensus sequence and often a CAAT box consensus sequence (Breathnach and Chambon, (1981) Annu. Rev. Biochem. 50:349). In plants, the CAAT box may be replaced by an AGGA box (Messing et al., (1983) in Genetic Engineering of Plants, T. Kosuge, C. Meredith and A. Hollaender (eds.), Plenum Press, pp. 211-227). In some embodiments, the promoter region may contain at least one intron (e.g., SEQ ID NO: 1 or SEQ ID NO: 2).
[0055] Promoters useful in the present invention can include, for example, constitutive, inducible, temporally-regulated, developmentally-regulated, chemically-regulated, tissue-preferred and / or tissue-specific promoters for use in preparing recombinant nucleic acid molecules, e.g., "synthetic nucleic acid constructs" or "protein-RNA complexes." These various types of promoters are known in the art.
[0056] The selection of promoter can vary depending on the time and space requirements of expression, and can vary based on the host cell to be transformed.The promoters for many different organisms are well known in the art.Based on the extensive knowledge existing in the art, suitable promoters can be selected for specific target host organisms.Therefore, for example, much is known about the upstream promoters of genes that are highly constitutively expressed in model organisms, and this knowledge can be easily accessed and implemented in other systems as appropriate.
[0057] In some embodiments, promoters functional in plants can be used with the constructs of the present invention. Non-limiting examples of promoters useful for driving expression in plants include the promoter of the Rubisco small subunit gene 1 (PrbcS1), the promoter of the actin gene (Pactin), the promoter of the nitrate reductase gene (Pnr), and the promoter of the duplicated carbonic anhydrase gene 1 (Pdca1) (see Walker et al. Plant Cell Rep. 23:727-735 (2005); Li et al. Gene 403:132-142 (2007); Li et al. Mol Biol. Rep. 37:1143-1154 (2010)). PrbcS1 and Pactin are constitutive promoters, while Pnr and Pdca1 are inducible promoters. Pnr is induced by nitrate and repressed by ammonium (Li et al. Gene 403:132-142 (2007)), and Pdca1 is induced by salt (Li et al. Mol Biol. Rep. 37:1143-1154 (2010)).
[0058] Examples of constitutive promoters useful in plants include, but are not limited to, the cestrum virus promoter (cmp) (U.S. Pat. No. 7,166,770), the rice actin 1 promoter (Wang et al. (1992) Mol. Cell. Biol. 12:3399-3406; and U.S. Pat. No. 5,641,876), the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812), the CaMV 19S promoter (Lawton et al. (1987) Plant Mol. Biol. 9:315-324), the nos promoter (Ebert et al. (1987) Proc. Natl. Acad. Sci USA 84:5745-5749), the Adh promoter (Walker et al. (1987) Proc. Natl. Acad. Sci USA 84:5745-5749), the α-terminal β ... al. (1987) Proc. Natl. Acad. Sci. USA 84:6624-6629), sucrose synthase promoter (Yang & Russell (1990) Proc. Natl. Acad. Sci. USA 87:4144-4148), and ubiquitin promoter. Constitutive promoters derived from ubiquitin have accumulated in many cell types. Ubiquitin promoters have been cloned from several plant species for use in transgenic plants (e.g., sunflower (Binet et al., 1991. Plant Science 79:87-94), maize (Christensen et al., 1989. Plant Molec. Biol. 12:619-632), and Arabidopsis (Norris et al. 1993. Plant Molec. Biol. 21:895-906)). The maize ubiquitin promoter (UbiP) has been developed in a transgenic monocotyledonous plant system, and its sequence and a vector constructed for monocotyledonous plant transformation are disclosed in European Patent Publication EP 0 342 926. The ubiquitin promoter is suitable for expression of the nucleotide sequence of the present invention in transgenic plants, particularly monocotyledonous plants.Additionally, the promoter expression cassette described by McElroy et al. (Mol. Gen. Genet. 231:150-160 (1991)) can be readily modified for expression of the nucleotide sequences of the present invention and is particularly suitable for use in monocotyledonous hosts.
[0059] In some embodiments, tissue-specific / tissue-preferred promoters can be used for expression of heterologous polynucleotides in plant cells. Tissue-specific or preferred expression patterns include, but are not limited to, green tissue-specific or preferred, root-specific or preferred, stem-specific or preferred, flower-specific or preferred, or pollen-specific or preferred. Promoters suitable for expression in green tissues include many of those regulating genes involved in photosynthesis, many of which have been cloned from both monocotyledonous and dicotyledonous plants. In one embodiment, a promoter useful in the present invention is the maize PEPC promoter from the phosphoenol carboxylase gene (Hudspeth & Grula, Plant Molec. Biol. 12:579-589 (1989)). Non-limiting examples of tissue-specific promoters include those associated with genes encoding seed storage proteins (e.g., β-conglycinin, cruciferin, napin, and phaseolin), zein or oil body proteins (e.g., oleosin), or proteins involved in fatty acid biosynthesis (including acyl carrier protein, stearoyl-ACP desaturase, and fatty acid desaturase (fad2-1)), and other nucleic acids expressed during embryo development (e.g., Bce4; see, e.g., Kridl et al. (1991) Seed Sci. Res. 1:209-219; and EP 255378). Tissue-specific or tissue-preferred promoters useful for expression of the nucleotide sequences of the invention in plants, particularly maize, include, but are not limited to, those that direct expression in roots, pith, leaves, or pollen. Such promoters are disclosed, for example, in WO 93 / 07278, incorporated herein by reference for its disclosure of promoters.Other non-limiting examples of tissue-specific or tissue-preferred promoters useful in the present invention include the cotton Rubisco promoter disclosed in U.S. Pat. No. 6,040,504; the rice sucrose synthase promoter disclosed in U.S. Pat. No. 5,604,121; the root-specific promoter described by de Framond (FEBS 290:103-106 (1991); European Patent EP 0 452 269 (Ciba-Geigy)); the stem-specific promoter driving expression of the maize trpA gene described in U.S. Pat. No. 5,625,136 (Ciba-Geigy); the cestrum yellow leaf curling virus promoter disclosed in WO 01 / 73087; and, without limitation, the ProOsLPS10 and ProOsLPS11 promoters from rice (Nguyen et al. Plant Biotechnol. Reports 9(5):297-306(2015)), ZmSTK2_USP from maize (Wang et al. Genome 60(6):485-495(2017)), LAT52 and LAT59 from tomato (Twell et al. Development 109(3):705-713(1990)), Zm13 (U.S. Patent No. 10,421,972), the PLA2-δ promoter from Arabidopsis thaliana (U.S. Patent No. 7,141,424), and / or the ZmC5 promoter from maize (International PCT Publication No. WO1999 / 042587).
[0060] Further examples of plant tissue-specific / tissue-preferred promoters include, but are not limited to, root hair-specific cis-element (RHE) (Kim et al. The Plant Cell 18:2958-2970 (2006)), root-specific promoters RCc3 (Jeong et al. Plant Physiol. 153:185-197 (2010)) and RB7 (U.S. Pat. No. 5,459,252), lectin promoter (Lindstrom et al. (1990) Der. Genet. 11:160-167; and Vodkin (1983) Prog. Clin. Biol. Res. 138:87-98), maize alcohol dehydrogenase 1 promoter (Dennis et al. (1984) Nucleic Acids Res. 12:3983-4000), S-adenosyl-L-methionine synthase (SAMS) (Vander Mijnsbrugge et al. (1996) Plant and Cell Physiology, 37(8):1108-1115), maize light-harvesting complex promoter (Bansal et al. (1992) Proc. Natl. Acad. Sci. USA 89:3654-3658), maize heat shock protein promoter (O'Dell et al. (1985) EMBOJ. 5:451-458; and Rochester et al. (1986) EMBOJ. 5:451-458), pea small subunit RuBP carboxylase promoter (Cashmore, "Nuclear genes encoding the small subunit of ribulose-1,5-bisphosphate carboxylase," pp. 29-39 In: Genetic Engineering of Plants (Hollaender ed., Plenum Press 1983); and Poulsen et al. al. (1986) Mol. Gen. Genet. 205:193-200), the mannopine synthase promoter of the Ti plasmid (Langridge et al. (1989) Proc. Natl. Acad. Sci. USA 86:3219-3223), the Ti plasmid nopaline synthase promoter (Langridge et al.(1989), supra), the petunia chalcone isomerase promoter (van Tunen et al. (1988) EMBOJ.7:1257-1263), the bean glycine-rich protein 1 promoter (Keller et al. (1989) Genes Dev.3:1639-1646), the truncated CaMV 35S promoter (O'Dell et al. (1985) Nature 313:810-812), the potato patatin promoter (Wenzler et al. (1989) Plant Mol. Biol.13:347-354), a root cell promoter (Yamamoto et al. (1990) Nucleic Acids Res.18:7449), and the maize zein promoter (Kriz et al. (1987) Mol. Gen. Genet.207:90-98; Langridge et al. al. (1983) Cell 34:1015-1022; Reina et al. (1990) Nucleic Acids Res. 18:6425; Reina et al. (1990) Nucleic Acids Res. 18:7449; and Wandelt et al. (1989) Nucleic Acids Res. 17:2354), the globulin-1 promoter (Belanger et al. (1991) Genetics 129:863-872), the α-tubulin cab promoter (Sullivan et al. (1989) Mol. Gen. Genet. 215:431-440), the PEPCase promoter (Hudspeth & Grula (1989) Plant Mol. Biol. 12:579-589), the R gene complex-associated promoter (Chandler et al. (1989) Plant Cell 1:1175-1183), and the chalcone synthase promoter (Franken et al. (1991) EMBO J. 10:2605-2612).
[0061] Useful for species-specific expression is the pea vicilin promoter (Czako et al. (1992) Mol. Gen. Genet. 235:33-40; as well as the species-specific promoters disclosed in U.S. Pat. No. 5,625,136). Useful promoters for expression in mature leaves are those that are switched on during senescence, such as the SAG promoter from Arabidopsis (Gan et al. (1995) Science 270:1986-1988).
[0062] Additionally, promoters functional in chloroplasts can be used. Non-limiting examples of such promoters include the bacteriophage T3 gene 9 5'UTR and other promoters disclosed in U.S. Patent No. 7,579,516. Other promoters useful in the present invention include, but are not limited to, the S-E9 small subunit RuBP carboxylase promoter and the Kunitz trypsin inhibitor gene promoter (Kti3).
[0063] Additional regulatory elements useful in the present invention include, but are not limited to, introns, enhancers, termination sequences and / or 5' and 3' untranslated regions.
[0064] Introns useful in the present invention may be introns identified and isolated in plants and inserted into expression cassettes used in plant transformation. As will be understood by those skilled in the art, introns can contain sequences necessary for self-excision, which are integrated in-frame within the nucleic acid construct / expression cassette. Introns can be used as spacers to separate multiple protein-coding sequences within a single nucleic acid construct, or introns can be used within a single protein-coding sequence, for example, to stabilize mRNA. When used within a protein-coding sequence, they are inserted "in-frame" with the included excision site. Introns can also be associated with promoters to improve or modify expression. By way of example, promoter / intron combinations useful in the present invention include, but are not limited to, the maize Ubi1 promoter and intron combination.
[0065] Non-limiting examples of introns useful in the present invention include introns from the ADHI gene (e.g., Adh1-S introns 1, 2, and 6), an intron from the ubiquitin gene (Ubi1), an intron from the Rubisco small subunit (rbcS) gene, an intron from the Rubisco large subunit (rbcL) gene, an intron from the actin gene (e.g., the actin-1 intron), an intron from the pyruvate dehydrogenase kinase gene (pdk), an intron from the nitrate reductase gene (nr), an intron from the double carbonic anhydrase gene 1 (Tdca1), an intron from the psbA gene, an intron from the atpA gene, or any combination thereof.
[0066] In some embodiments, the polynucleotides and / or nucleic acid constructs of the present invention may be "expression cassettes" or may be included within an expression cassette. As used herein, "expression cassette" refers to a recombinant nucleic acid molecule comprising, for example, a nucleic acid construct of the present invention (e.g., a polynucleotide encoding a type V CRISPR-Cas effector protein, a polynucleotide encoding a type V CRISPR-Cas fusion protein, a polynucleotide encoding a deaminase (e.g., cytosine deaminase and / or adenine deaminase), a polynucleotide encoding a deaminase fusion protein, a polynucleotide encoding a peptide tag, a polynucleotide encoding an affinity polypeptide, a recruitment guide nucleic acid, and / or a guide nucleic acid), wherein the nucleic acid construct is operably associated with at least control sequences (e.g., primers). Thus, some embodiments of the present invention provide, for example, expression cassettes designed to express the nucleic acid constructs of the present invention. When an expression cassette contains more than one polynucleotide, the polynucleotides may be operably linked to a single promoter that drives the expression of all polynucleotides, or the polynucleotides may be operably linked to one or more different promoters (e.g., three polynucleotides may be driven by any combination of one, two, or three promoters). Thus, for example, the polynucleotide encoding a V-type CRISPR-Cas fusion protein, the polynucleotide encoding a deaminase fusion protein, and the guide nucleic acid contained in the expression cassette may each be operably associated with a single promoter, or may be operably associated with any combination of separate promoters (e.g., two or three promoters). As another example, the polynucleotide encoding a V-type CRISPR-Cas effector protein, the polynucleotide encoding a deaminase fusion protein, and the recruitment guide nucleic acid contained in the expression cassette may each be operably associated with a single promoter, or may be operably associated with any combination of separate promoters (e.g., two or three promoters).
[0067] In some embodiments, expression cassettes comprising the polynucleotides / nucleic acid constructs of the present invention may be optimized for expression in an organism (eg, an animal, a plant, a bacterium, etc.).
[0068] An expression cassette comprising a nucleic acid construct of the invention may be chimeric, meaning that at least one of its components is heterologous to at least one of its other components (e.g., a promoter from a host organism operably linked to a polynucleotide of interest to be expressed in the host organism, where the polynucleotide of interest is from an organism different from the host or is not normally found in association with that promoter). Expression cassettes may be of natural origin, but have been obtained in a recombinant form useful for heterologous expression.
[0069] The expression cassette can also include a transcriptional and / or translational termination region (i.e., a termination region) and / or an enhancer region that is functional in the selected host cell. A variety of transcription terminators and enhancers are known in the art and available for use in expression cassettes. The transcription terminator is responsible for transcription termination and correct mRNA polyadenylation. The termination region and / or enhancer region can be native to the transcription initiation region, native to the gene encoding the CRISPR-Cas effector protein or the gene encoding the deaminase, native to the host cell, or native to another source (e.g., foreign or heterologous to the promoter, the gene encoding the CRISPR-Cas effector protein, or the gene encoding the deaminase, the host cell, or any combination thereof).
[0070] The expression cassettes of the present invention may also include a polynucleotide encoding a selectable marker that can be used to select transformed host cells. As used herein, a "selectable marker" refers to a polynucleotide sequence that, when expressed, confers a distinct phenotype on host cells expressing the marker, such that such transformed cells can be distinguished from those that do not possess the marker. Such polynucleotide sequences may encode either a selectable marker or a screenable marker, depending on whether the marker confers a characteristic that can be selected for by chemical means, e.g., by using a selection agent (e.g., an antibiotic), or whether the marker is a simple characteristic that can be identified through observation or testing, e.g., by screening (e.g., fluorescence). Many examples of suitable selectable markers are known in the art and can be used in the expression cassettes described herein.
[0071] The expression cassettes, nucleic acid molecules / constructs, and polynucleotide sequences described herein may be used in reference to vectors. The term "vector" refers to a composition for transferring, delivering, or introducing a nucleic acid(s) into a cell. A vector includes a nucleic acid construct containing the nucleotide sequence(s) to be transferred, delivered, or introduced. Vectors for use in transforming host organisms are well known in the art. Non-limiting examples of general classes of vectors include viral vectors, plasmid vectors, phage vectors, phagemid vectors, cosmid vectors, fosmid vectors, bacteriophages, artificial chromosomes, minicircles, or Agrobacterium binary vectors in double- or single-stranded, linear, or circular forms, which may or may not be self-transmissible or motile. In some embodiments, viral vectors include, but are not limited to, retroviral, lentiviral, adenoviral, adeno-associated, or herpes simplex viral vectors. As defined herein, a vector can transform a prokaryotic or eukaryotic host by integration into the cellular genome or can exist extrachromosomally (e.g., an autonomously replicating plasmid with an origin of replication). Also included is a shuttle vector, which refers to a DNA vehicle that is naturally or by design capable of replicating in two different host organisms, which may be selected from actinomycetes and related species, bacteria, and eukaryotes (e.g., higher plants, mammals, yeast, or fungal cells). In some embodiments, the nucleic acid in the vector is under the control of, and operably linked to, an appropriate promoter or other regulatory elements for transcription in the host cell. The vector may be a bifunctional expression vector that functions in multiple hosts. In the case of genomic DNA, this may include its own promoter and / or other regulatory elements, and in the case of cDNA, it may be under the control of an appropriate promoter and / or other regulatory elements for expression in the host cell. Thus, the nucleic acid construct of the present invention and / or expression cassettes containing it may be contained within a vector described herein and known in the art.
[0072] As used herein, "contact," "contacting," "contacted," and grammatical variations thereof refer to bringing together components of a desired reaction under conditions appropriate for carrying out the desired reaction (e.g., transformation, transcriptional regulation, genome editing, nicking, and / or cleavage). Thus, for example, a target nucleic acid may be contacted with a nucleic acid construct of the invention encoding, for example, a type V CRISPR-Cas fusion protein, a deaminase fusion protein, and a guide nucleic acid under conditions in which the type V CRISPR-Cas fusion protein is expressed, the type V CRISPR-Cas fusion protein forms a complex with the guide nucleic acid, the complex hybridizes to the target nucleic acid, and the deaminase fusion protein is recruited to the type V CRISPR-Cas effector protein (and thus to the target nucleic acid), thereby modifying the target nucleic acid. In some embodiments, a Type V CRISPR-Cas protein (which may be a Type V CRISPR-Cas fusion protein), a guide nucleic acid, and a deaminase (which may be a deaminase fusion protein) contact a target nucleic acid, thereby modifying the nucleic acid. In some embodiments, the Type V CRISPR-Cas protein, guide nucleic acid, and / or deaminase may be in the form of a complex (e.g., a ribonucleoprotein, e.g., an assembled ribonucleoprotein complex), and the complex contacts the target nucleic acid. In some embodiments, the complex or a component thereof (e.g., the guide nucleic acid) hybridizes to the target nucleic acid, thereby modifying the target nucleic acid (e.g., through the action of the Type V CRISPR-Cas protein and / or the deaminase). In some embodiments, the deaminase or deaminase fusion protein and the Type V CRISPR-Cas effector protein localize to the target nucleic acid (which may be via covalent and / or non-covalent interactions).
[0073] As used herein, "modifying" or "modification" in reference to a target nucleic acid includes editing (e.g., mutation), covalent modification, exchange / substitution of nucleic acid / nucleotide bases, deletion, cleavage, nicking, and / or transcriptional regulation of the target nucleic acid.
[0074] As used herein, "recruit," "recruiting," or "recruitment" refers to attracting one or more polypeptide(s) or polynucleotide(s) to another polypeptide or polynucleotide (e.g., to a specific location in a genome) using protein-protein interactions, RNA-protein interactions, and / or chemical interactions. Protein-protein interactions can include, but are not limited to, peptide tags (e.g., epitopes, multimerization epitopes) and corresponding affinity polypeptides, RNA recruitment motifs and corresponding affinity polypeptides, and / or chemical interactions. Exemplary chemical interactions that may be useful for polypeptides and polynucleotides for recruitment purposes include, but are not limited to, rapamycin-induced dimerization of FRB-FKBP; biotin-streptavidin interactions; SNAP tags (Hussain et al. Curr Pharm Des. 19(30):5437-42(2013)); Halo tags (Los et al. ACS Chem Biol. 3(6):373-82(2008)); CLIP tags (Gautier et al. Chemistry & Biology 15:128-136(2008)); compound-induced DmrA-DmrC heterodimers (Tak et al. Nat Methods 14(12):1163-1166(2017)); and bifunctional ligand approaches (fusing two protein-binding chemicals together) (Vos et al. Curr Opin Chemical Biology 28:194-201(2015)) (e.g., dihydrofolate reductase (DHFR) (Kopyteck et al. Cell Chem Biol 7(5):313-321(2000)).
[0075] "Introducing," "introduce," "introduced" (and grammatical variations thereof) in the context of a polynucleotide of interest means presenting a nucleotide sequence of interest (e.g., a polynucleotide, a nucleic acid construct, and / or a guide nucleic acid) to a host organism or a cell of the organism (e.g., a host cell; e.g., a plant cell) in a manner that allows the nucleotide sequence to access the interior of the cell. Thus, for example, nucleic acid constructs of the invention encoding the type V CRISPR-Cas fusion proteins and deaminase fusion proteins described herein and a guide nucleic acid can be introduced into a cell of an organism, thereby transforming the cell with the type V CRISPR-Cas effector protein fusion protein, the deaminase fusion protein, and the guide nucleic acid.
[0076] As used herein, the term "transformation" refers to the introduction of heterologous nucleic acid into a cell. Cellular transformation may be stable or transient. Thus, in some embodiments, a host cell or host organism may be stably transformed with a polynucleotide / nucleic acid molecule of the present invention. In some embodiments, a host cell or host organism may be transiently transformed with a nucleic acid construct of the present invention.
[0077] "Transient transformation" in the context of a polynucleotide means that the polynucleotide is introduced into a cell and does not integrate into the genome of the cell.
[0078] By "stably introduce" or "stably introduced" in the context of a polynucleotide introduced into a cell, it is intended that the introduced polynucleotide is stably integrated into the genome of the cell, and thus the cell is stably transformed by the polynucleotide.
[0079] As used herein, "stable transformation" or "stably transformed" means that a nucleic acid molecule is introduced into a cell and integrated into the cell's genome. Thus, the integrated nucleic acid molecule can be inherited by its progeny, more specifically, by multiple generations of progeny. As used herein, "genome" includes the nuclear and plastid genomes, and therefore includes, for example, the integration of a nucleic acid into the chloroplast or mitochondrial genome. As used herein, stable transformation can also refer to a transgene that is maintained extrachromosomally, for example, as a minichromosome or a plasmid.
[0080] Transient transformation can be detected, for example, by enzyme-linked immunosorbent assay (ELISA) or Western blot, which can detect the presence of peptides or polypeptides encoded by one or more transgenes introduced into an organism. Stable transformation of cells can be detected, for example, by Southern blot hybridization assay of the genomic DNA of the cells using a nucleic acid sequence that specifically hybridizes with the nucleotide sequence of the transgene introduced into the organism (e.g., a plant). Stable transformation of cells can be detected, for example, by Northern blot hybridization assay of the RNA of the cells using a nucleic acid sequence that specifically hybridizes with the nucleotide sequence of the transgene introduced into the host organism. Stable transformation of cells can also be detected by polymerase chain reaction (PCR) or other amplification reactions well known in the art, which, for example, use specific primer sequences that hybridize with the target sequence(s) of the transgene, resulting in amplification of the transgene sequence, which can be detected according to standard methods. Transformation can also be detected by direct sequencing and / or hybridization protocols well known in the art.
[0081] Thus, in some embodiments, the nucleotide sequences, polynucleotides, nucleic acid constructs, and / or expression cassettes of the invention may be transiently expressed and / or stably integrated into the genome of a host organism. Thus, in some embodiments, the nucleic acid constructs of the invention may be transiently introduced into a cell along with a guide nucleic acid, and thus the DNA is not maintained within the cell.
[0082] Nucleic acid constructs of the invention can be introduced into cells by any method known to those of skill in the art. In some embodiments of the invention, cell transformation comprises nuclear transformation. In other embodiments, cell transformation comprises plastid transformation (e.g., chloroplast transformation). In further embodiments, recombinant nucleic acid constructs of the invention can be introduced into cells via conventional breeding techniques.
[0083] Procedures for transforming both eukaryotes and prokaryotes are well known and routine in the art and are described throughout the literature (see, e.g., Jiang et al. 2013. Nat. Biotechnol. 31:233-239; Ran et al. Nature Protocols 8:2281-2308 (2013)).
[0084] Thus, nucleotide sequences can be introduced into a host organism or its cells by any number of methods well known in the art. The method of the present invention does not rely on a specific method for introducing one or more nucleotide sequences into an organism, but only requires access to the inside of at least one cell of the organism. When more than one nucleotide sequence is introduced, they can be assembled as part of a single nucleic acid construct, or can be located on the same or different nucleic acid constructs as separate nucleic acid constructs. Thus, nucleotide sequences can be introduced into cells of interest in a single transformation event and / or in separate transformations, or, if relevant, the nucleotide sequences can be incorporated into plants, for example, as part of a breeding protocol.
[0085] The present invention relates to improved base-editing nucleic acid constructs. In some embodiments, the present invention provides a nucleic acid construct comprising: (a) a type V CRISPR-Cas fusion protein comprising a type V CRISPR-Cas effector protein fused to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide capable of binding to the peptide tag; and (c) a guide nucleic acid. In some embodiments, the present invention provides a nucleic acid construct comprising: (a) a Cas12a fusion protein comprising a Cas12a effector protein fused to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the peptide tag; and (c) a guide nucleic acid.
[0086] In some embodiments, the Cas12a (Cpf1) effector protein is selected from the group consisting of LbCpf1 [Lachnospiraceae bacterium], AsCpf1 [Acidaminococcus sp.], BpCpf1 [Butyrivibrio proteoclasticus], CMtCpf1 [Candidatus Methanoplasma termitum], EeCpf1 [Eubacterium eligens], FnCpf1 (Francisella novicida U112), Lb2Cpf1 [Lachnospiraceae bacterium], >Lb3Cpf1 [Lachnospiraceae bacterium], LiCpf1 [Leptospira inadai], MbCpf1 [Moraxella bovoculi 237], PbCpf1 [Parcubacteria bacterium], GWC2011_GWC2_44_17], PcCpf1 [Porphyromonas crevioricanis], PdCpf1 [Prevotella disiens], PeCpf1 [Peregrinibacteria bacterium GW2011_GWA_33_10], PmCpf1 [Porphyromonas macacae], and / or SsCpf1 [Smithella sp. SC_K08D17] (e.g., having the sequence of any one of SEQ ID NOs: 3 to 22). In some embodiments, the Cas12a effector protein may be Lachnospiraceae bacterium ND2006 Cas12a (LbCas12a) (LbCpf1) (e.g., having the sequence of any one of SEQ ID NOs: 3 and 9-11), Acidaminococcus sp. Cpf1 (AsCas12a) (AsCpf1) (e.g., having the sequence of any one of SEQ ID NOs: 4), and / or enAsCas12a (e.g., having the sequence of any one of SEQ ID NOs: 20-22).
[0087] In some embodiments, the present invention provides a nucleic acid construct comprising: (a) a type V CRISPR-Cas fusion protein comprising a type V CRISPR-Cas effector protein fused to an affinity polypeptide capable of binding to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to the peptide tag; and (c) a guide nucleic acid. In some embodiments, the present invention provides a nucleic acid construct comprising: (a) a Cas12a fusion protein comprising a Cas12a effector protein fused to an affinity polypeptide that binds to the peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to the peptide tag; and (c) a guide nucleic acid.
[0088] In some embodiments, a nucleic acid construct is provided, comprising: (a) a type V CRISPR-Cas effector protein; (b) a recruitment guide nucleic acid comprising a guide nucleic acid linked to an RNA recruitment motif; and (c) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif. In some embodiments, a nucleic acid construct is provided, comprising: (a) a Cas12a effector protein; (b) a recruitment guide nucleic acid comprising a guide nucleic acid linked to an RNA recruitment motif; and (c) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif.
[0089] In some embodiments, a V-type CRISPR-Cas fusion protein is provided, comprising: (a) a V-type CRISPR-Cas effector protein fused to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the peptide tag; and (c) a V-type Clustered Regularly Interspaced Short Palindromic Determinant Protein (CRISPR-Cas) fusion protein comprising a guide nucleic acid comprising a spacer sequence and a repeat sequence. A CRISPR-associated Repeats (CRISPR) (Cas) (CRISPR-Cas) system is provided, wherein a guide nucleic acid is capable of forming a complex with a type V CRISPR-Cas effector protein of a type V CRISPR-Cas fusion protein, wherein a spacer sequence is capable of hybridizing to a target nucleic acid, thereby guiding the type V CRISPR-Cas fusion protein to the target nucleic acid, and wherein a deaminase fusion protein is recruited to the type V CRISPR-Cas fusion protein and the target nucleic acid by binding of an affinity polypeptide to a peptide tag fused to the type V CRISPR-Cas fusion protein, thereby enabling the system to modify (e.g., cleave or edit) the target nucleic acid.In some embodiments, a V-type Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) system is provided, comprising: (a) a Cas12a fusion protein comprising a Cas12a effector protein fused to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the peptide tag; and (c) a guide nucleic acid comprising a spacer sequence and a repeat sequence, wherein the guide nucleic acid is capable of forming a complex with the Cas12a effector protein of the Cas12a fusion protein, and the spacer sequence is capable of hybridizing to the target nucleic acid, thereby guiding the Cas12a fusion protein to the target nucleic acid, and wherein the deaminase fusion protein is recruited to the Cas12a fusion protein and the target nucleic acid by binding of the affinity polypeptide to the peptide tag fused to the Cas12a fusion protein, thereby enabling the system to modify (e.g., cleave or edit) the target nucleic acid.
[0090] In some embodiments, the present invention provides a V-type Clustered Regularly Interspaced Short Palindromic Determinant Protein (CRISPR-Cas) comprising: (a) a V-type CRISPR-Cas effector protein; (b) a recruitment guide nucleic acid comprising a guide nucleic acid linked to an RNA recruitment motif; and (c) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif. The present invention provides a Crisp Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) system, wherein a recruitment guide nucleic acid comprises a spacer sequence and a repeat sequence, the guide nucleic acid is capable of forming a complex with a V-type CRISPR-Cas effector protein, and the spacer sequence is capable of hybridizing to the target nucleic acid, thereby guiding the V-type CRISPR-Cas effector protein to the target nucleic acid, and wherein a deaminase fusion protein is recruited to the V-type CRISPR-Cas effector protein and the target nucleic acid by binding of an affinity polypeptide to an RNA recruitment motif fused to the recruitment guide nucleic acid, thereby enabling the system to modify (e.g., cleave or edit) the target nucleic acid.In some embodiments, the invention provides a V-type Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) system comprising: (a) a Cas12a effector protein; (b) a recruitment guide nucleic acid comprising a guide nucleic acid linked to an RNA recruitment motif; and (c) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif, wherein the recruitment guide nucleic acid comprises a spacer sequence and a repeat sequence, and the guide nucleic acid is capable of forming a complex with the Cas12a effector protein, and the spacer sequence is capable of hybridizing to a target nucleic acid, thereby guiding the Cas12a effector protein to the target nucleic acid, and wherein the deaminase fusion protein is recruited to the Cas12a effector protein and the target nucleic acid by binding of the affinity polypeptide to the RNA recruitment motif fused to the recruitment guide nucleic acid, thereby enabling the system to modify (e.g., cleave or edit) the target nucleic acid.
[0091] In some embodiments, a nucleic acid construct of the present invention (e.g., a polynucleotide encoding a Type V CRISPR-Cas effector protein, a polynucleotide encoding a Type V CRISPR-Cas fusion protein, a polynucleotide encoding a deaminase, a polynucleotide encoding a deaminase fusion protein, a polynucleotide encoding a peptide tag, a polynucleotide encoding an affinity polypeptide, an RNA recruitment motif, a recruitment guide nucleic acid and / or a guide nucleic acid and / or an expression cassette and / or vector comprising the same) may be operably linked to at least one regulatory sequence, wherein the at least one regulatory sequence may be codon-optimized for expression in plants. In some embodiments, the at least one regulatory sequence may be, for example, a promoter, an operon, a terminator, or an enhancer. In some embodiments, the at least one regulatory sequence may be a promoter. In some embodiments, the regulatory sequence may be an intron. In some embodiments, the at least one regulatory sequence may be, for example, a promoter operably associated with an intron or an intron-containing promoter region. In some embodiments, the at least one regulatory sequence may be, for example, a ubiquitin promoter and its associated introns (e.g., Medicago truncatula and / or Zea mays and their associated introns). In some embodiments, the at least one regulatory sequence may be a terminator nucleotide sequence and / or an enhancer nucleotide sequence.
[0092] In some embodiments, a nucleic acid construct of the invention may be operably associated with a promoter region, where the promoter region comprises an intron, and the promoter region may be a ubiquitin promoter and intron (e.g., an alfalfa or maize ubiquitin promoter and intron, e.g., SEQ ID NO: 1 or SEQ ID NO: 2). In some embodiments, a nucleic acid construct of the invention operably associated with an intron-containing promoter region may be codon-optimized for expression in plants.
[0093] In some embodiments, the nucleic acid constructs of the present invention may further encode one or more polypeptides of interest, which may be codon-optimized for expression in plants.
[0094] Polypeptides of interest useful in the present invention include, but are not limited to, polypeptides or protein domains having deaminase activity, nickase activity, recombinase activity, transposase activity, methylase activity, glycosylase (DNA glycosylase) activity, glycosylase inhibitor activity (e.g., uracil DNA glycosylase inhibitor (UGI)), demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, restriction endonuclease activity (e.g., Fok1), nucleic acid binding activity, methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, polymerase activity, ligase activity, helicase activity, nuclear localization sequence or activity, and / or photolyase activity. In some embodiments, the polypeptide of interest is a Fok1 nuclease or a uracil DNA glycosylase inhibitor. When encoded in a nucleic acid (polynucleotide, expression cassette, and / or vector), the encoded polypeptide or protein domain may be codon-optimized for expression in an organism. In some embodiments, the polypeptide of interest may be linked to a type V CRISPR-Cas effector protein, providing a type V CRISPR-Cas fusion protein comprising the type V CRISPR-Cas effector protein and the polypeptide of interest. In some embodiments, a type V CRISPR-Cas fusion protein comprising a type V CRISPR-Cas effector protein linked to a peptide tag or affinity polypeptide may be linked to the polypeptide of interest (e.g., the type V CRISPR-Cas effector protein may be linked to both, e.g., a peptide tag (or affinity polypeptide) and, e.g., a polypeptide of interest, e.g., UGI). In some embodiments, the polypeptide of interest may be a uracil glycosylase inhibitor (e.g., a uracil DNA glycosylase inhibitor (UGI)).
[0095] In some embodiments, nucleic acid constructs of the invention that encode a type V CRISPR-Cas fusion protein, a deaminase fusion protein, and include a guide nucleic acid may further encode a polypeptide of interest, which may be codon-optimized for expression in an organism. In some embodiments, nucleic acid constructs of the invention that encode a type V CRISPR-Cas effector protein, a deaminase fusion protein, and include a recruitment guide nucleic acid may further encode a polypeptide of interest, which may be codon-optimized for expression in an organism (e.g., a plant).
[0096] As used herein, a "Type V CRISPR-Cas effector protein" or "Type V Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) effector protein" is a protein or polypeptide or domain thereof of a Type V CRISPR-Cas system that cleaves, cuts, or nicks nucleic acid, binds to nucleic acid (e.g., target nucleic acid and / or guide nucleic acid), and / or identifies, recognizes, or binds to a guide nucleic acid as defined herein. In some embodiments, a Type V CRISPR-Cas effector protein may be or be part of an enzyme (e.g., a nuclease, endonuclease, nickase, etc.) and / or may act as an enzyme. In some embodiments, the term "V-type CRISPR-Cas effector protein" refers to a V-type CRISPR-Cas nuclease polypeptide or a domain comprising nuclease activity or a domain with reduced or eliminated nuclease activity, and / or a domain comprising nickase activity or a domain with reduced or eliminated nickase activity, and / or a domain comprising single-strand DNA cleavage activity (ssDNAse activity) or a domain with reduced or eliminated ssDNAse activity, and / or a domain comprising self-processing RNAse activity or a domain with reduced or eliminated self-processing RNAse activity. The V-type CRISPR-Cas effector protein can bind to a target nucleic acid. In some embodiments, the V-type CRISPR-Cas effector protein can be a Cas12 effector protein.
[0097] In some embodiments, type V CRISPR-Cas effector proteins useful in the present invention may contain a mutation within their nuclease active site (e.g., RuvC, HNH, e.g., the RuvC site of a Cas12a nuclease domain). Type V CRISPR-Cas effector proteins that have a mutation within their nuclease active site and therefore no longer contain nuclease activity are commonly referred to as "dead" (e.g., dCas12a). In some embodiments, type V CRISPR-Cas effector proteins that have a mutation within their nuclease active site may have impaired or reduced activity (e.g., nickase, such as Cas12a nickase) compared to the same type V CRISPR-Cas effector protein without the mutation.
[0098] Type V CRISPR-Cas effector proteins useful in embodiments of the present invention can be any type V CRISPR-Cas nuclease. Type V CRISPR-Cas nucleases useful in the present invention as effector proteins can include, but are not limited to, Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c nucleases. In some embodiments, type V CRISPR-Cas nuclease polypeptides or domains useful in embodiments of the present invention can be Cas12a polypeptides or domains. In some embodiments, Type V CRISPR-Cas effector proteins useful in embodiments of the present invention may be nickases, and may be Cas12a nickases.
[0099] In some embodiments, the type V CRISPR-Cas effector protein can be a Cas12a effector protein. Cas12a differs from the more well-known type II CRISPR Cas9 in several respects. For example, Cas9 recognizes a G-rich protospacer adjacent motif (PAM) (3'-NGG) that is 3' to its guide RNA (gRNA, sgRNA, crRNA, crDNA, CRISPR array) binding site (protospacer, target nucleic acid, target DNA), while Cas12a recognizes a T-rich PAM (5'-TTN, 5'-TTTN) located 5' to the target nucleic acid. In fact, the binding orientations of Cas9 and Cas12a to their guide RNAs with respect to their N- and C-termini are almost opposite. Furthermore, the Cas12a enzyme uses a single guide RNA (gRNA, CRISPR array, crRNA) rather than the dual guide RNAs (sgRNA (e.g., crRNA and tracr RNA)) found in the native Cas9 system, and Cas12a processes its own gRNA. Additionally, Cas12a nuclease activity generates protruding DNA double-strand breaks instead of the blunt ends generated by Cas9 nuclease activity, and Cas12a relies on a single RuvC domain to cleave both DNA strands, whereas Cas9 utilizes an HNH domain and a RuvC domain for cleavage.
[0100] A type V CRISPR-Cas effector protein may be a CRISPR-Cas12a polypeptide or CRISPR-Cas12a domain derived from any known or later identified Cas12a (formerly known as Cpf1) (see, e.g., U.S. Patent No. 9,790,490, which is incorporated by reference for its disclosure of Cpf1 (Cas12a) sequences). The terms "Cas12a," "Cas12a polypeptide," or "Cas12a domain" refer to an RNA-guided polypeptide comprising a Cas12a polypeptide, or a fragment thereof, including the guide nucleic acid binding domain of Cas12a and / or the active, inactive, or partially active DNA cleavage domain of Cas12a, and / or the RNA-guided polypeptide may have nuclease activity. In some embodiments, a Cas12a useful in the present invention may contain a mutation within a nuclease active site (e.g., the RuvC site of a Cas12a domain). A Cas12a domain or Cas12a polypeptide that has a mutation in its nuclease active site and therefore no longer contains nuclease activity is commonly referred to as a deadCas12a (e.g., dCas12a). In some embodiments, a Cas12a domain or Cas12a polypeptide that has a mutation in its nuclease active site may have impaired activity (e.g., impaired nickase activity).
[0101] In some embodiments, a V-type CRISPR-Cas effector protein (e.g., a Cas12a polypeptide) may be optimized for expression in an organism, such as an animal, plant, fungus, archaea, or bacterium. In some embodiments, a V-type CRISPR-Cas effector protein (e.g., a Cas12a polypeptide) may be optimized for expression in a plant.
[0102] Any deaminase or domain thereof or polypeptide useful for base editing may be used in the present invention. As used herein, "cytosine deaminase" and "cytidine deaminase" refer to a polypeptide or domain thereof that catalyzes or is capable of catalyzing cytosine deamination, in that the polypeptide or domain catalyzes or is capable of catalyzing the removal of an amine group from a cytosine base. Thus, cytosine deaminase can cause the conversion of cytosine to thymidine (through a uracil intermediate), resulting in a C to T conversion or a G to A conversion in the complementary strand within the genome. Thus, in some embodiments, the cytosine deaminase encoded by the polynucleotide of the present invention causes a C to T conversion in the sense (e.g., "+"; template) strand of the target nucleic acid or a G to A conversion in the antisense (e.g., "-", complementary) strand of the target nucleic acid. In some embodiments, the cytosine deaminase encoded by the polynucleotide of the present invention results in the conversion of a C to a T, G, or A in the complementary strand within the genome.
[0103] Cytosine deaminases useful in the present invention may be any known or later-identified cytosine deaminases from any organism (see, e.g., U.S. Pat. No. 10,167,457 and Thuronyi et al. Nat. Biotechnol. 37:1070-1079 (2019), each of which is incorporated herein by reference for its disclosure of cytosine deaminases). Cytosine deaminases can catalyze the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. Thus, in some embodiments, deaminases or deaminase domains useful in the present invention may be cytidine deaminase domains capable of catalyzing the hydrolytic deamination of cytosine to uracil. In some embodiments, the cytosine deaminase may be a variant of a naturally occurring cytosine deaminase, including but not limited to, primate (e.g., human, monkey, chimpanzee, gorilla), dog, cow, rat, or mouse. Thus, in some embodiments, cytosine deaminases useful in the present invention may be about 70% to about 100% identical to a wild-type cytosine deaminase (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to a naturally occurring cytosine deaminase, and any range or value therein).
[0104] In some embodiments, a cytosine deaminase useful in the present invention may be an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the cytosine deaminase may be APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, APOBEC3H deaminase, APOBEC4 deaminase, human activation-induced deaminase (hAID), rAPOBEC1, FERNY, and / or CDA1, and may also be pmCDA1, atCDA1 (e.g., At2g19570), and evolved versions thereof. Evolved deaminases are disclosed, for example, in U.S. Pat. No. 10,113,163, Gaudelli et al. Nature 551(7681):464-471 (2017)) and Thuronyi et al. (Nature Biotechnology 37:1070-1079 (2019)), each of which is incorporated by reference herein for their disclosure of deaminases and evolved deaminases. In some embodiments, the cytosine deaminase may be an APOBEC1 deaminase, which may have the amino acid sequence of SEQ ID NO: 23. In some embodiments, the cytosine deaminase may be an APOBEC3A deaminase, which may have the amino acid sequence of SEQ ID NO: 24. In some embodiments, the cytosine deaminase may be a CDA1 deaminase, which may have the amino acid sequence of SEQ ID NO: 25. In some embodiments, the cytosine deaminase may be FERNY deaminase, which may be FERNY having the amino acid sequence of SEQ ID NO: 26. In some embodiments, the cytosine deaminase may be rAPOBEC1 deaminase, which may be rAPOBEC1 deaminase having the amino acid sequence of SEQ ID NO: 27.In some embodiments, the cytosine deaminase may be a hAID deaminase, which may be a hAID having the amino acid sequence of SEQ ID NO:28 or SEQ ID NO:29.
[0105] In some embodiments, cytosine deaminases useful in the invention can be about 70% to about 100% identical (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% identical) to the amino acid sequence of a naturally occurring cytosine deaminase (e.g., an "evolved deaminase") (see, e.g., SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32). In some embodiments, a cytosine deaminase useful in the invention may be about 70% to about 99.5% identical (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical) to the amino acid sequence of any one of SEQ ID NOs: 23-32 (e.g., at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of SEQ ID NOs: 23-32). In some embodiments, a polynucleotide encoding a cytosine deaminase may be codon-optimized for expression in an organism (e.g., a plant), and the codon-optimized polypeptide may be about 70% to 99.5% identical to the reference polynucleotide.
[0106] As used herein, "adenine deaminase" and "adenosine deaminase" refer to a polypeptide or domain thereof that catalyzes or is capable of catalyzing the hydrolytic deamination of adenine or adenosine (e.g., removal of an amine group from adenine). In some embodiments, adenine deaminase can catalyze the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, adenosine deaminase can catalyze the hydrolytic deamination of adenine or adenosine in DNA. In some embodiments, adenine deaminase encoded by a nucleic acid construct of the present invention can cause an A→G conversion in the sense (e.g., "+"; template) strand of a target nucleic acid or a T→C conversion in the antisense (e.g., "-", complementary) strand of a target nucleic acid. Adenine deaminases useful in the present invention can be any known or later identified adenine deaminase from any organism (see, e.g., U.S. Pat. No. 10,113,163, incorporated herein by reference for its disclosure of adenine deaminases).
[0107] In some embodiments, the adenosine deaminase can be a variant of a naturally occurring adenine deaminase. Thus, in some embodiments, the adenosine deaminase can be about 70% to 100% identical to a wild-type adenine deaminase (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to a naturally occurring adenine deaminase, and any range or value therein). In some embodiments, the deaminase or deaminase is not naturally occurring and may be referred to as an engineered, mutated, or evolved adenosine deaminase. Thus, for example, an engineered, mutated, or evolved adenine deaminase polypeptide or adenine deaminase may be about 70% to 99.9% identical to a naturally occurring adenine deaminase polypeptide (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 1109%, 1110%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 149%, 150%, 151%, 152%, 153%, 154%, 155%, %, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical, and any range or value therein. In some embodiments, the adenosine deaminase may be derived from bacteria (e.g., Escherichia coli, Staphylococcus aureus, Haemophilus influenzae, Caulobacter crescentus, etc.). In some embodiments, the polynucleotide encoding the adenine deaminase polypeptide may be codon-optimized for expression in plants.
[0108] In some embodiments, the adenine deaminase is a wild-type tRNA-specific adenosine deaminase, e.g., tRNA-specific adenosine deaminase (TadA), and / or a mutated / evolved adenosine deaminase, e.g., a mutated / evolved tRNA-specific adenosine deaminase (TadA). * ). In some embodiments, TadA may be derived from E. coli. In some embodiments, TadA may be modified, e.g., truncated, and may have one or more N-terminal and / or C-terminal amino acids deleted relative to full-length TadA (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal and / or C-terminal amino acid residues deleted compared to full-length TadA. In some embodiments, the TadA polypeptide or TadA domain does not include an N-terminal methionine. In some embodiments, wild-type E. coli TadA comprises the amino acid sequence of SEQ ID NO: 33. In some embodiments, mutated / evolved E. coli TadA * comprises the amino acid sequence of SEQ ID NO: 34-37 (e.g., SEQ ID NO: 34, 35, 36, or 37). In some embodiments, TadA / TadA * The polynucleotide encoding may be codon-optimized for expression in plants. In some embodiments, the adenine deaminase may comprise all or part of the amino acid sequence of any one of SEQ ID NOs: 33-43.
[0109] A "uracil glycosylase inhibitor" or "UGI" useful in the present invention can be any protein or polypeptide, or domain thereof, capable of inhibiting the uracil DNA glycosylase base excision repair enzyme. In some embodiments, UGI comprises wild-type UGI or a fragment thereof. In some embodiments, UGI useful in the present invention can be about 70% to about 100% identical (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% identical, and any range or value therein), to the amino acid sequence of a naturally occurring UGI. In some embodiments, the UGI may comprise the amino acid sequence of SEQ ID NO:44 or a polypeptide having about 70% to about 99.5% identity to the amino acid sequence of SEQ ID NO:44 (e.g., at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to the amino acid sequence of SEQ ID NO:44). For example, in some embodiments, the UGI may comprise a fragment of the amino acid sequence of SEQ ID NO: 44 that is 100% identical to a portion of the contiguous nucleotides of the amino acid sequence of SEQ ID NO: 44 (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 contiguous nucleotides; e.g., about 10, 15, 20, 25, 30, 35, 40, 45, to about 50, 55, 60, 65, 70, 75, 80 contiguous nucleotides). In some embodiments, the UGI may be a variant of a known UGI (e.g., SEQ ID NO: 44) having about 70% to about 99.5% identity to the known UGI (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% identity, and any range or value therein).In some embodiments, a polynucleotide encoding a UGI may be codon-optimized for expression in a plant (e.g., a plant), and the codon-optimized polypeptide may be about 70% to about 99.5% identical to the reference polynucleotide.
[0110] In some embodiments, nucleic acid constructs of the present invention, e.g., a type V CRISPR-Cas fusion protein comprising a type V CRISPR-Cas effector protein fused to a peptide tag, a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the peptide tag, and a guide nucleic acid; or a type V CRISPR-Cas fusion protein comprising a type V CRISPR-Cas effector protein, a recruitment guide nucleic acid comprising a guide nucleic acid linked to an RNA recruitment motif, and a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif, may further comprise / encode a polypeptide of interest. In some embodiments, the polypeptide of interest may be a uracil glycosylase inhibitor (UGI) (e.g., uracil DNA glycosylase inhibitor) polypeptide or domain, where the UGI may be codon-optimized for expression in an organism (e.g., a plant). In some embodiments, the glycosylase inhibitor may be fused to the type V CRISPR-Cas effector protein and / or the deaminase. In some embodiments, glycosylase inhibitors can be recruited to Type V CRISPR-Cas effector proteins and / or deaminases via the methods and constructs disclosed herein for protein-protein, protein-RNA, and / or chemical recruitment. Thus, by way of example, glycosylase inhibitors can be recruited to Type V CRISPR-Cas effector proteins and / or deaminases utilizing peptide tags / affinity polypeptides, RNA recruitment motifs / affinity polypeptides, and / or biotin-streptavidin interactions (or other chemical interactions) described herein.
[0111] The nucleic acid constructs of the present invention comprising a Type V CRISPR-Cas effector protein or a fusion protein thereof can be used in combination with a guide nucleic acid (e.g., gRNA, CRISPR array, CRISPR RNA, crRNA) or recruited guide nucleic acid designed to function with the encoded Type V CRISPR-Cas effector protein to modify a target nucleic acid. Guide nucleic acids and / or recruited guide nucleic acids useful in the present invention may comprise at least one spacer sequence and at least one repeat sequence. The guide nucleic acid and recruited guide nucleic acid are capable of forming a complex with a Type V CRISPR-Cas effector protein of the invention (e.g., a Type V CRISPR-Cas effector protein encoded and expressed by a nucleic acid construct of the invention), and the spacer sequence is capable of hybridizing to the target nucleic acid, thereby guiding the complex (e.g., the Type V CRISPR-Cas effector protein to the target nucleic acid), whereby the target nucleic acid can be modified (e.g., cleaved or edited) and / or modulated (e.g., transcriptionally modulated), optionally by a deaminase (e.g., a cytosine deaminase and / or an adenine deaminase that may be present in and / or recruited to the complex).
[0112] In some embodiments, a nucleic acid construct encoding a deaminase fusion protein comprising a type V CRISPR-Cas effector protein (e.g., Cas12a, Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c) fused to a peptide tag (e.g., a type V CRISPR-Cas effector fusion protein), and a deaminase fused to an affinity polypeptide that binds the peptide tag, is used to transfect a type V CRISPR-Cas effector protein (e.g., a type V CRISPR-Cas effector fusion protein) with a deaminase fused to an affinity polypeptide that binds the peptide tag. A target nucleic acid can be modified using in combination with a RISPR-Cas guide nucleic acid, where the guide nucleic acid binds to the target nucleic acid and guides a Type V CRISPR-Cas effector protein to the target nucleic acid, and the deaminase fusion protein is recruited to the Type V CRISPR-Cas effector protein and then recruited to the target nucleic acid via binding of the affinity polypeptide of the deaminase to the peptide tag of the Type V CRISPR-Cas effector protein, thereby enabling the deaminase of the deaminase fusion protein to deaminate cytosine bases in the target nucleic acid, thereby modifying (e.g., editing) the target nucleic acid.
[0113] In some embodiments, a nucleic acid construct encoding a deaminase fusion protein comprising a Type V CRISPR-Cas effector protein (e.g., Cas12a, Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c) and a deaminase fusion protein fused to an affinity polypeptide that binds to an RNA recruitment motif is used to recruit a CRISPR-Cas effector protein to an RNA. A target nucleic acid can be modified using in combination with a recruitment guide nucleic acid comprising a guide nucleic acid linked to a recruitment motif, where the recruitment guide nucleic acid binds to the target nucleic acid and guides a Type V CRISPR-Cas effector protein to the target nucleic acid, and a deaminase fusion protein is recruited to the target nucleic acid via binding of an affinity polypeptide to the RNA recruitment motif of the recruitment guide nucleic acid, thereby enabling the deaminase of the deaminase fusion protein to deaminate cytosine bases in the target nucleic acid, thereby modifying (e.g., editing) the target nucleic acid.
[0114] As used herein, the terms "guide nucleic acid," "guide RNA," "gRNA," "CRISPR RNA / DNA," "crRNA," or "crDNA" refer to a nucleic acid comprising at least one spacer sequence that is complementary to (and hybridizes with) a target DNA (e.g., a protospacer) and at least one repeat sequence (e.g., a repeat of a V-type CRISPR-Cas system, or a fragment or portion thereof, including but not limited to, Cas12a, Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c, or fragments thereof), wherein the repeat sequence can be linked to the 5' and / or 3' end of the spacer sequence. The design of gRNAs of the present invention is based on the Type V CRISPR-Cas system. In some embodiments, the guide nucleic acid comprises DNA. In some embodiments, the guide nucleic acid comprises RNA. In some embodiments, Type V CRISPR-Cas effector proteins, such as Cas12 agRNAs, may comprise a repeat sequence (e.g., a full-length repeat sequence or a portion thereof ("handle"); e.g., a pseudoknot-like structure) and a spacer sequence from 5' to 3'.
[0115] As used herein, a "recruitment guide nucleic acid" or "recruitment guide RNA" refers to a guide nucleic acid as defined herein that includes an RNA recruitment motif. In some embodiments, the RNA recruitment motif may be linked to the 3' or 5' end of the recruitment guide nucleic acid. In some embodiments, the RNA recruitment motif may be inserted within the recruitment guide nucleic acid (e.g., within a hairpin loop). An RNA recruitment motif useful in the present invention may be any RNA motif that can be recognized by an affinity polypeptide, for example, the RNA recruitment motif is capable of being bound by an affinity polypeptide. RNA recruitment motifs and their corresponding affinity polypeptides may include, but are not limited to, telomerase Ku-binding motifs (e.g., Ku-binding hairpins) and Ku affinity polypeptides (e.g., Ku heterodimers); telomerase Sm7-binding motifs and Sm7 affinity polypeptides; MS2 phage operator stem-loop and MS2 Coat protein (MCP) affinity polypeptides; PP7 phage operator stem-loop and PP7 Coat protein (PCP) affinity polypeptides; SfMu phage Com stem-loop and Com RNA-binding protein affinity polypeptides; PUF binding sites (PBSs) and corresponding Pumilio / fem-3 mRNA-binding factors (PUFs); and / or synthetic RNA aptamers and corresponding aptamer ligands. See, e.g., WO2018 / 129129 and U.S. Patent Application Publication Nos. 20190218261 and 20180094257, each of which is incorporated herein by reference for its disclosure of RNA recruitment motifs.
[0116] In some embodiments, the recruitment guide nucleic acid may be linked to one RNA recruitment motif or two or more RNA recruitment motifs (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10 or more copies of the RNA recruitment motif; e.g., about 2 to about 5, about 2 to about 8, about 2 to about 10, about 3 to about 5, about 3 to about 8, about 5 to about 8, about 5 to about 10, etc. recruitment motifs), where the two or more RNA recruitment motifs may be the same or different RNA recruitment motifs. Exemplary RNA recruitment motifs and corresponding affinity polypeptides that may be useful in the present invention include, but are not limited to, SEQ ID NOs: 45-55.
[0117] In some embodiments, the guide nucleic acid and / or recruitment guide nucleic acid may comprise more than one repeat sequence-spacer sequence (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more repeat-spacer sequences) (e.g., repeat-spacer-repeat, e.g., repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer, etc.). The guide nucleic acids or recruitment guide nucleic acids of the invention are synthetic, man-made, and not found in nature. gRNAs can be quite long and can be used as aptamers (as in the MS2 recruitment strategy) or other RNA structures hanging off of spacers (e.g., recruitment guide nucleic acids).
[0118] As used herein, "repeat sequence" refers to, for example, any repeat sequence of a wild-type V-type CRISPR Cas locus (e.g., the Cas12a locus, the Cas12b locus, the Cas12c locus (C2c3), the Cas12d locus (CasY), the Cas12e locus (CasX), the Cas12g locus, the Cas12h locus, the Cas12i locus, the C2c1 locus, the C2c4 locus, the C2c5 locus, the C2c8 locus, the C2c9 locus, the C2c10 locus, the Cas14a locus, the Cas14b locus, and / or the Cas14c locus, or fragments thereof), or a repeat sequence of a synthetic crRNA that is functional with a V-type CRISPR-Cas effector protein encoded by a nucleic acid construct of the present invention. Repeat sequences useful in the present invention can be any known or later identified repeat sequence of a V-type CRISPR-Cas locus, or can be synthetic repeats designed to be functional in a V-type CRISPR-Cas system. The repeat sequence may comprise a hairpin and / or stem-loop structure. In some embodiments, the repeat sequence may form a pseudoknot-like structure at its 5' end (i.e., a "handle"). Thus, in some embodiments, the repeat sequence can be identical or substantially identical to a repeat sequence from a wild-type V-type CRISPR-Cas locus. Repeat sequences from wild-type CRISPR-Cas loci can be determined by established algorithms, for example, using CRISPRfinder provided by CRISPRdb (see Grissa et al. Nucleic Acids Res. 35 (Web Server issue): W52-7). In some embodiments, the repeat sequence, or a portion thereof, may be linked at its 3' end to the 5' end of a spacer sequence, thereby forming a repeat-spacer sequence (e.g., guide RNA, crRNA).
[0119] In some embodiments, the repeat sequence comprises, consists essentially of, or consists of at least 10 nucleotides (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50-100 or more nucleotides, or any range or value therein; e.g., about), depending on the particular repeat and regardless of whether the guide nucleic acid containing the repeat is processed. In some embodiments, the repeat sequence comprises, consists essentially of, or consists of about 10 to about 20, about 10 to about 30, about 10 to about 45, about 10 to about 50, about 10 to about 100, about 15 to about 30, about 15 to about 40, about 15 to about 45, about 15 to about 50, about 50 to about 100, about 20 to about 30, about 20 to about 40, about 20 to about 50, about 20 to about 100, about 30 to about 40, about 30 to about 50, about 30 to about 100, about 40 to about 80, about 40 to about 100, about 50 to about 100 or more nucleotides.
[0120] The repeat sequence linked to the 5' end of the spacer sequence can include a portion of the repeat sequence (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more consecutive nucleotides of the wild-type repeat sequence). In some embodiments, the portion of the repeat sequence linked to the 5' end of the spacer sequence can be about 5 to about 10 contiguous nucleotides in length (e.g., about 5, 6, 7, 8, 9, 10 nucleotides) and can have at least 90% sequence identity (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more) to an identical region (e.g., the 5' end) of a wild-type CRISPR Cas repeat nucleotide sequence. In some embodiments, the portion of the repeat sequence can include a pseudoknot-like structure (e.g., a "handle") at its 5' end.
[0121] As used herein, a "spacer sequence" is a nucleotide sequence that is complementary to a target nucleic acid (e.g., a target DNA) (e.g., a protospacer). A spacer sequence can be fully complementary or substantially complementary (e.g., at least about 70% complementary (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more)) to a target nucleic acid. Thus, in some embodiments, a spacer sequence can have 1, 2, 3, 4, or 5 mismatches compared to a target nucleic acid, and the mismatches can be contiguous or non-contiguous. In some embodiments, the spacer sequence can have 70% complementarity to the target nucleic acid. In other embodiments, the spacer nucleotide sequence can have 80% complementarity to the target nucleic acid. In still other embodiments, the spacer nucleotide sequence can have 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% complementarity to the target nucleic acid (protospacer), etc. In some embodiments, the spacer sequence is 100% complementary to the target nucleic acid. The spacer sequence can have a length of about 15 nucleotides to about 30 nucleotides (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides, or any range or value therein). Thus, in some embodiments, a spacer sequence may have perfect or substantial complementarity over a region of the target nucleic acid (e.g., a protospacer) that is at least about 15 to about 30 nucleotides in length. In some embodiments, the spacer is about 20 nucleotides in length. In some embodiments, the spacer is about 21, 22, or 23 nucleotides in length.
[0122] In some embodiments, the 5' region of the spacer sequence of the guide nucleic acid can be identical to the target DNA, while the 3' region of the spacer can be substantially complementary to the target DNA (e.g., V-type CRISPR-Cas), or the 3' region of the spacer sequence of the guide nucleic acid can be identical to the target DNA, such that the overall homology of the spacer sequence to the target DNA is less than 100%. Thus, for example, in a guide for a V-type CRISPR-Cas system, for example, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in the 5' region of a 20-nucleotide spacer sequence (i.e., the seed region) can be 100% complementary to the target DNA, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary (e.g., at least about 70% complementary) to the target DNA. In some embodiments, the first 1 to 8 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8 nucleotides, and any range therein) at the 5' end of the spacer sequence can be 100% complementary to the target DNA, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary to the target DNA (e.g., at least about 50% complementary (e.g., 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more)). The recruitment guide nucleic acid further comprises one or more recruitment motifs described herein, which can be linked to the 5' end of the guide nucleic acid, and the 3' end of the guide nucleic acid or one or more RNA recruitment motif(s) can be inserted within the recruitment guide nucleic acid (e.g., within a hairpin loop).
[0123] In some embodiments, the seed region of the spacer can be about 8 to about 10 nucleotides in length, about 5 to about 6 nucleotides in length, or about 6 nucleotides in length.
[0124] As used herein, "target nucleic acid," "target DNA," "target nucleotide sequence," "target region," or "target region within a genome" refers to a region of an organism's genome that is fully complementary (100% complementary) or substantially complementary (e.g., at least 70% complementary (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more)) to a spacer sequence in a guide nucleic acid of the invention. Target regions (known as protospacer adjacent motifs (PAMs)) useful for Type V CRISPR-Cas systems may be located adjacent to the spacer (or target) sequence. These PAM DNA sequences are typically described by referring to their sequence and location relative to the non-target strand of CRISPR complex.These PAM sequences can be 3' (for example, type V CRISPR-Cas system) or 5' (for example, type II CRISPR-Cas system) relative to the end of the protospacer sequence.The target region (also referred to as protospacer) can be selected from any region of at least 15 consecutive nucleotides (for example, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides, etc.) that is located immediately adjacent to the PAM sequence.
[0125] "Protospacer sequence" refers to the target double-stranded DNA, and specifically to the portion of the target DNA (e.g., or target region within a genome) that is perfectly or substantially complementary to (and hybridizes with) the spacer sequence of a CRISPR repeat-spacer sequence (e.g., guide nucleic acid, CRISPR array, crRNA).
[0126] In the case of V-type CRISPR-Cas (e.g., Cas12a) systems, the protospacer sequence is flanked (immediately adjacent) to a protospacer adjacent motif (PAM). For V-type CRISPR-Cas systems, the PAM is located at the 5' end of the non-target strand and the 3' end of the target strand (see below for an example). JPEG2025157268000002.jpg31166
[0127] The guide structure and PAM are described by R. Barrangou (Genome Biol. 16:247 (2015)).
[0128] Standard V-type CRISPR-Cas12a PAMs are T-rich. In some embodiments, the standard Cas12a PAM sequence can be 5'-TTN, 5'-TTTN, or 5'-TTTV. In some embodiments, non-standard PAMs can be used, but may be less efficient.
[0129] Additional PAM sequences can be determined by those skilled in the art through proven experimental and computational methods.Thus, for example, experimental methods include targeting the sequence flanked by all possible nucleotide sequences, and identifying the sequence members that are not subjected to targeting, for example, by transforming target plasmid DNA (Esvelt et al.2013.Nat.Methods10:1116-1121; Jiang et al.2013.Nat.Biotechnol.31:233-239). In some embodiments, computational methods can include performing a BLAST search of natural spacers to identify the original target DNA sequence in a bacteriophage or plasmid, and aligning these sequences to determine conserved sequences flanking the target sequence (Briner and Barrangou. 2014. Appl. Environ. Microbiol. 80:994-1001; Mojica et al. 2009. Microbiology 155:733-740).
[0130] The "peptide tags" described herein can be used to recruit one or more polypeptides. A peptide tag can be any polypeptide capable of being bound by a corresponding affinity polypeptide. A peptide tag can also be referred to as an "epitope," and when provided in multiple copies, as a "multimerization epitope." Examples of peptide tags include, but are not limited to, a GCN4 peptide tag (e.g., Sun-Tag), a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or strep tag II, a V5 tag, and / or a VSV-G epitope. In some embodiments, a peptide tag can also include a phosphorylated tyrosine in a specific sequence context recognized by an SH2 domain, a characteristic consensus sequence containing phosphoserine recognized by 14-3-3 proteins, a proline-rich peptide motif recognized by an SH3 domain, a PDZ protein interaction domain or a PDZ signal sequence, and an AGO hook motif derived from plants. Peptide tags are disclosed in WO2018 / 136783 and U.S. Patent Application Publication No. 2017 / 0219596, which are incorporated by reference for their disclosure of peptide tags. Peptide tags that may be useful in the present invention may include, but are not limited to, SEQ ID NO: 59 and SEQ ID NO: 60. Affinity polypeptides useful in peptide tags include, but are not limited to, SEQ ID NO: 61.
[0131] Any epitope that can be linked to a polypeptide and for which there is a corresponding affinity polypeptide that can be linked to another polypeptide can be used as a peptide tag in the present invention. In some embodiments, the peptide tag can comprise one or more copies of the peptide tag (e.g., peptide repeat units, multimerization epitopes (e.g., tandem repeats)) (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more peptide tag(s)). In some embodiments, the affinity polypeptide that interacts / binds with the peptide tag can be an antibody. In some embodiments, the antibody can be an scFv antibody. In some embodiments, the affinity polypeptide that binds to the peptide tag may be synthetic (e.g., evolved affinity interactions), including, but not limited to, an affibody, anticalin, monobody, and / or DARPin (see, e.g., Sha et al., Protein Sci. 26(5):910-924(2017)); Gilbreth (Curr Opin Struc Biol 22(4):413-420(2013)); U.S. Patent No. 9,982,053, each of which is incorporated by reference in their entirety for teachings related to affibodies, anticalins, monobodies, and / or DARPins.
[0132] In some embodiments, the guide nucleic acid can be linked to an RNA recruitment motif, and the polypeptide to be recruited (e.g., a deaminase) can be fused to an affinity polypeptide that binds to the RNA recruitment motif, such that the guide nucleic acid binds to the target nucleic acid and the RNA recruitment motif binds to the affinity polypeptide, thereby recruiting the polypeptide to the guide nucleic acid and contacting the target nucleic acid with the polypeptide (e.g., a deaminase). In some embodiments, two or more polypeptides can be recruited to the guide nucleic acid, thereby contacting the target nucleic acid with two or more polypeptides (e.g., a deaminase).
[0133] In some embodiments, components for recruiting polypeptides and nucleic acids may include, but are not limited to, those that function through chemical interactions that may include rapamycin-induced dimerization of FRB-FKBP; biotin-streptavidin; SNAP tags; Halo tags; CLIP tags; compound-induced DmrA-DmrC heterodimers; and / or bifunctional ligands (e.g., fusions of two protein-binding chemicals; for example, dihydrofolate reductase (DHFR)).
[0134] The peptide tag may comprise or be present in one copy or in two or more copies of a peptide tag (e.g., a multimerizing peptide tag or multimerizing epitope) (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 9, 20, 21, 22, 23, 24, or 25 or more peptide tags). When multimerized, peptide tags can be fused directly to each other or linked to each other via one or more amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more amino acids, such as about 3 to about 10, about 4 to about 10, about 5 to about 10, about 5 to about 15, or about 5 to about 20 amino acids, and any value or range therein). Thus, in some embodiments, a Type V CRISPR-Cas fusion protein of the invention can comprise a Type V CRISPR-Cas effector protein fused to one peptide tag or to two or more peptide tags, where the two or more peptide tags can be fused to each other via one or more amino acid residues. In some embodiments, peptide tags useful in the present invention may be a single copy of a GCN4 peptide tag or epitope, or may be a multimerized GCN4 epitope comprising from about 2 to about 25 or more copies of the peptide tag (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more copies of a GCN4 epitope, or any range therein).
[0135] In some embodiments, the peptide tag can be fused to a type V CRISPR-Cas protein. In some embodiments, the peptide tag can be fused or linked to the C-terminus of a type V CRISPR-Cas effector protein to form a type V CRISPR-Cas fusion protein. In some embodiments, the peptide tag can be fused or linked to the N-terminus of a type V CRISPR-Cas effector protein to form a type V CRISPR-Cas fusion protein.
[0136] In some embodiments, when the peptide tag comprises more than one peptide tag, the amount and / or spacing of the epitopes may be optimized within the peptide tag to maximize peptide tag occupancy and minimize, for example, steric interference of the deaminase domains with each other.
[0137] An "affinity polypeptide" (e.g., a "recruitment polypeptide") refers to any polypeptide capable of binding to its corresponding peptide tag or RNA recruitment motif. An affinity polypeptide for a peptide tag can be, for example, an antibody and / or a single-chain antibody that specifically binds to the peptide tag. In some embodiments, an antibody for a peptide tag can be, but is not limited to, an scFv antibody. In some embodiments, an affinity polypeptide can be fused or linked to the N-terminus of a deaminase (e.g., cytosine deaminase or adenine deaminase) and recruit the deaminase to a recruitment guide nucleic acid or a type V CRISPR-Cas effector protein. In some embodiments, the affinity polypeptide is stable under reducing conditions in a cell or cell extract.
[0138] The nucleic acid constructs and / or guide nucleic acids and / or recruitment guide nucleic acids of the invention may be included within one or more expression cassettes as described herein, in some embodiments, the nucleic acid constructs of the invention may be included within the same or separate expression cassette or vector as that containing the guide nucleic acid and / or recruitment guide nucleic acid.
[0139] When used in combination with a guide nucleic acid and a recruiting guide nucleic acid, the nucleic acid constructs of the invention (and expression cassettes and vectors comprising same) can be used to modify a target nucleic acid and / or its expression. The target nucleic acid can be contacted with the nucleic acid constructs of the invention and / or expression cassettes and / or vectors comprising same before, simultaneously with, or after contacting the guide / recruiting guide nucleic acid (and / or expression cassettes and vectors comprising same) with the target nucleic acid.
[0140] The present invention further provides methods for modifying a target nucleic acid using the compositions, complexes (e.g., ribonucleocomplexes), systems, nucleic acid constructs, expression cassettes, and / or vectors of the present invention. The methods can be performed in an in vivo system (e.g., in a cell or organism) or in an in vitro system (e.g., cell-free). In some embodiments, the present invention provides methods for modifying a target nucleic acid, comprising contacting the target nucleic acid with (a) a Type V CRISPR-Cas effector protein; (b) a deaminase (the target nucleic acid may be contacted with two or more deaminases); and (c) a guide nucleic acid, wherein the deaminase is recruited to the Type V CRISPR-Cas effector protein (e.g., recruited via protein-protein interaction, RNA-protein interaction, and / or chemical interaction), wherein the Type V CRISPR-Cas effector protein, deaminase, and guide nucleic acid may be co-expressed, thereby modifying the target nucleic acid.
[0141] In some embodiments, methods of modifying a target nucleic acid are provided, comprising contacting the target nucleic acid with (a) a type V CRISPR-Cas fusion protein comprising a type V CRISPR-Cas effector protein fused to a peptide tag (e.g., an epitope or a multimerization epitope); (b) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the peptide tag (the target nucleic acid may be contacted with two or more deaminase fusion proteins); and (c) a guide nucleic acid, wherein the type V CRISPR-Cas fusion protein, deaminase fusion protein, and guide nucleic acid may be co-expressed, thereby modifying the target nucleic acid. In some embodiments, the peptide tag may be a single copy of a peptide tag or two or more copies (e.g., two or more epitopes) of a peptide tag, as described herein. In some embodiments, the peptide tag may be, for example, a GCN4 peptide tag (e.g., Sun-Tag) comprising from 1 to about 25 repeat units. In some embodiments, the affinity polypeptide may be an antibody. In some embodiments, multiple deaminases can be contacted with a target nucleic acid and recruited by a type V CRISPR-Case fusion protein. In some embodiments, a target nucleic acid can be contacted with more than one guide nucleic acid, which can include the same or different spacers and / or repeats from each other, thereby allowing targeting of different sites on the target nucleic acid and / or interacting with different type V CRISPR-Cas effector proteins (e.g., Cas12a, Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c).
[0142] In some embodiments, the present invention provides a method for modifying a target nucleic acid, comprising contacting the target nucleic acid with (a) a type V CRISPR-Cas effector protein; (b) a recruitment guide nucleic acid comprising a guide RNA linked to an RNA recruitment motif; and (c) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif, wherein the target nucleic acid may be contacted with two or more deaminase fusion proteins, and the type V CRISPR-Cas effector protein, the deaminase fusion protein, and the recruitment guide nucleic acid are co-expressed, thereby modifying the target nucleic acid. Any RNA recruitment motif(s) described herein may be used with the method of the present invention. In some embodiments, multiple deaminases may be contacted with the target nucleic acid and recruited by the recruitment guide nucleic acid. In some embodiments, a target nucleic acid may be contacted with more than one recruitment guide nucleic acid, which may comprise the same or different RNA recruitment motifs and / or spacers and / or repeats from each other, thereby allowing targeting of different sites on the target nucleic acid (e.g., 2, 3, 4, 5, or more different sites), allowing interaction with different Type V CRISPR-Cas effector proteins, and recruiting multiple polypeptides, which may be the same or different.
[0143] In some embodiments, deaminases useful for modifying target nucleic acids can be cytosine deaminases and / or adenine deaminases described herein. In some embodiments, the cytosine deaminase can be an apolipoprotein B mRNA editing catalytic polypeptide-like (APOBEC) domain, human activation-induced deaminase (hAID), FERNY deaminase, and / or CDA1 deaminase. In some embodiments, the APOBEC deaminase can be an APOBEC3A deaminase. In some embodiments, the adenine deaminase can be TadA (tRNA-specific adenosine deaminase) and / or TadA *(evolved tRNA-specific adenosine deaminase). In some embodiments, the methods of the invention may further comprise introducing / expressing a glycosylase inhibitor and / or a polynucleotide encoding a glycosylase inhibitor (e.g., uracil-DNA glycosylase inhibitor (UGI)), and the methods may comprise introducing or expressing two or more glycosylase inhibitors. In some embodiments, the glycosylase inhibitor may be fused to a Type V CRISPR-Cas effector protein and / or deaminase. In some embodiments, the glycosylase inhibitor may be recruited to a Type V CRISPR-Cas effector protein and / or deaminase via methods and constructs disclosed herein for protein-protein recruitment, protein-RNA recruitment, and / or chemical recruitment. Thus, by way of example, glycosylase inhibitors can be recruited to Type V CRISPR-Cas effector proteins and / or deaminases utilizing the peptide tags / affinity polypeptides, RNA recruitment motifs / affinity polypeptides and / or biotin-streptavidin interactions (or other chemical interactions) described herein.
[0144] The methods of the invention may comprise contacting a target nucleic acid with a CRISPR Cas effector protein, deaminase, and / or a fusion protein thereof and / or polypeptide of interest of the invention, or the target nucleic acid may be contacted with a polynucleotide encoding a CRISPR Cas effector protein, deaminase, and / or a fusion protein thereof and / or polypeptide of interest of the invention, which polypeptide may be comprised within one or more expression cassettes and / or vectors described herein, which expression cassettes and / or vectors may comprise one or more guide nucleic acids / recruiting guide nucleic acids.
[0145] As described herein, the nucleic acids of the present invention and / or expression cassettes and / or vectors comprising the same may be codon-optimized for expression in an organism. Organisms useful in the present invention may be any organism or cell thereof for which nucleic acid modification may be useful. Organisms may include, but are not limited to, any animal, any plant, any fungus, any archaea, or any bacterium. In some embodiments, the organism may be a plant or cell thereof.
[0146] The target nucleic acid of any plant or plant part can be modified using the nucleic acid construct of the present invention. Any plant (or grouping of plants into, for example, a genus or higher classification), including angiosperms, gymnosperms, monocotyledons, dicotyledons, C3, C4, CAM plants, bryophytes, ferns and / or ferns other than Pteridophytes, microalgae, and / or macroalgae, can be modified using the nucleic acid construct of the present invention. Plants and / or plant parts useful in the present invention can be plants and / or plant parts of any plant species / variety / cultivar. As used herein, the term "plant part" includes, but is not limited to, embryos, pollen, ovules, seeds, leaves, stems, buds, flowers, branches, fruits, grains, ears, cobs, husks, petioles, roots, root tips, anthers, plant cells (including intact plant cells in plants and / or plant parts, plant protoplasts, plant tissues, plant cell tissue cultures, plant calluses, plant clumps, etc.). As used herein, "shoot" refers to the part above ground, including leaves and stems. Furthermore, as used herein, "plant cell" refers to the structural and physiological unit of a plant, including the cell wall, and may also refer to a protoplast. A plant cell can be in the form of an isolated single cell, or can be a cultured cell, or can be part of a higher organized unit, such as a plant tissue or plant organ.
[0147] Non-limiting examples of plants useful in the present invention include turfgrasses (e.g., bluegrass, bentgrass, ryegrass, fescue), feather reed grass, tufted hair grass, miscanthus, arundo, switchgrass, vegetable crops including artichoke, kohlrabi, arugula, leeks, asparagus, lettuce (e.g., head, leaf, romaine), malanga, melons (e.g., muskmelon, watermelon, cranberry, honeydew, cantaloupe), brassica crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, collards, kale, Chinese cabbage, bok choy), cardoni, carrots, napa cabbage, and okra. , onion, celery, parsley, chickpea, parsnip, chicory, pepper, potato, cucurbits (e.g., cucumber, zucchini, eggplant, pumpkin, honeydew melon, watermelon, cantaloupe), radish, dry bulb onion, rutabaga, eggplant, burdock, endive, shallot, endive, garlic, spinach, leek, eggplant, leafy vegetables, beet (sugar beet and fodder beet), sweet potato, chard, horseradish, tomato, turnip, and spices;Fruit crops, such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, quince, figs, nuts (e.g., chestnuts, pecans, pistachios, hazelnuts, peanuts, walnuts, macadamia nuts, almonds, etc.), citrus fruits (e.g., clementines, kumquats, oranges, grapefruit, tangerines, mandarins, lemons, limes, etc.), blueberries, black raspberries, boysenberries, cranberries, currants, gooseberries, loganberries, raspberries, strawberries, blackberries, grapes (wine and table), avocados, bananas, kiwi, persimmons, pomegranates, pineapples, tropical fruits, pears fruit, melon, mango, papaya, and lychee, crop plants such as clover, alfalfa, timothy grass, evening primrose, meadowfoam, corn / maize (field, sweet, popcorn), hops, jojoba, buckwheat, safflower, quinoa, wheat, rice, barley, rye, millet, sorghum, oats, lychee, sorghum, tobacco, kapok, legumes (beans (e.g., green and dry), lentils, peas, soybeans), oil plants (rapeseed, canola, mustard, poppy, olives, sunflower, coconut, castor oil plant, cocoa beans, peanuts, oil palm), duckweed, Arabidopsis, fiber plants (cotton, flax, hemp, jute), hemp (e.g., Cannabis sativa), sativa, Cannabis indica, and Cannabis ruderalis), Lauraceae (cinnamon, camphor), or plants, such as coffee, sugarcane, tea, and rubber plants; and / or bedding plants, such as flowering plants, cacti, succulents, and / or ornamental plants (e.g., roses, tulips, violets), and trees, such as forest trees (broad-leaved and evergreen trees, e.g., conifers;For example, elm, ash, oak, maple, fir, spruce, cedar, pine, birch, cypress, eucalyptus, willow), as well as shrubs and other seedlings. In some embodiments, the nucleic acid constructs of the invention and / or expression cassettes and / or vectors encoding same may be used to modify corn, soybean, wheat, canola, rice, tomato, pepper, sunflower, bramble, blackberry, black raspberry, and / or cherry.
[0148] In some embodiments, the invention provides cells (e.g., plant cells, animal cells, bacterial cells, archaeal cells, etc.) comprising a polypeptide, polynucleotide, nucleic acid construct, expression cassette, or vector of the invention.
[0149] The present invention further includes a kit or kits for carrying out the methods of the present invention. The kits of the present invention may include reagents, buffers, and instruments for mixing, measuring, selecting, labeling, etc., as well as instructions suitable for modifying the target nucleic acid.
[0150] In some embodiments, the present invention provides kits comprising one or more nucleic acid constructs of the present invention described herein, and / or expression cassettes and / or vectors and / or cells comprising same, and may include instructions for their use. In some embodiments, the kits may further comprise CRISPR-Cas guide nucleic acids and / or recruitment guide nucleic acids (corresponding to CRISPR-Cas effector proteins encoded by polynucleotides of the present invention) and / or expression cassettes and / or vectors and / or cells comprising same. In some embodiments, the guide nucleic acids and / or recruitment guide nucleic acids may be provided on the same expression cassette and / or vector as one or more nucleic acid constructs of the present invention. In some embodiments, the guide nucleic acids and / or recruitment guide nucleic acids may be provided on a separate expression cassette or vector from that comprising one or more nucleic acid constructs of the present invention.
[0151] Thus, in some embodiments, kits are provided that include (a) a nucleic acid construct comprising a polynucleotide(s) provided herein and (b) a promoter that drives expression of the polynucleotide(s) of (a). In some embodiments, the kits may further include a nucleic acid construct encoding a guide nucleic acid and / or a recruitment guide nucleic acid, wherein the construct includes a cloning site for cloning a nucleic acid sequence identical or complementary to a target nucleic acid sequence into the backbone of the guide nucleic acid and / or the recruitment guide nucleic acid.
[0152] In some embodiments, a nucleic acid construct of the invention may be an mRNA, which may encode one or more introns within the encoded polynucleotide(s). In some embodiments, a nucleic acid construct of the invention, and / or an expression cassette and / or vector comprising same, may further encode one or more selectable markers useful for identifying transformants (e.g., nucleic acids encoding antibiotic resistance genes, herbicide resistance genes, etc.).
[0153] The present invention will now be described with reference to the following examples. It should be understood that these examples are not intended to limit the scope of the claimed invention, but rather are intended to be illustrative of particular embodiments. Any variations in the exemplified methods that occur to those skilled in the art are intended to be within the scope of the present invention. [Example]
[0154] Example 1. dCpf1 fused with SunTag for C to T base editing Eight copies of the GCN4 epitope were fused to the C-terminus of catalytically inactivated LbCpf1 (dLbCpf1) (dLbCas12a), as provided in SEQ ID NO: 62. In separate plasmids, antibodies targeted to the GCN4 epitope were fused to various deaminases (scFv deaminases), including rAPOBEC1, hAPOBEC3A, hAPOBEC3B, hAID, and pmCDA1 (SEQ ID NOs: 63-67).
[0155] Plasmids encoding dLbCpf1-SunTag, scFv-deaminase, UGI, and guide RNAs were transfected into HEK293T cells. UGI was co-expressed to temporarily suppress base excision repair during base editing events, thereby improving efficiency and product purity (Nishida et al. Science 353(6305)(2016)(DOI:10.1126 / science.aaf8729)). A total of four guide RNAs targeting different endogenous genes were used to examine C to T base editing (Table 1). After 3 days, genome editing was quantified by next-generation sequencing (NGS). While most deaminases tested showed variable and generally low C to T editing efficiencies, APOBEC3A consistently demonstrated high C to T editing rates, ranging from approximately 7% to approximately 19% of the treated cell population (unsorted transfected cells) (Figures 1-4). Fusion of SunTag to Cpf1 is also functional in APOBEC3A recruitment, as SunTag fusions provide up to five-fold higher editing efficiency compared to unfused dCpf1 (Figures 1-4). Efficient editing is observed within approximately positions 6-15 (positions -4, -3, -2, and -1 are designated PAM(TTTV)) (Figures 1-4).
[0156] JPEG2025157268000003.jpg71166
[0157] Figures 1-4 show that SunTag-fused dCpf1 can efficiently recruit a co-expressed deaminase domain in cells and can be used to generate significant levels of C-to-T base editing. This is the first demonstration of base editing using Cpf1 without the use of a fusion architecture. The multiple additional components described herein can be co-expressed either alone (expressed under a single promoter) or in any combination.
[0158] Example 2. dCpf1 fused with SunTag for A to G base editing Eight copies of the GCN4 epitope were fused to the C-terminus of catalytically inactivated LbCpf1 (dLbCpf1) (dLbCas12a) as provided in SEQ ID NO: 62. In a separate plasmid, an antibody targeted to the GCN4 epitope was fused to an evolved adenine deaminase (TadA8e) (scFv-deaminase) as provided in SEQ ID NO: 72.
[0159] Plasmids encoding dLbCpf1-SunTag, scFv-TadA8e (Richter et al. Nat Biotechnol. 2020, 38(7):883-891), and guide RNAs were transfected into HEK293T cells. A total of five different guide RNAs containing one of the spacer sequences (SEQ ID NOs: 73-77) were used to target endogenous genes (positions 1-5) and examine A-to-G base editing (Figure 5). After 3 days, genome editing was quantified by next-generation sequencing (NGS).
[0160] Significant adenine deamination was observed in the predicted targeting window around positions 8-10 (positions -4, -3, -2, -1 are designated as PAM(TTTV)).
[0161] The constructs and methods of the present invention are broadly applicable to many different types of systems, including in vitro and in vivo systems, and can utilize multiple different Type V CRISPR Cas effector proteins and multiple different deaminases, making them applicable in systems of multiple types of organisms, including animals (e.g., mammals) and plants. By multiplexing guides, editing could be simultaneously targeted to multiple different loci within a genome.
[0162] Example 3. SunTag-fused dCpf1 for C to T base editing in soybean plants T-DNA vectors containing expression cassettes for GCN4 epitope-tagged dCpf1 (LbCpf1 and EnAsCpf1), single-chain antibody-fused APOBEC3A (scFv-A3A), and uracil glycosylase inhibitor (UGI) were constructed. These components were expressed via a single promoter utilizing fusion and P2A linkers, or via multiple promoters driving the expression of individual components (Table 2). Specifically, the DaMV promoter was used to express the CRISPR-SunTag components, and the Mt.Ubq2 promoter was used to express the deaminase and UGI components (Table 2). The T-DNA vectors also contained an antibiotic selection cassette and guide RNA cassette driven by the U6 promoter containing sequences targeting either target #1 or target #2 within the target soybean gene. The T-DNA vectors were transformed into Agrobacteria and treated with soybean dried excised embryos to induce plant transformation. Leaves grown on antibiotic selection medium were sampled approximately 4 weeks after transformation, and their genetic composition was analyzed using Illumina high-throughput amplicon sequencing after DNA extraction.
[0163] Base editing activity (C to T changes in the target spacer sequence) was detected in all constructs tested at various rates, ranging from 20% to 88% (Table 2). For each construct, high-throughput sequencing of edited plants indicated that each plant sample contained, on average, 0.77% to 17.48% edited DNA per plant (Table 2). Editing was observed for both tested target genes (Target #1 and Target #2) and both tested CRISPR enzymes (LbCpf1 and EnAsCpf1) (Table 2). These results indicate that this method for base editing is suitable for other target genes and other type V CRISPR systems.
[0164] JPEG2025157268000004.jpg142166
[0165] The foregoing is illustrative of the present invention and is not to be construed as limiting thereof. The present invention is defined by the following claims, equivalents of which are included herein.
Claims
1. 1. A method for modifying a target nucleic acid, comprising: The method comprises: The target nucleic acid is: (a) V-shaped Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) effector proteins; (b) a deaminase, wherein the target nucleic acid may be contacted with two or more deaminases; and (c) guide nucleic acid contacting the wherein the deaminase is recruited to the Type V CRISPR-Cas effector protein (e.g., via protein-protein interactions, RNA-protein interactions, and / or chemical interactions), thereby modifying the target nucleic acid, and wherein the Type V CRISPR-Cas effector protein, the deaminase, and the guide nucleic acid may be co-expressed. method.
2. 1. A method for modifying a target nucleic acid, comprising: The method comprises: The target nucleic acid is: (a) a V-type Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) fusion protein comprising a V-type CRISPR-Cas effector protein fused to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the peptide tag; wherein the target nucleic acid may be contacted with two or more deaminase fusion proteins; and (c) guide nucleic acid contacting the wherein the Type V CRISPR-Cas fusion protein, the deaminase fusion protein and the guide nucleic acid may be co-expressed, thereby modifying the target nucleic acid. method.
3. 1. A method for modifying a target nucleic acid, comprising: The method comprises: The target nucleic acid is: (a) a V-type Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) fusion protein comprising a V-type CRISPR-Cas effector protein fused to an affinity polypeptide that binds to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to the peptide tag; wherein the target nucleic acid may be contacted with two or more deaminase fusion proteins; and (c) guide nucleic acid contacting the wherein the Type V CRISPR-Cas fusion protein, the deaminase fusion protein and the guide nucleic acid may be co-expressed, thereby modifying the target nucleic acid. method.
4. 1. A method for modifying a target nucleic acid, comprising: The method comprises: The target nucleic acid is: (a) V-shaped Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) effector proteins; (b) a recruitment guide nucleic acid comprising a guide RNA linked to an RNA recruitment motif; and (c) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif; wherein the target nucleic acid may be contacted with two or more deaminase fusion proteins; contacting the wherein the Type V CRISPR-Cas effector protein, the deaminase fusion protein and the recruitment guide nucleic acid may be co-expressed, thereby modifying the target nucleic acid. method.
5. The method of claim 2 or 3, wherein the peptide tag comprises two or more copies of the peptide tag (e.g., two or more epitopes).
6. 6. The method of any one of claims 2, 3 or 5, wherein the peptide tag is a GCN4 peptide repeat unit (e.g., Sun-Tag), a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or strep tag II, a V5 tag, and / or a VSV-G epitope.
7. The method of any one of claims 2, 3, 5, or 6, wherein the affinity polypeptide is an antibody.
8. The method of claim 7, wherein the antibody is an scFv antibody.
9. 5. The method of claim 4, wherein the recruitment guide nucleic acid is linked to two or more RNA recruitment motifs, wherein the two or more RNA recruitment motifs may be the same RNA recruitment motif or different RNA recruitment motifs.
10. 10. The method of claim 4 or 9, wherein the RNA recruitment motif is a telomerase Ku-binding motif (e.g., a Ku-binding hairpin) and the affinity polypeptide is a Ku polypeptide (e.g., a Ku heterodimer); the RNA recruitment motif is a telomerase Sm7-binding motif and the affinity polypeptide is an Sm7 polypeptide; the RNA recruitment motif is an MS2 phage operator stem-loop and the affinity polypeptide is an MS2 Coat protein (MCP); the RNA recruitment motif is a PP7 phage operator stem-loop and the affinity polypeptide is a PP7 Coat protein (PCP); the RNA recruitment motif is an SfMu phage Com stem-loop and the affinity polypeptide is a Com RNA-binding protein; the RNA recruitment motif is a Pumilio / fem-3 mRNA-binding factor (PUF) and the affinity polypeptide is a PUF-binding site (PBS) polypeptide; and / or the RNA recruitment motif is a synthetic RNA aptamer and the affinity polypeptide is a corresponding aptamer ligand.
11. 11. The method of any one of claims 1 to 10, wherein the Type V CRISPR-Cas effector protein comprises a mutation within a nuclease active site.
12. The method according to any one of claims 1 to 11, wherein the deaminase is cytosine deaminase and / or adenine deaminase.
13. 13. The method of any one of claims 1 to 12, wherein the two or more deaminase fusion proteins comprise the same or different deaminases.
14. 14. The method of Claim 13, wherein a first deaminase fusion protein of the two or more deaminase fusion proteins comprises a cytosine deaminase and a second deaminase fusion protein of the two or more deaminase fusion proteins comprises an adenine deaminase.
15. 15. The method of any one of claims 12 to 14, wherein the cytosine deaminase is apolipoprotein B mRNA editing catalytic polypeptide-like (APOBEC), human activation-induced deaminase (hAID), FERNY deaminase, and / or CDA1 deaminase, and the cytosine deaminase may comprise or be any one of the sequences of SEQ ID NOs: 23 to 32.
16. 16. The method of claim 15, wherein the cytosine deaminase is APOBEC3A, and the APOBEC3A may comprise or be the sequence of (e.g., SEQ ID NO: 24).
17. The adenine deaminase is TadA (tRNA-specific adenosine deaminase) and / or TadA * (an evolved tRNA-specific adenosine deaminase), wherein the adenine deaminase may comprise or be any one of the sequences of SEQ ID NOs: 33-43.
18. The method of any one of claims 1 to 17, further comprising the step of introducing a glycosylase inhibitor, and optionally introducing two or more glycosylase inhibitors.
19. 19. The method of claim 18, wherein the glycosylase inhibitor is a uracil DNA glycosylase inhibitor (UGI), and the UGI may comprise or be the sequence of SEQ ID NO:
44.
20. 20. The method of claim 18 or 19, wherein the glycosylase inhibitor is recruited to the Type V CRISPR-Cas effector protein and / or the deaminase, and may be via a protein-protein interaction, an RNA-protein interaction, or a chemical interaction.
21. 21. The method of any one of claims 18 to 20, wherein the glycosylase inhibitor is fused to the Type V CRISPR-Cas effector protein and / or the deaminase.
22. 22. The method of any one of claims 1 to 21, wherein the Type V CRISPR-Cas effector protein, the Type V CRISPR-Cas fusion protein, the deaminase, and / or the deaminase fusion protein are encoded by a polynucleotide.
23. 23. The method of any one of claims 18 to 22, wherein the glycosylase inhibitor is encoded by a polynucleotide.
24. 24. The method of claim 22 or 23, wherein the polynucleotide encoding the Type V CRISPR-Cas effector protein, the polynucleotide encoding the deaminase, and the guide nucleic acid are comprised within one or more expression cassette(s) and / or vector(s).
25. 25. The method of any one of claims 22 to 24, wherein the polynucleotide encoding the Type V CRISPR-Cas fusion protein, the polynucleotide encoding the deaminase fusion protein, and the guide nucleic acid are comprised within one or more expression cassette(s) and / or vector(s).
26. 26. The method of any one of claims 22-25, wherein the polynucleotide encoding the Type V CRISPR-Cas effector protein, the polynucleotide encoding the deaminase fusion protein, and the recruitment guide nucleic acid are comprised within one or more expression cassette(s) and / or vector(s).
27. 27. The method of any one of claims 23 to 26, wherein the polynucleotide encoding the glycosylase inhibitor is contained within a single expression cassette or vector, and may be within the same expression cassette or vector containing one or more of the polynucleotide encoding the type V CRISPR-Cas effector protein, the polynucleotide encoding the deaminase, and the guide nucleic acid; may be within the same expression cassette or vector containing one or more of the polynucleotide encoding the type V CRISPR-Cas fusion protein, the polynucleotide encoding the deaminase fusion protein, and / or the guide nucleic acid; and / or may be within the same expression cassette or vector containing one or more of the polynucleotide encoding the type V CRISPR-Cas effector protein, the polynucleotide encoding the deaminase fusion protein, and / or the recruitment guide nucleic acid.
28. 28. The method of any one of claims 1 to 27, wherein the Type V CRISPR-Cas effector protein is linked to a polypeptide of interest.
29. 29. The method of claim 28, wherein the polypeptide of interest comprises at least one polypeptide having nickase activity, recombinase activity, transposase activity, methylase activity, glycosylase (DNA glycosylase) activity, glycosylase inhibitor activity (e.g., uracil DNA glycosylase inhibitor (UGI)), demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, restriction endonuclease activity (e.g., Fok1), nucleic acid binding activity, methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, polymerase activity, ligase activity, helicase activity, and / or photolyase activity.
30. 30. The method of claim 28 or 29, wherein the polypeptide of interest is encoded by a polynucleotide.
31. 31. The method of any one of claims 1 to 30, wherein the type V CRISPR-Cas effector protein is a Cas12a (Cpf1) polypeptide, a Cas12b polypeptide, a Cas12c (C2c3) polypeptide, a Cas12d (CasY) polypeptide, a Cas12e (CasX) polypeptide, a Cas12g polypeptide, a Cas12h polypeptide, a Cas12i polypeptide, a C2c4 polypeptide, a C2c5 polypeptide, a C2c8 polypeptide, a C2c9 polypeptide, a C2c10 polypeptide, a Cas14a polypeptide, a Cas14b polypeptide, and / or a Cas14c polypeptide.
32. The type V CRISPR-Cas effector protein is a Cas12a (Cpf1) polypeptide, and the Cas12a (Cpf1) polypeptide may have the sequence of any one of SEQ ID NOs: 3 to 19, or may be encoded by a polynucleotide having the sequence of any one of SEQ ID NOs: 20 to 22. The Cas12a polypeptide may be a Lachnospiraceae bacterium ND2006 Cas12a (LbCas12a) (LbCpf1) polypeptide (e.g., having the sequence of any one of SEQ ID NOs: 3 or 9 to 11), an Acidaminococcus sp. (e.g., having the sequence of SEQ ID NO: 4), or a polynucleotide encoding the Cas12a (Cpf1) polypeptide.
32. The method of any one of claims 1 to 31, wherein the polypeptide is a Cpf1 (AsCas12a) (AsCpf1) polypeptide and / or an enAsCas12a polypeptide (e.g., encoded by a polynucleotide having the sequence of any one of SEQ ID NOs: 20 to 22).
33. 33. The method of any one of claims 1 to 32, wherein the target nucleic acid is in an organism which may be an animal, a plant, a fungus, an archaea, or a bacterium.
34. 29. The method of claim 28, wherein the polynucleotide encoding the Type V CRISPR-Cas effector protein, the polynucleotide encoding the Type V CRISPR-Cas fusion protein, the polynucleotide encoding the deaminase, the polynucleotide encoding the deaminase fusion protein, the polynucleotide encoding the glycosylase inhibitor, and / or the polynucleotide encoding the polypeptide of interest is codon-optimized for expression in the organism comprising the target nucleic acid.
35. (a) a V-type Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) fusion protein comprising a V-type CRISPR-Cas effector protein fused to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the peptide tag; and (c) guide nucleic acid Including, Nucleic acid constructs.
36. (a) V-shaped Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) effector proteins; (b) a recruitment guide nucleic acid comprising a guide RNA linked to an RNA recruitment motif; and (c) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif. Including, Nucleic acid constructs.
37. A V-shaped Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) system, below: (a) a type V CRISPR-Cas fusion protein comprising a type V CRISPR-Cas effector protein fused to a peptide tag; (b) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the peptide tag; and (c) a guide nucleic acid comprising a spacer sequence and a repeat sequence; Including, the guide nucleic acid is capable of forming a complex with the type-V CRISPR-Cas effector protein of the type-V CRISPR-Cas fusion protein, and the spacer sequence is capable of hybridizing to a target nucleic acid, thereby guiding the type-V CRISPR-Cas fusion protein to the target nucleic acid, wherein the deaminase fusion protein is recruited to the type-V CRISPR-Cas fusion protein and target nucleic acid by binding of the affinity polypeptide to the peptide tag fused to the type-V CRISPR-Cas fusion protein, thereby enabling the system to modify (e.g., cleave or edit) the target nucleic acid. system.
38. A V-shaped Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) (CRISPR-Cas) system, below: (a) Type V CRISPR-Cas effector protein; (b) a recruitment guide nucleic acid comprising a guide RNA linked to an RNA recruitment motif; and (c) a deaminase fusion protein comprising a deaminase fused to an affinity polypeptide that binds to the RNA recruitment motif; Including, the recruitment guide nucleic acid comprises a spacer sequence and a repeat sequence, the recruitment guide nucleic acid is capable of forming a complex with the V-type CRISPR-Cas effector protein, the guide RNA is capable of hybridizing to a target nucleic acid, thereby guiding the V-type CRISPR-Cas effector protein to the target nucleic acid, and the deaminase fusion protein is recruited to the V-type CRISPR-Cas effector protein and target nucleic acid by binding of the affinity polypeptide to the RNA recruitment motif fused to the recruitment guide nucleic acid, thereby enabling the system to modify (e.g., cleave or edit) the target nucleic acid. system.
39. 39. The nucleic acid construct of claim 35 or 36, or the CRISPR-Cas system of claim 37 or 38, wherein the type V CRISPR-Cas effector protein is a Cas12a (Cpf1) polypeptide, a Cas12b polypeptide, a Cas12c (C2c3) polypeptide, a Cas12d (CasY) polypeptide, a Cas12e (CasX) polypeptide, a Cas12g polypeptide, a Cas12h polypeptide, a Cas12i polypeptide, a C2c4 polypeptide, a C2c5 polypeptide, a C2c8 polypeptide, a C2c9 polypeptide, a C2c10 polypeptide, a Cas14a polypeptide, a Cas14b polypeptide, and / or a Cas14c polypeptide.
40. The type V CRISPR-Cas effector protein is a Cas12a (Cpf1) polypeptide, and the Cas12a (Cpf1) polypeptide may have the sequence of any one of SEQ ID NOs: 3 to 19, or may be encoded by a polynucleotide having the sequence of any one of SEQ ID NOs: 20 to 22, wherein the Cas12a polypeptide is selected from the group consisting of Lachnospiraceae bacterium ND2006 Cas12a (LbCas12a) (LbCpf1) polypeptide (e.g., having the sequence of any one of SEQ ID NOs: 3 or 9 to 11), Acidaminococcus sp. (e.g., having the sequence of SEQ ID NO: 4), and the like. The nucleic acid construct of claim 35 or 36, or the CRISPR-Cas system of claim 37 or 38, which may be a Cpf1 (AsCas12a) (AsCpf1) polypeptide and / or an enAsCas12a polypeptide (e.g., encoded by a polynucleotide having the sequence of any one of SEQ ID NOs: 20-22).
41. 41. An expression cassette comprising the nucleic acid construct of any one of claims 35, 36, 39, or 40.
42. An expression cassette comprising the CRISPR-Cas system according to any one of claims 37 to 40.
43. 42. A cell comprising the nucleic acid construct of any one of claims 35, 36, 39, or 40 or the expression cassette of claim 41.
44. A cell comprising the CRISPR-Cas system of any one of claims 37 to 40 or the expression cassette of claim 42.