Compositions and methods for allelic RNA-encoded DNA replacement

The combination of Type V CRISPR-Cas effector proteins and reverse transcriptase with extended guide nucleic acids enhances nucleic acid editing capabilities, addressing limitations of current tools by enabling broader and more efficient nucleotide incorporation.

JP7785002B2Active Publication Date: 2025-12-12PAIRWISE PLANTS SERVICES INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022538683
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-05
Filing Date
2020-11-05
Publication Date
2025-12-12
Estimated Expiration
2040-11-05

AI Technical Summary

Technical Problem

Current base editing tools are limited in their ability to edit nucleic acids beyond converting cytosine and adenine to thymine and guanine, have a small editing window, and require high PAM density, limiting their applicability to trait-related targets.

Method used

A method involving Type V CRISPR-Cas effector proteins, reverse transcriptase, and extended guide nucleic acids is used to modify target nucleic acids, allowing for broader editing capabilities by introducing additional nucleotides through a combination of CRISPR-Cas effector proteins and reverse transcriptase.

Benefits of technology

This approach expands the scope of nucleic acid editing, enabling efficient modification of a wider range of organisms by incorporating desired nucleotides into the genome, overcoming limitations of existing base editing tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007785002000016
    Figure 0007785002000016
  • Figure 0007785002000017
    Figure 0007785002000017
  • Figure 0007785002000018
    Figure 0007785002000018
Patent Text Reader

Abstract

The present invention relates to recombinant nuclear constructs comprising Type V CRISPR-Cas effector proteins, reverse transcriptase, and extended guide nucleic acids, and methods of use thereof for modifying nucleic acids in plants.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Description of Electronic Filing of Sequence Listings In lieu of a paper copy, a Sequence Listing in ASCII text format, filed under 37 CFR § 1.821, with filename 1499-12WO_ST25.txt, 768,694 bytes in size, created on November 5, 2020, and submitted via EFS-Web, is provided. This Sequence Listing is hereby incorporated by reference herein for its disclosure.

[0002] Priority statement This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Application No. 62 / 930,836, filed November 5, 2019, the entire contents of which are incorporated herein by reference.

[0003] The present invention relates to recombinant nuclear constructs comprising Type V CRISPR-Cas effector proteins, reverse transcriptase, and extended guide nucleic acids, and methods of use thereof for modifying nucleic acids in plants. [Background technology]

[0004] Base editing has been shown to be an efficient method for changing cytosine and adenine residues into thymine and guanine, respectively.Although these tools are powerful, they have certain limitations, such as bystander bases, a small base editing window that limits the accessibility of trait-related targets unless enzymes with high PAM density are available for compensation, the limited ability to convert cytosine and adenine into residues other than thymine and guanine, respectively, and the inability to edit thymine or guanine residues.Therefore, the current tools available for base editing are limited.Therefore, new editing tools are needed to make nucleic acid editing more useful by increasing the scope of possible editing for a larger number of organisms. Summary of the Invention [Means for solving the problem]

[0005] In a first aspect, a method of modifying a target nucleic acid is provided, comprising contacting the target nucleic acid with (a) a Type V CRISPR-Cas effector protein or a Type II CRISPR-Cas effector protein, (b) a reverse transcriptase, and (c) an extended guide nucleic acid (e.g., extended Type II or Type V CRISPR RNA, extended Type II or Type V CRISPR DNA, extended Type II or Type V crRNA, extended Type II or Type V crDNA), thereby modifying the target nucleic acid.

[0006] In a second aspect, there is provided a method of modifying a target nucleic acid, comprising contacting the target nucleic acid at a first site with (a)(i) a first CRISPR-Cas effector protein and (ii) a first extended guide nucleic acid (e.g., extended CRISPR RNA, extended CRISPR DNA, extended crRNA, extended crDNA), and (b)(i) a second CRISPR-Cas effector protein, (ii) a first reverse transcriptase, and (ii) the first guide nucleic acid, thereby modifying the target nucleic acid.

[0007] In a third aspect, there is provided a method for modifying a target nucleic acid in a plant or plant cell, comprising the steps of introducing an expression cassette of the invention into the plant or plant cell, thereby modifying the target nucleic acid in the plant or plant cell, and producing the plant or plant cell comprising the modified target nucleic acid.

[0008] In a fourth aspect, there is provided a complex comprising: (a) a type V CRISPR-Cas effector protein or a type II CRISPR-Cas effector protein; (b) a reverse transcriptase; and (c) an extended guide nucleic acid (e.g., extended CRISPR RNA, extended CRISPR DNA, extended crRNA, extended crDNA, e.g., a target allele guide (tag) nucleic acid (i.e., tag DNA, tag RNA)).

[0009] In a fifth aspect, there is provided a codon-optimized expression cassette for expression in an organism, comprising, from 5' to 3': (a) a polynucleotide encoding a plant-specific promoter sequence (e.g., ZmUbi1, MtUb2, RNA polymerase II (Pol II)); (b) a plant-codon-optimized polynucleotide encoding a type V CRISPR-Cas nuclease (e.g., Cpf1 (Cas12a), dCas12a, etc.); (c) a linker sequence; and (d) a plant-codon-optimized polynucleotide encoding a reverse transcriptase.

[0010] In a sixth aspect, there is provided an expression cassette that is codon-optimized for expression in an organism, comprising: (a) a polynucleotide encoding a promoter sequence; and (b) an extended RNA guide sequence that comprises at its 3' end a primer binding site and an extension that comprises an edit desired to be incorporated into a target nucleic acid (e.g., a reverse transcriptase template), optionally contained in the expression cassette and optionally operably linked to a Pol II promoter.

[0011] The present invention further provides cells, including plant cells, bacterial cells, archaeal cells, fungal cells, and animal cells, that contain target nucleic acids modified by the methods of the invention, as well as organisms, including plants, bacteria, archaeal cells, fungi, and animals, that contain the cells. Additionally, the present invention provides kits that include the polynucleotides, polypeptides, and expression cassettes of the invention.

[0012] These and other aspects of the invention are described in more detail below in the description of the invention. [Brief explanation of the drawings]

[0013] [Figure 1]Figure 1 provides a schematic diagram showing the generation of a DNA sequence from reverse transcription from crRNA and subsequent incorporation into the nick site. The extended guide crRNA (tag RNA) is bound to Cpf1 nickase (cas12a nickase) (nCpf1, top left). Alternatively, the extension encoding the editing template can be located 5' to the crRNA. The 3' end of the crRNA is complementary to the DNA at the nick site (non-bold paired lines, top left). nCpf1 can be covalently bound to reverse transcriptase (RT), or RT can be recruited to nCpf1, in which case multiple reverse transcriptase proteins can be recruited to nCpf1. RT polymerizes the DNA from the 3' end of the DNA nick on the second strand, creating a DNA sequence complementary to the crRNA (bold paired lines, curly brackets, top right), followed by complementary nucleotides (non-bold paired lines, top right), with nucleotides that are not complementary to the genome. Upon dissociation, the resulting DNA has an extended ssDNA with a 3' overhang, which is largely the same sequence as the original DNA (non-bold pair of lines, bottom right), but with some non-native nucleotides (bold pair of lines, curly braces, bottom right). This flap is in equilibrium with a structure with a 5' overhang, in which a mismatched nucleotide has been incorporated into the DNA (bottom left). The equilibrium can also be driven toward the left-hand structure by reducing mismatch repair, removing the 5' flap during repair and replication, and nicking the first strand as described herein. [Figure 2]Figure 2 provides a schematic diagram illustrating a method for reducing mismatch repair. To drive the equilibrium more favorably toward forming an end product with a modified nucleotide (bold, curly brackets), a nickase is directed (via a guide nucleic acid) to cleave the first strand of the target nucleic acid (e.g., the target strand or the bottom strand) at a region outside the RT-editing region (lightning bolt), a fixed distance from the nick in the second strand (e.g., the target strand or the top strand). NCpf1:crRNA molecules can be on one or both sides of the editing bubble. Nicking the first strand (dashed line) signals to the cell that the newly incorporated nucleotide is the correct nucleotide during mismatch repair and replication, thus favoring the end product with the new nucleotide. Another possible method for driving the equilibrium toward the desired product is removal of the 5' flap. [Figure 3] FIG. 3 is a diagram illustrating an alternative method of modifying nucleic acids using the compositions of the invention, whereby two nicks are introduced into the second strand, allowing the sequence introduced by RT to replace the doubly nicked WT sequence and thereby be more efficiently incorporated into the genome. [Figure 4] LbCas12a_R1138A is an in vitro validated nickase separated on a 1% TAE-agarose gel. A supercoiled 2.8 kB plasmid ran at an apparent size of 2.0 kB (lane 2) until a double-strand break was generated by wild-type LbCas12a (lane 3). [Figure 5] FIG. 5 shows the configuration of the REDRAW editor tested in E. coli (see Example 1). [Figure 6] FIG. 6 shows the conformations of the tag RNAs tested in the first library. [Figure 7] FIG. 7 shows the structure of an example hairpin sequence designed for use in REDRAW editing. [Figure 8]Figure 8 shows Sanger sequencing results demonstrating a TGA>CTG edit in the non-functional aadA gene that restores antibiotic resistance. Editing was observed in colonies in selection 10 with the protein configuration SV40-MMLV-RT-XTEN-nLbCas12a-SV40 (SEQ ID NO: 71). [Figure 9] 9 shows Sanger sequencing results demonstrating an AAA>CGT edit in the rpsL gene in the E. coli genome that confers resistance to the antibiotic streptomycin. Editing was observed in colonies in selection 2.5 with the protein configuration SV40-MMLV-RT-XTEN-nRVRLbCas12a(H759A)-SV40 (SEQ ID NO: 79). [Figure 10] Figure 10 shows Sanger sequencing results demonstrating TGA>GAT editing in the non-functional aadA gene, restoring antibiotic resistance. Editing was observed in colonies in selection 2.25 with the protein configuration SV40-nLbCas12a-XTEN-MMLV-RT-SV40 (SEQ ID NO: 73). [Figure 11] Figure 11 shows Sanger sequencing results demonstrating a TGA>GAT edit in the non-functional aadA gene that restores antibiotic resistance. Editing was observed in colonies in selection 2.31 with the protein configuration SV40-MMLV-RT-XTEN-nLbCas12a(H759A)-SV40 (SEQ ID NO: 83). [Figure 12]Figure 12 shows an example of an editing method performed in human cells (see Example 2). Panel A shows a double-stranded target nucleic acid. The Cas12a complex (which includes an extended guide nucleic acid, not shown) is recruited to the first strand (target strand, bottom strand), and the 5' flap in the second strand (top strand, non-target strand) is optionally removed with a 5'-3' exonuclease (panel B). Panel C shows the reverse transcriptase MMuLV-RT (5M) extending from a primed site or primer (complementary to the primer binding site) on the target nucleic acid (dashed line = extension). Panels D and E show the division of the DNA intermediate and the creation of a newly edited DNA strand via mismatch repair and DNA ligation. [Figure 13] Figure 13 shows precise editing at the FANCF1 site in HEK293T cells using various guide conformations. The construct name is Cas12a(H759A)+RT(5M)+RecE FANCF1. [Figure 14] Figure 14 shows precise editing at the DMNT1 site in HEK293T cells using various guide conformations. The construct is named Cas12a(H759A)+RT(5M)+DMNT1. [Figure 15] FIG. 15 shows the effect of exonuclease transfection on precise editing activity at DMNT1 sites (normalized to no exonuclease treatment, pUC19=1). DETAILED DESCRIPTION OF THE INVENTION

[0014] A brief description of arrays SEQ ID NOs: 1 to 20 are examples of Cas12a amino acid sequences useful in the present invention.

[0015] SEQ ID NO:21 and SEQ ID NO:22 are exemplary regulatory sequences encoding a promoter and an intron.

[0016] SEQ ID NOs: 23-25 ​​provide examples of peptide tags and affinity polypeptides.

[0017] SEQ ID NOs: 26-36 provide examples of RNA recruitment motifs and corresponding affinity polypeptides.

[0018] SEQ ID NOs: 37-52 provide examples of single-stranded RNA-binding domains (RBDs).

[0019] SEQ ID NO: 53 and SEQ ID NO: 97 provide examples of reverse transcriptase sequences (M-MuLV).

[0020] SEQ ID NOs: 54-56 provide examples of the location of protospacer adjacent motifs in type V CRISPR-Cas12a nucleases.

[0021] SEQ ID NO: 57 and SEQ ID NO: 58 provide examples of constructs of the invention.

[0022] SEQ ID NO:59 and SEQ ID NO:60 provide examples of CRISPR RNAs and examples of protospacers.

[0023] SEQ ID NO: 61 and SEQ ID NO: 62 provide examples of introns.

[0024] SEQ ID NOs: 63-86 provide examples of REDRAW editing constructs.

[0025] SEQ ID NO: 87 provides an example of a tag RNA with an 11 base pair (bp) primer binding sequence and a 96 bp reverse transcriptase template.

[0026] SEQ ID NOs: 88-91 provide sequences of example plasmids.

[0027] SEQ ID NOs: 92 to 94 provide the sequences of the tag RNAs associated with editing shown in Figures 9 to 11, respectively.

[0028] SEQ ID NO: 96 provides an example of an LbCas12a with an H759A mutation and flanked by NLSs on both sides.

[0029] SEQ ID NOs: 98-101 provide examples of 5'-3' exonuclease polypeptides.

[0030] SEQ ID NO: 102 and SEQ ID NO: 103 provide examples of DMNT1 target sites and target spacers.

[0031] SEQ ID NO: 104 and SEQ ID NO: 105 provide examples of FANCF1 target sites and target spacers.

[0032] Detailed Description The present invention will be described hereinafter with reference to the accompanying drawings and examples, which illustrate embodiments of the invention. This description is not intended to be a detailed catalog of all the various ways in which the invention may be practiced or all the features that may be added to the invention. For example, features illustrated with respect to one embodiment may be incorporated into other embodiments, and features illustrated with respect to a particular embodiment may be omitted from that embodiment. Thus, the present invention contemplates that some embodiments of the invention may exclude or omit any feature or combination of features described herein. Furthermore, numerous modifications and additions to the various embodiments suggested herein will be apparent to those skilled in the art in view of this disclosure and do not depart from the invention. Therefore, the following description is intended to illustrate some particular embodiments of the invention, but does not exhaustively specify all permutations, combinations, and variations thereof.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. The terminology used in the description of the invention herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0034] All publications, patent applications, patents, and other references cited herein are incorporated by reference in their entirety for the teachings relevant to the sentence and / or paragraph in which the reference is presented.

[0035] Unless the context dictates otherwise, it is specifically contemplated that the various features of the invention described herein can be used in any combination. Furthermore, in some embodiments of the invention, the invention also contemplates that any feature or combination of features described herein can be excluded or omitted. By way of example, if the specification states that a composition includes components A, B, and C, it is specifically contemplated that any one or combination of A, B, or C can be omitted and rejected, either alone or in any combination.

[0036] As used in the description of this invention and the appended claims, the singular forms "a," "an," and "the" are intended to include the singular forms as well, unless the content clearly dictates otherwise.

[0037] Also, as used herein, "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted as the alternative ("or").

[0038] As used herein, the term "about," when referring to a measurable value such as an amount or concentration, is meant to encompass variations of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the specified value, and the specified value. For example, "about X," where X is a measurable value, means including variations of X and ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of X. Ranges provided herein for measurable values ​​can include any other ranges and / or individual values ​​therein.

[0039] Phrases used herein such as "X to Y" and "about X to Y" should be interpreted as including X and Y. Phrases used herein such as "about X to Y" mean "about X to about Y", and phrases such as "about X to Y" mean "about X to about Y".

[0040] The recitation of ranges of values ​​herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise stated herein, and each separate value is incorporated into the specification as if it were individually listed in the present specification. For example, if the range 10 to 15 is disclosed, then 11, 12, 13, and 14 are also disclosed.

[0041] As used herein, the terms "comprise", "comprises" and "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0042] As used herein, the transitional phrase "consisting essentially of" means that the scope of a claim should be construed to include the specified materials or steps recited in the claim and that do not materially affect the basic and novel feature(s) of the claimed invention. Thus, when used in the claims of the present invention, the term "consisting essentially of" is not intended to be construed as equivalent to "comprising."

[0043] As used herein, the terms "increase," "increasing," "enhance," "enhancing," "improve," and "improving" (and grammatical variations thereof) describe an increase of at least about 25%, 50%, 75%, 100%, 150%, 200%, 300%, 400%, 500%, or more compared to a control.

[0044] As used herein, the terms "reduce," "reduced," "reducing," "reduction," "diminish," and "decrease" (and grammatical variations thereof) describe a decrease of, for example, at least about 5%, 10%, 15%, 20%, 25%, 35%, 50%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% compared to a control. In certain embodiments, a decrease may result in no or essentially no detectable activity or amount (i.e., an insignificant amount, e.g., less than about 10%, or even less than 5%).

[0045] A "heterologous" or "recombinant" nucleotide sequence is a nucleotide sequence that is not naturally associated with the host cell into which it is introduced, and includes non-naturally occurring multiple copies of a naturally occurring nucleotide sequence.

[0046] A "native" or "wild-type" nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence refers to a naturally occurring or endogenous nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence. Thus, for example, a "wild-type mRNA" is an mRNA that naturally occurs in or is endogenous to a reference organism. A "homologous" nucleic acid sequence is a nucleotide sequence that is naturally associated with a host cell into which it is introduced.

[0047] As used herein, the terms "nucleic acid," "nucleic acid molecule," "nucleotide sequence," and "polynucleotide" refer to RNA or DNA, whether linear or branched, single-stranded or double-stranded, or a hybrid thereof. This term also encompasses RNA / DNA hybrids. When synthetically producing dsRNA, less common bases such as inosine, 5-methylcytosine, 6-methyladenine, and hypoxanthine can also be used for antisense, dsRNA, and ribozyme pairing. For example, polynucleotides containing C-5 propyne analogs of uridine and cytidine have been shown to bind to RNA with high affinity and to be potent antisense inhibitors of gene expression. Other modifications, such as modifications to the phosphodiester backbone or the 2'-hydroxyl in the ribose sugar group of RNA, can also be made.

[0048] As used herein, the term "nucleotide sequence" refers to a heteropolymer of nucleotides or the sequence of these nucleotides from the 5' to 3' end of a nucleic acid molecule, including DNA or RNA molecules, including cDNA, DNA fragments or portions, genomic DNA, synthetic (e.g., chemically synthesized) DNA, plasmid DNA, mRNA, and antisense RNA, any of which can be single-stranded or double-stranded. The terms "nucleotide sequence," "nucleic acid," "nucleic acid molecule," "nucleic acid construct," "oligonucleotide," and "polynucleotide" are also used interchangeably herein to refer to a heteropolymer of nucleotides. The nucleic acid molecules and / or nucleotide sequences provided herein are presented herein in the 5' to 3' direction, from left to right, and are represented using the standard code for representing nucleotide letters, as set forth in the U.S. Sequencing Rules, 37 CFR 1.821-1.825, and World Intellectual Property Organization (WIPO) Standard ST.25. As used herein, the term "5' region" can refer to the region of a polynucleotide closest to the 5' end of the polynucleotide. Thus, for example, an element in the 5' region of a polynucleotide can be located anywhere from the first nucleotide at the 5' end of the polynucleotide to a nucleotide located in the middle of the polynucleotide. As used herein, the term "3' region" can refer to the region of a polynucleotide that is closest to the 3' end of the polynucleotide. Thus, for example, an element in the 3' region of a polynucleotide can be located anywhere from the first nucleotide at the 3' end of the polynucleotide to a nucleotide located in the middle of the polynucleotide.

[0049] As used herein, the term "gene" refers to a nucleic acid molecule that can be used to produce mRNA, antisense RNA, miRNA, anti-microRNA antisense oligodeoxyribonucleotide (AMO), etc. A gene may or may not be used to produce a functional protein or gene product. A gene can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and / or 5' and 3' untranslated regions). A gene may be "isolated," which means a nucleic acid that is substantially or essentially free from components normally found associated with the nucleic acid in its natural state. Such components include other cellular material, culture medium from recombinant production, and / or various chemicals used in the chemical synthesis of nucleic acids.

[0050] The term "mutation" refers to point mutations (e.g., single base pair insertions or deletions resulting in missense or nonsense, or frameshifts), insertions, deletions, and / or truncations. When a mutation is a substitution of a residue in an amino acid sequence for another, or a deletion or insertion of one or more residues in the sequence, the mutation is typically described by identifying the original residue, then the position of the residue in the sequence, and then the identity of the newly substituted residue.

[0051] As used herein, the term "complementary" or "complementarity" refers to the natural binding of polynucleotides by base pairing under permissive salt and temperature conditions. For example, the sequence "AGT" (5' to 3') binds to the complementary sequence "TCA" (3' to 5'). Complementarity between two single-stranded molecules can be "partial," where only a portion of the nucleotides bind, or complete, where complete complementarity exists between the single-stranded molecules. The degree of complementarity between nucleic acid strands has a significant impact on the efficiency and strength of hybridization between nucleic acid strands.

[0052] As used herein, "complementary" can mean 100% complementarity with a reference nucleotide sequence, or it can mean less than 100% complementarity (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, etc. complementarity).

[0053] A "portion" or "fragment" of a nucleotide sequence of the invention is a fragment that is reduced in length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides removed) relative to a reference nucleic acid or nucleotide sequence and is identical or nearly identical (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110, 111, 112, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides removed) to a reference nucleic acid or nucleotide sequence. "A nucleotide sequence that is identical to a nucleic acid fragment of a nucleic acid fragment of a nucleotide sequence ... As an example, the repeat sequence of a guide nucleic acid of the invention may comprise a portion of a wild-type V-type CRISPR-Cas repeat sequence (e.g., a wild-type CRISPR-Cas repeat, e.g., a repeat from a CRISPR Cas system such as Cas12a (Cpf1), Cas12b (C2c1), Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c). In some embodiments, the repeat sequence of a guide nucleic acid of the invention may comprise a portion of a wild-type CRISPR-Cas9 repeat sequence.

[0054] Various nucleic acids or proteins that share homology are referred to herein as "homologues." The term homologue includes homologous sequences from the same species and other species, as well as orthologous sequences from the same species and other species. "Homology" refers to the level of similarity between two or more nucleic acid and / or amino acid sequences in terms of percent positional identity (i.e., sequence similarity or identity). Homology also refers to the concept of similar functional properties between various nucleic acids or proteins. Thus, the compositions and methods of the present invention further include homologues to the nucleotide and polypeptide sequences of the present invention. As used herein, "orthologous" refers to homologous nucleotide and / or amino acid sequences in different species that arose during speciation from a common ancestral gene. A homologue of a nucleotide sequence of the present invention has substantial sequence identity (e.g., at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%) to said nucleotide sequence of the present invention.

[0055] As used herein, "sequence identity" refers to the degree to which two optimally aligned polynucleotide or polypeptide sequences are invariant over the alignment window of components, e.g., nucleotides or amino acids. "Identity" can be readily calculated by known methods, including, but not limited to, those described in Computational Molecular Biology (Lesk, A.M., ed.), Oxford University Press, New York (1988), Biocomputing: Informatics and Genome Projects (Smith, D.W., ed.), Academic Press, New York (1993), Computer Analysis of Sequence Data, Part I (Griffin, A.M. and Griffin, H.G., eds.), Humana Press, New Jersey (1994), Sequence Analysis in Molecular Biology (von Heinje, G., ed.), Academic Press (1987), and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.), Stockton Press, New York (1991).

[0056] As used herein, the term "percent sequence identity" or "percent identity" refers to the percentage of identical nucleotides in a linear polynucleotide sequence of a reference ("query") polynucleotide molecule (or its complement) compared to a test ("subject") polynucleotide molecule (or its complement) when the two sequences are optimally aligned. In some embodiments, "percent identity" can refer to the percentage of identical amino acids in an amino acid sequence compared to a reference polypeptide.

[0057] The phrases "substantially identical" or "substantial identity," as used herein in the context of two nucleic acid molecules, nucleotide sequences, or protein sequences, refer to two or more sequences or subsequences that have at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or more nucleotide or amino acid residue identity when compared and aligned for maximum correspondence as measured using one of the following sequence comparison algorithms or by visual inspection: In some embodiments of the invention, substantial identity exists over a region of contiguous nucleotides of a nucleotide sequence of the invention that is about 10 nucleotides to about 20 nucleotides, about 10 nucleotides to about 25 nucleotides, about 10 nucleotides to about 30 nucleotides, about 15 nucleotides to about 25 nucleotides, about 30 nucleotides to about 40 nucleotides, about 50 nucleotides to about 60 nucleotides, about 70 nucleotides to about 80 nucleotides, about 90 nucleotides to about 100 nucleotides, or more nucleotides in length, or any range therein up to the full length of the sequence. In some embodiments, nucleotide sequences can be substantially identical over at least about 20 nucleotides (e.g., about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 nucleotides). In some embodiments, a substantially identical nucleotide or protein sequence performs substantially the same function as the nucleotide (or encoded protein sequence) to which it is substantially identical.

[0058] In sequence comparison, typically, one sequence serves as a reference sequence to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are input into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. The sequence comparison algorithm then calculates the percent sequence identity of the test sequence(s) relative to the reference sequence based on the designated program parameters.

[0059] Optimal alignment of sequences for aligning a comparison window is well known to those skilled in the art and can be performed by tools such as Smith and Waterman's local homology algorithm, Needleman and Wunsch's homology alignment algorithm, Pearson and Lipman's similarity search method, and optionally by computer implementations of these algorithms, such as GAP, BESTFIT, FASTA, and TFASTA, available as part of the GCG® Wisconsin Package® (Accelrys Inc., San Diego, CA). The "fractional identity" of the aligned segments of a test sequence and a reference sequence is the number of identical components shared by the two aligned sequences divided by the total number of components in the reference sequence segment, for example, the entire reference sequence or a smaller, defined portion of the reference sequence. Percent sequence identity is expressed as the fractional identity multiplied by 100. Comparison of one or more polynucleotide sequences can be to a full-length polynucleotide sequence or a portion thereof, or to a longer polynucleotide sequence. For purposes of the present invention, "percent identity" may also be determined using BLASTX version 2.0 for translated nucleotide sequences and BLASTN version 2.0 for polynucleotide sequences.

[0060] Two nucleotide sequences can also be considered to be substantially complementary if the two sequences hybridize to each other under stringent conditions. In some representative embodiments, two nucleotide sequences considered to be substantially complementary hybridize to each other under highly stringent conditions.

[0061] "Stringent hybridization conditions" and "stringent hybridization wash conditions" in the context of nucleic acid hybridization experiments such as Southern and Northern hybridizations are sequence-dependent and vary under various environmental parameters. An extensive guide to nucleic acid hybridization can be found in Tijssen Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, Part I, Chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assays," Elsevier, New York (1993). In general, highly stringent hybridization and wash conditions are those that, for a particular sequence, achieve a thermal melting point (T m ) is chosen to be approximately 5°C lower than

[0062] T m is the temperature (under defined ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched probe. Very stringent conditions are defined as the temperature T mAn example of stringent hybridization conditions for hybridization of complementary nucleotide sequences with more than 100 complementary residues on a filter in a Southern or Northern blot is 50% formamide with 1 mg heparin at 42°C, with hybridization carried out overnight. An example of highly stringent wash conditions is 0.15 M NaCl at 72°C for approximately 15 minutes. An example of stringent wash conditions is a 0.2x SSC wash at 65°C for 15 minutes (see Sambrook, infra, for a description of SSC buffer). Often, a low stringency wash precedes a high stringency wash to remove background probe signal. For example, an example of a moderate stringency wash for a duplex of more than 100 nucleotides is 1x SSC at 45°C for 15 minutes. An example of low stringency washing for duplexes of, for example, more than 100 nucleotides is 4-6×SSC at 40°C for 15 minutes. For short probes (e.g., about 10-50 nucleotides), stringent conditions typically include a salt concentration of less than about 1.0 M Na ion, typically about 0.01-1.0 M Na ion (or other salt), pH 7.0-8.3, and a temperature typically of at least about 30°C. Stringent conditions can also be achieved with the addition of destabilizing agents such as formamide. Generally, a signal-to-noise ratio of 2-fold (or higher) than that observed for an unrelated probe in a particular hybridization assay indicates detection of specific hybridization. Nucleotide sequences that do not hybridize to each other under stringent conditions are still substantially identical if the proteins they encode are substantially identical. This may occur, for example, when copies of nucleotide sequences are generated using the maximum codon degeneracy permitted by the genetic code.

[0063] Polynucleotides and / or recombinant nucleic acid constructs of the invention can be codon optimized for expression. In some embodiments, polynucleotides, nucleic acid constructs, expression cassettes, and / or vectors of the invention (e.g., those comprising / encoding a CRISPR-Cas effector protein (e.g., a type V CRISPR-Cas effector protein), reverse transcriptase, flap endonuclease, 5'-3' exonuclease, etc.) are codon optimized for expression in an organism (e.g., a particular species), optionally an animal, plant, fungus, archaea, or bacterium. In some embodiments, a codon-optimized nucleic acid construct, polynucleotide, expression cassette, and / or vector of the invention has about 70% to about 99.9% (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%) identity or more to a nucleic acid construct, polynucleotide, expression cassette, and / or vector that is not codon-optimized.

[0064] In any of the embodiments described herein, the polynucleotides or nucleic acid constructs of the invention may be operably associated with various promoters and / or other regulatory elements for expression in plants and / or plant cells. Thus, in some embodiments, the polynucleotides or nucleic acid constructs of the invention may further comprise one or more promoters, introns, enhancers, and / or terminators operably linked to one or more nucleotide sequences. In some embodiments, the promoter may be operably associated with an intron (e.g., the Ubi1 promoter and intron). In some embodiments, a promoter associated with an intron may be referred to as a "promoter region" (e.g., the Ubi1 promoter and intron).

[0065] As used herein with reference to a polynucleotide, "operably linked" or "operably associated" means that the indicated elements are functionally related and, generally, physically related to each other. Thus, as used herein, the terms "operably linked" or "operably associated" refer to nucleotide sequences on a single nucleic acid molecule that are functionally related. Thus, a first nucleotide sequence operably linked to a second nucleotide sequence refers to a situation in which the first nucleotide sequence is placed in a functional relationship with the second nucleotide sequence. For example, a promoter is operably associated with a nucleotide sequence if it effects the transcription or expression of the nucleotide sequence. Those skilled in the art will understand that a control sequence (e.g., a promoter) need not necessarily be contiguous with the nucleotide sequence to which it is operably associated, so long as it functions to direct the expression of the sequence. Thus, for example, an intervening untranslated but transcribed nucleic acid sequence can be present between the promoter and the nucleotide sequence, and the promoter can still be considered "operably linked" to the nucleotide sequence.

[0066] As used herein, the term "linked" with reference to a polypeptide refers to the attachment of one polypeptide to another. A polypeptide may be linked (at the N-terminus or C-terminus) to another polypeptide directly (e.g., via a peptide bond) or through a linker.

[0067] The term "linker" is recognized in the glycotechnology field and refers to a chemical group or molecule that links two molecules or moieties, such as two domains of a fusion protein, such as a DNA-binding polypeptide or domain and a peptide tag and / or a reverse transcriptase and an affinity polypeptide that binds to a peptide tag, or a DNA endonuclease polypeptide or domain and a peptide tag and / or an affinity polypeptide that binds to a reverse transcriptase and a peptide tag. A linker can consist of a single linking molecule or can consist of multiple linking molecules. In some embodiments, a linker can be a chemical moiety such as an organic molecule, a group, a polymer, or a bivalent organic moiety. In some embodiments, a linker can be an amino acid or a peptide. In some embodiments, a linker is a peptide.

[0068] In some embodiments, peptide linkers useful in the invention are from about 2 to about 100 or more amino acids in length, e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids in length (e.g., about 2 to about 40, about 2 to about 50, about 2 to about 60, about 4 to about 40, about 4 to about 50, about 4 to about 60, about 5 to about 40, about 5 to about 50, about 5 to about 60, about 9 to about 40, about 9 to about 9 about 50, about 9 to about 60, about 10 to about 40, about 10 to about 50, about 10 to about 60, or about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62 , 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids in length (e.g., about 105, 110, 115, 120, 130, 140, 150 or more amino acids in length). In some embodiments, the peptide linker can be a GS linker.

[0069] As used herein, the terms "linked" or "fused" with reference to polynucleotides refer to the attachment of one polynucleotide to another. In some embodiments, two or more polynucleotide molecules can be linked by a linker, which can be a chemical moiety such as an organic molecule, a group, a polymer, or a divalent organic moiety. A polynucleotide can be linked or fused (at the 5' or 3' end) to another polynucleotide via a covalent or non-covalent linkage or bond, including, for example, Watson-Crick base pairing, or through one or more linking nucleotides. In some embodiments, a polynucleotide motif of a particular structure can be inserted within another polynucleotide sequence (e.g., an extension of a hairpin structure in a guide RNA). In some embodiments, the linking nucleotide can be a naturally occurring nucleotide. In some embodiments, the linking nucleotide can be a non-naturally occurring nucleotide.

[0070] A "promoter" is a nucleotide sequence that controls or regulates the transcription of a nucleotide sequence (e.g., a coding sequence) operably associated with the promoter. The coding sequence controlled or regulated by a promoter may encode a polypeptide and / or functional RNA. Typically, a "promoter" refers to a nucleotide sequence that contains a binding site for RNA polymerase II and directs the initiation of transcription. In general, promoters are found 5' or upstream to the start of the coding region of the corresponding coding sequence. A promoter may contain other elements that act as regulators of gene expression, such as a promoter region. These include a TATA box consensus sequence and often a CAAT box consensus sequence (Breathnach and Chambon, (1981), Annu. Rev. Biochem., 50:349). In plants, the CAAT box may be replaced by an AGGA box (Messing et al., (1983), Genetic Engineering of Plants, T. Kosuge, C. Meredith, and A. Hollaender (eds.), Plenum Press, pp. 211-227). In some embodiments, the promoter region may contain at least one intron (see, e.g., SEQ ID NO: 21, SEQ ID NO: 22).

[0071] Promoters useful in the present invention can include, for example, constitutive, inducible, temporally-regulated, developmentally-regulated, chemically-regulated, tissue-preferred and / or tissue-specific promoters for use in preparing recombinant nucleic acid molecules, such as "synthetic nucleic acid constructs" or "protein-RNA complexes." These various types of promoters are known in the art.

[0072] The choice of promoter may vary depending on the time and space requirements for expression, and may also vary based on the host cell to be transformed.The promoters for many different organisms are well known in the art.Based on the extensive knowledge existing in the art, suitable promoters can be selected for specific target host organisms.Therefore, for example, much is known about the promoters upstream of the genes that are highly constitutively expressed in model organisms, and such knowledge can be easily accessed and implemented in other systems as needed.

[0073] In some embodiments, promoters functional in plants can be used in the constructs of the present invention. Non-limiting examples of promoters useful for driving expression in plants include the promoter of RubisCo small subunit gene 1 (PrbcS1), the promoter of actin gene (Pactin), the promoter of nitrate reductase gene (Pnr), and the promoter of duplicated carbonic anhydrase gene 1 (Pdca1) (see Walker et al., Plant Cell Rep., 23:727-735 (2005); Li et al., Gene, 403:132-142 (2007); Li et al., Mol Biol. Rep., 37:1143-1154 (2010)). PrbcS1 and Pactin are constitutive promoters, while Pnr and Pdca1 are inducible promoters. Pnr is induced by nitrate and repressed by ammonium (Li et al., Gene, 403:132-142 (2007)), and Pdca1 is induced by salt (Li et al., Mol. Biol. Rep., 37:1143-1154 (2010)). In some embodiments, promoters useful in the present invention are RNA polymerase II (Pol II) promoters. In some embodiments, the U6 promoter or 7SL promoter from maize (Zea mays) may be useful in the constructs of the present invention. In some embodiments, the U6c promoter and / or 7SL promoter from maize may be useful for driving expression of a guide nucleic acid. In some embodiments, the U6c promoter, U6i promoter, and / or 7SL promoter from soybean (Glycine max) may be useful in the constructs of the present invention. In some embodiments, the U6c promoter, U6i promoter, and / or 7SL promoter from soybean may be useful for driving expression of a guide nucleic acid.

[0074] Examples of constitutive promoters useful in plants include, but are not limited to, the cestrum virus promoter (cmp) (U.S. Pat. No. 7,166,770), the rice actin 1 promoter (Wang et al., (1992), Mol. Cell. Biol., 12:3399-3406 and U.S. Pat. No. 5,641,876), the CaMV 35S promoter (Odell et al., (1985), Nature, 313:810-812), the CaMV 19S promoter (Lawton et al., (1987), Plant Mol. Biol., 9:315-324), the nos promoter (Ebert et al., (1987), Proc. Natl. Acad. Sci. USA, 84:5745-5749), the Adh promoter (Walker et al. (1987), Proc. Natl. Acad. Sci. USA, 84:6624-6629), the sucrose synthase promoter (Yang and Russell (1990), Proc. Natl. Acad. Sci. USA, 87:4144-4148), and the ubiquitin promoter. Constitutive promoters derived from ubiquitin accumulate in many cell types. Ubiquitin promoters have been cloned from several plant species for use in transgenic plants, such as sunflower (Binet et al., 1991, Plant Science, 79:87-94), maize (Christensen et al., 1989, Plant Molec. Biol., 12:619-632), and Arabidopsis (Norris et al., 1993, Plant Molec. Biol., 21:895-906). The maize ubiquitin promoter (UbiP) has been developed in transgenic monocotyledonous plant systems, and its sequence and vectors constructed for the transformation of monocotyledonous plants are disclosed in Patent Publication EP 0 342 926. The ubiquitin promoter is suitable for expressing the nucleotide sequences of the present invention in transgenic plants, particularly monocotyledonous plants.Additionally, the promoter expression cassettes described by McElroy et al. (Mol. Gen. Genet. 231:150-160 (1991)) can be readily modified for expression of the nucleotide sequences of the present invention and are particularly suitable for use in monocotyledonous hosts.

[0075] In some embodiments, tissue-specific / tissue-preferred promoters can be used to express heterologous polynucleotides in plant cells. Tissue-specific or preferred expression patterns include, but are not limited to, green tissue-specific or preferred, root-specific or preferred, stem-specific or preferred, flower-specific or preferred, or pollen-specific or preferred. Promoters suitable for expression in green tissues include many that regulate genes involved in photosynthesis, many of which have been cloned from both monocotyledonous and dicotyledonous plants. In one embodiment, a promoter useful in the present invention is the maize PEPC promoter from the phosphoenol carboxylase gene (Hudspeth and Grula, Plant Molec. Biol., 12:579-589 (1989)). Non-limiting examples of tissue-specific promoters include those involved in genes encoding storage proteins (e.g., β-conglycinin, cruciferin, napin, phaseolin), zein or oil body proteins (e.g., oleosin), or proteins involved in fatty acid biosynthesis (including acyl carrier protein, stearoyl-ACP desaturase, and fatty acid desaturase (fad2-1)), as well as other nucleic acids expressed during embryogenesis (e.g., Bce4; see, e.g., Kridl et al. (1991), Seed Sci. Res., 1:209-219 and EP Patent No. 255378). Tissue-specific or tissue-preferential promoters useful for expression of the nucleotide sequences of the invention in plants, particularly maize, include, but are not limited to, those directing expression in roots, pith, leaves, or pollen. Such promoters are disclosed, for example, in WO 93 / 07278, which is incorporated herein by reference in its entirety.Other non-limiting examples of tissue-specific or tissue-preferred promoters useful in the present invention include the cotton rubisco promoter disclosed in U.S. Pat. No. 6,040,504, the comecrose synthase promoter disclosed in U.S. Pat. No. 5,604,121, the root-specific promoter described by Framond (FEBS, 290:103-106 (1991); Ciba-Geigy, EP 0 452 269), the stem-specific promoter driving expression of the maize trpA gene described in U.S. Pat. No. 5,625,136 (Ciba-Geigy), the cestrum yellow leaf curling virus promoter disclosed in WO 01 / 73087, and, but not limited to, the ProOsLPS10 and ProOsLPS11 promoters from rice (Nguyen et al., Plant Physiol. 2004, 103:103-106 (2004)). Biotechnol. Reports, 9(5):297-306(2015)), ZmSTK2_USP from maize (Wang et al., Genome, 60(6):485-495(2017)), LAT52 and LAT59 from tomato (Twell et al., Development, 109(3):705-713(1990)), Zm13 (U.S. Patent No. 10,421,972), the PLA2-δ promoter from Arabidopsis thaliana (U.S. Patent No. 7,141,424), and / or the ZmC5 promoter from maize (International PCT Publication No. WO 1999 / 042587).

[0076] Further examples of plant tissue-specific / tissue-preferred promoters include, but are not limited to, the root hair-specific cis-element (RHE) (Kim et al., The Plant Cell, 18:2958-2970 (2006)), the root-specific promoters RCc3 (Jeong et al., Plant Physiol., 153:185-197 (2010)) and RB7 (U.S. Patent No. 5,459,252), the lectin promoter (Lindstrom et al., (1990), Der. Genet., 11:160-167 and Vodkin (1983), Prog. Clin. Biol. Res., 138:87-98), the maize alcohol dehydrogenase 1 promoter (Dennis et al., (1984), Nucleic Acids Res., 12:3983-4000), S-adenosyl-L-methionine synthase (SAMS) (Vander Mijnsbrugge et al., (1996), Plant and Cell Physiology, 37(8):1108-1115), maize light-harvesting complex promoter (Bansal et al., (1992), Proc. Natl. Acad. Sci. USA, 89:3654-3658), maize heat shock protein promoter (O'Dell et al., (1985), EMBO J., 5:451-458 and Rochester et al., (1986), EMBO J., 5:451-458), pea small subunit RuBP carboxylase promoter (Cashmore, "Nuclear genes encoding the small subunit of ribulose-l,5-bisphosphate carboxylase," pp. 29-39, Genetic Engineering of Plants (Hollaender, ed., Plenum Press, 1983 and Poulsen et al., (1986), Mol. Gen. Genet., 205:193-200), Ti plasmid mannopine synthase promoter (Langridge et al., (1989), Proc. Natl. Acad. Sci.USA 86:3219-3223), Ti plasmid nopaline synthase promoter (Langridge et al., (1989), supra), petunia chalcone isomerase promoter (van Tunen et al., (1988), EMBO J., 7:1257-1263), bean glycine-rich protein 1 promoter (Keller et al., (1989), Genes Dev., 3:1639-1646), truncated CaMV 35S promoter (O'Dell et al., (1985), Nature, 313:810-812), potato patatin promoter (Wenzler et al., (1989), Plant Mol. Biol., 13:347-354), root cell promoter (Yamamoto et al., (1990), Nucleic Acids Res., 18:7449), maize zein promoter (Kriz et al., (1987), Mol. Gen. Genet., 207:90-98; Langridge et al., (1983), Cell, 34:1015-1022; Reina et al., (1990), Nucleic Acids Res., 18:6425; Reina et al., (1990), Nucleic Acids Res., 18:7449; and Wandelt et al., (1989), Nucleic Acids Res., 17:2354), globulin-1 promoter (Belanger et al., (1991), Genetics, 129:863-872), α-tubulin cab promoter (Sullivan et al., (1989), Mol. Gen. Genet., 215:431-440), PEPCase promoter (Hudspeth & Grula, (1989), Plant Mol. Biol., 12:579-589), R gene complex-associated promoter (Chandler et al., (1989), Plant Cell, 1:1175-1183), and chalcone synthase promoter (Franken et al., (1991), EMBO J., 10:2605-2612).

[0077] Useful for seed-specific expression is the pea vicilin promoter (Czako et al., (1992), Mol. Gen. Genet., 235:33-40 and the seed-specific promoter disclosed in U.S. Pat. No. 5,625,136). Useful promoters for expression in mature leaves are those that are switched on at the onset of senescence, such as the SAG promoter from Arabidopsis (Gan et al., (1995), Science, 270:1986-1988).

[0078] Additionally, promoters functional in chloroplasts can be used. Non-limiting examples of such promoters include the bacteriophage T3 gene 9 5'UTR and other promoters disclosed in U.S. Patent No. 7,579,516. Other promoters useful in the present invention include, but are not limited to, the S-E9 small subunit RuBP carboxylase promoter and the Kunitz trypsin inhibitor gene promoter (Kti3).

[0079] Additional regulatory elements useful in the present invention include, but are not limited to, introns, enhancers, termination sequences, and / or 5' and 3' untranslated regions.

[0080] Introns useful in the present invention can be introns identified in plants, isolated from them, and then inserted into an expression cassette for use in plant transformation. As will be understood by those skilled in the art, introns can contain sequences necessary for self-excision and are incorporated in-frame within a nucleic acid construct / expression cassette. Introns can be used as spacers to separate multiple protein-coding sequences within a single nucleic acid construct, or they can be used within a single protein-coding sequence, for example, to stabilize mRNA. When used within a protein-coding sequence, they are inserted "in-frame" with an excision site. Introns can also be associated with promoters to improve or modify expression. By way of example, promoter / intron combinations useful in the present invention include, but are not limited to, the maize Ubi1 promoter and an intron.

[0081] Non-limiting examples of introns useful in the present invention include introns from the ADHI gene (e.g., Adh1-S introns 1, 2, and 6), ubiquitin gene (Ubi1), RuBisCO small subunit (rbcS) gene, RuBisCO large subunit (rbcL) gene, actin gene (e.g., actin-1 intron), pyruvate dehydrogenase kinase gene (pdk), nitrate reductase gene (nr), duplicated carbonic anhydrase gene 1 (Tdca1), psbA gene, atpA gene, or any combination thereof. Exemplary intron sequences include, but are not limited to, SEQ ID NO: 61 and SEQ ID NO: 62.

[0082] In some embodiments, polynucleotides and / or nucleic acid constructs of the invention can be or be contained within an "expression cassette." As used herein, "expression cassette" refers to a recombinant nucleic acid molecule that includes, for example, a nucleic acid construct of the invention (e.g., a CRISPR-Cas effector protein, a reverse transcriptase polypeptide or domain, a flap endonuclease polypeptide or domain (e.g., FEN)), and / or a 5'-3' exonuclease), wherein the nucleic acid construct is operably associated with one or more regulatory sequences (e.g., a promoter, a terminator, etc.). Thus, some embodiments of the present invention provide expression cassettes designed to express, for example, a nucleic acid construct of the invention (e.g., a nucleic acid construct of the invention encoding a CRISPR-Cas effector protein or domain, a reverse transcriptase polypeptide or domain, a flap endonuclease polypeptide or domain, and / or a 5'-3' exonuclease polypeptide or domain. When an expression cassette of the invention comprises multiple polynucleotides, the polynucleotides may be operably linked to a single promoter that drives expression of all of the polynucleotides, or the polynucleotides may be operably linked to one or more separate promoters (e.g., three polynucleotides in any combination, one, two, three, four, five, six, eight ... (The expression cassette may be driven by one or three promoters. When two or more separate promoters are used, the promoters may be the same promoter or different promoters. Thus, the polynucleotide encoding a CRISPR-Cas effector protein or domain, the polynucleotide encoding a reverse transcriptase polypeptide or domain, the polynucleotide encoding a flap endonuclease polypeptide or domain, and / or the polynucleotide encoding a 5'-3' exonuclease polypeptide or domain contained in the expression cassette may each be operably linked to a separate promoter, or may be operably linked to two or more promoters in any combination.

[0083] An expression cassette comprising a nucleic acid construct of the invention may be chimeric, meaning that at least one of its components is heterologous to at least one of its other components (e.g., a promoter from a host organism operably linked to a polynucleotide of interest to be expressed in a host cell, where the polynucleotide of interest is from an organism different from the host or is not normally found in association with the promoter). Expression cassettes may be naturally occurring but obtained in a recombinant form useful for heterologous expression.

[0084] The expression cassette can also optionally contain a transcriptional and / or translational termination region (i.e., a termination region) and / or an enhancer region that is functional in the selected host cell. A variety of transcription terminators and enhancers are known in the art and are available for use in expression cassettes. The transcription terminator is responsible for the termination of transcription and correct mRNA polyadenylation. The termination and / or enhancer regions may be native to the transcription initiation region, e.g., native to the gene encoding the CRISPR-Cas effector protein, the gene encoding the reverse transcriptase, the gene encoding the flap endonuclease, and / or the gene encoding the 5'-3' exonuclease, may be native to the host cell, or may be native to another source (e.g., foreign or heterologous to the promoter, the gene encoding the CRISPR-Cas effector protein, the gene encoding the reverse transcriptase, the gene encoding the flap endonuclease, and / or the gene encoding the 5'-3' exonuclease, the host cell, or any combination thereof).

[0085] The expression cassettes of the present invention can also include a polynucleotide encoding a selectable marker that can be used to select transformed host cells. As used herein, a "selectable marker" refers to a polynucleotide sequence that, when expressed, confers a distinct phenotype on host cells expressing the marker, thus allowing such transformed cells to be distinguished from those that do not possess the marker. Such polynucleotide sequences may encode either a selectable or a screenable marker, depending on whether the marker confers a trait that can be selected for by chemical means, such as by using a selection agent (e.g., an antibiotic), or whether the marker is simply a substance that can be identified through observation or testing, such as by screening (e.g., fluorescence). Numerous examples of suitable selectable markers are known in the art and can be used in the expression cassettes described herein.

[0086] In addition to expression cassettes, the nucleic acid molecules / constructs and polynucleotide sequences described herein can be used in connection with vectors. The term "vector" refers to a composition for transferring, delivering, or introducing a nucleic acid (or multiple nucleic acids) into a cell. A vector includes a nucleic acid construct containing the nucleotide sequence(s) to be transferred, delivered, or introduced. Vectors for use in transforming host organisms are well known in the art. Non-limiting examples of general classes of vectors include viral vectors, plasmid vectors, phage vectors, phagemid vectors, cosmid vectors, fosmid vectors, bacteriophages, artificial chromosomes, minicircles, or Agrobacterium binary vectors, which may or may not be double-stranded or single-stranded, linear or circular, and self-infecting or mobilizable. In some embodiments, viral vectors may include, but are not limited to, retroviral, lentiviral, adenoviral, adeno-associated, or herpes simplex viral vectors. As defined herein, vectors are capable of transforming prokaryotic or eukaryotic hosts either by integration into a cellular genome or by being present extrachromosomally (e.g., an autonomously replicating plasmid with an origin of replication). Also included are shuttle vectors, which refer to DNA vehicles capable of replicating, naturally or by design, in two different host organisms, which may be selected from actinomycetes and related species, bacteria, and eukaryotes (e.g., higher plant, mammalian, yeast, or fungal cells). In some embodiments, the nucleic acid in the vector is under the control of and operably linked to an appropriate promoter or other regulatory elements for transcription in the host cell. The vector may be a bifunctional expression vector that functions in multiple hosts. In the case of genomic DNA, it may contain its own promoter and / or other regulatory elements, and in the case of cDNA, it may be under the control of an appropriate promoter and / or other regulatory elements for expression in the host cell.Thus, a nucleic acid construct or polynucleotide of the invention and / or an expression cassette comprising same may be included in a vector, as described herein and known in the art.

[0087] As used herein, "contact," "contacting," "contacted," and grammatical variations thereof refer to placing components of a desired reaction together under conditions suitable for carrying out the desired reaction (e.g., transformation, transcriptional regulation, genome editing, nicking, and / or cleavage). As an example, a target nucleic acid may be contacted with a type II or type V CRISPR-Cas effector protein and reverse transcriptase, or a nucleic acid construct encoding the same, under conditions in which the CRISPR-Cas effector protein and reverse transcriptase are expressed and the CRISPR-Cas effector protein binds to the target nucleic acid, and the reverse transcriptase is fused to the CRISPR-Cas effector protein or recruited to the CRISPR-Cas effector protein (e.g., via a peptide tag fused to the CRISPR-Cas effector protein and an affinity tag fused to the reverse transcriptase), thereby placing the reverse transcriptase in proximity to the target nucleic acid and thereby modifying the target nucleic acid. Other methods for recruiting reverse transcriptase may be used, utilizing other protein-protein interactions, as well as RNA-protein interactions and chemical interactions.

[0088] "Modifying" or "modification," as used herein with reference to a target nucleic acid, includes editing (e.g., mutating), covalently modifying, exchanging / substituting, deleting, truncating, nicking, and / or transcriptionally regulating the target nucleic acid. In some embodiments, modifications can include indels of any size and / or single base changes (SNPs) of any type.

[0089] "Introducing," "introduce," "introduced" (and grammatical variations thereof) in the context of a polynucleotide of interest means presenting a nucleotide sequence of interest (e.g., a polynucleotide, a nucleic acid construct, and / or a guide nucleic acid) to a host organism or a cell of said organism (e.g., a host cell, e.g., a plant cell) in a manner such that the nucleotide sequence is accessible to the interior of the cell.

[0090] As used herein, the terms "transformation" or "transfection" may be used interchangeably and refer to the introduction of heterologous nucleic acid into a cell. Cellular transformation can be stable or transient. Thus, in some embodiments, a host cell or host organism may be stably transformed with a polynucleotide / nucleic acid molecule of the present invention. In some embodiments, a host cell or host organism may be transiently transformed with a nucleic acid construct of the present invention.

[0091] "Transient transformation" in the context of polynucleotides means introducing a polynucleotide into a cell without integrating it into the genome of the cell.

[0092] By "stably introducing" or "stably introduced" in the context of a polynucleotide introduced into a cell is intended that the introduced polynucleotide is stably integrated into the genome of the cell, and thus the cell is stably transformed with the polynucleotide.

[0093] As used herein, "stable transformation" or "stably transformed" refers to the introduction of a nucleic acid molecule into a cell and its integration into the cell's genome. The integrated nucleic acid molecule can therefore be inherited by its progeny, more particularly by progeny in multiple successive generations. As used herein, "genome" includes nuclear, mitochondrial, and plastid genomes, and thus includes, for example, the integration of a nucleic acid into the genome of a chloroplast or mitochondrion. As used herein, stable transformation can also refer to a transgene maintained extrachromosomally, for example, on a minichromosome or plasmid.

[0094] Transient transformation can be detected, for example, by enzyme-linked immunosorbent assay (ELISA) or Western blot, which can be detected by the presence of peptides or polypeptides encoded by one or more transgenes introduced into the organism. Stable transformation of cells can be detected, for example, by Southern blot hybridization assay of the genomic DNA of the cells using a nucleic acid sequence that specifically hybridizes with the nucleotide sequence of the transgene introduced into the organism (e.g., a plant). Stable transformation of cells can be detected, for example, by Northern blot hybridization assay of the RNA of the cells using a nucleic acid sequence that specifically hybridizes with the nucleotide sequence of the transgene introduced into the host organism. Stable transformation of cells can also be detected, for example, by polymerase chain reaction (PCR) or other amplification reactions known in the art using specific primer sequences that hybridize with the target sequence(s) of the transgene, resulting in amplification of the transgene sequence, which can be detected according to standard methods. Transformation can also be detected by direct sequencing and / or hybridization protocols well known in the art.

[0095] Thus, in some embodiments, the nucleotide sequences, polynucleotides, nucleic acid constructs, and / or expression cassettes of the invention may be transiently expressed and / or stably integrated into the genome of a host organism. Thus, in some embodiments, the nucleic acid constructs of the invention (e.g., one or more expression cassettes encoding a DNA-binding polypeptide or domain, an endonuclease polypeptide or domain, a reverse transcriptase polypeptide or domain, a flap endonuclease polypeptide or domain, and / or a nucleic acid-modifying polypeptide or domain) may be transiently introduced into a cell using a guide nucleic acid, such that the DNA is not maintained in the cell.

[0096] The nucleic acid constructs of the present invention can be introduced into cells by any method known to those of skill in the art. In some embodiments of the present invention, transformation of a cell comprises nuclear transformation. In other embodiments, transformation of a cell comprises plastid transformation (e.g., chloroplast transformation). In further embodiments, the recombinant nucleic acid constructs of the present invention can be introduced into cells via conventional breeding techniques.

[0097] Procedures for transforming both eukaryotic and prokaryotic organisms are well known and routine in the art and are described throughout the literature (see, e.g., Jiang et al., 2013, Nat. Biotechnol., 31:233-239; Ran et al., Nature Protocols, 8:2281-2308 (2013)).

[0098] Thus, nucleotide sequences can be introduced into a host organism or its cells by a number of methods well known in the art. The methods of the present invention do not rely on a particular method for introducing one or more nucleotide sequences into an organism, relying only on their access to the interior of at least one cell of the organism. When multiple nucleotide sequences are introduced, they can be assembled as part of a single nucleic acid construct or as separate nucleic acid constructs, and can be located on the same or different nucleic acid constructs. Thus, the nucleotide sequences can be introduced into the target cell in a single transformation event and / or separate transformation events, or, if relevant, the nucleotide sequences can be incorporated into a plant, for example, as part of a breeding protocol.

[0099] Base editing has been shown to be an efficient method for changing cytosine and adenine residues to thymine and guanine, respectively. While powerful, these tools have limitations, such as bystander bases, a small base editing window, and limited PAM.

[0100] There are several critical steps for accurate template editing in cells, each of which has its own rate limiting effect, which, combined with low efficiency, severely hinders the ability to effectively edit. For example, one step requires inducing cells to initiate repair events at the target site. This is typically achieved by creating a double-strand break (DSB) or nick using an exogenously provided sequence-specific nuclease or nickase. Another step requires the local availability of a homologous template for repair. This step requires that the template be proximal to the DSB at exactly the right time to commit the DSB to the template editing pathway. Specifically, this step is widely considered to be the rate-limiting step of current editing technologies. Another step is the efficient incorporation of sequences from the template into the interrupted or nicked target. Prior to the present invention, this step was typically provided by endogenous DNA repair enzymes in cells. The efficiency of this step is low and difficult to manipulate. The present invention circumvents many of the major obstacles to the efficiency of the process of template editing by co-localizing in a coordinate manner the functional groups required to perform the steps described above.

[0101] Figure 1 illustrates the generation of a DNA sequence from reverse transcription from crRNA and subsequent integration into a nick site using the methods and constructs of the present invention. The extended crRNA is shown in blue and is bound to the second strand nickase Cpf1 (Cas12a) (nCpf1, top left). As described in more detail herein, nCpf1 can be covalently linked to a reverse transcriptase (RT), e.g., via a peptide, or the RT can be recruited to nCpf1 (e.g., via the use of affinity polypeptides that bind to peptide tag motifs / tags or via chemical interactions as described herein), in which case multiple reverse transcriptase proteins (RTs) can be integrated. n) can be recruited. The 3' end of the sgRNA is complementary to the DNA at the nick site (unbold paired line, top left). RT then polymerizes the DNA from the 3' end of the DNA nick, resulting in a DNA sequence complementary to the RNA with a non-genomic nucleotide (bold paired line, bracket, top right), followed by the complementary nucleotide (unbold paired line, top right). Upon dissociation, the resulting DNA has an extended ssDNA with a 3' overhang that is largely the same sequence as the original DNA (unbold paired line, bottom right), but with some non-native nucleotides (bold paired line, bracket, bottom right). This flap is in equilibrium with a structure with a 5' overhang, in which a mismatched nucleotide has been incorporated into the DNA (bottom left). This equilibrium favors perfect pairing on the right side, but drive can be reduced in various ways, including, for example, nicking the second strand (e.g., the target strand or the bottom strand). The left-hand structure may be preferentially cleaved by cellular flap endonucleases involved in DNA lagging strand synthesis, which are highly conserved between mammalian and plant cells (the amino acid sequence of Homo sapiens FEN1 is over 50% identical to both maize and soybean FEN1). In some embodiments, a flap endonuclease may be introduced to drive the equilibrium toward the 3' flap containing the non-native / mismatched nucleotide. Longer 5' flaps are often removed by Dna2 protein in eukaryotic cells, again driving the equilibrium toward the 3' flap (desired) product (see, e.g., Nucleic Acids Res., 2012 Aug; 40(14):6774-86).

[0102] Further in the process of the present invention, and as illustrated in Figure 2, Cpf1 nickase can be targeted to a region outside the RT-editing region (lightning bolt) as described herein to reduce mismatch repair and drive the equilibrium more favorably toward forming end products with modified nucleotides (bold, brackets). NCpf1:crRNA molecules can be on one or both sides of the editing bubble. Nicking the first strand (e.g., the target strand or bottom strand in Figure 2) (dashed line) indicates to the cell that the newly incorporated nucleotide is the correct nucleotide during mismatch repair and replication, thus favoring end products with the new nucleotide.

[0103] Reverse transcriptase (RT) variants can have significant effects on the temperature sensitivity and processivity of editing systems. Natural and rationally and non-rationally engineered (i.e., directed evolution) variants of RT can be useful to optimize activity and processivity profiles at plant-preferred temperatures.

[0104] Protein domain fusions with RT polypeptides can have a significant impact on the temperature sensitivity and processivity of editing systems. RT enzymes can be improved for temperature sensitivity, processivity, and template affinity through fusion with ssRNA-binding domains (RBDs). These RBDs can have sequence specificity, nonspecificity, or sequence preference (see, e.g., SEQ ID NOS: 37-52). A range of affinity distributions can be beneficial for editing in various cellular and in vitro environments. RBDs can modify both specificity and binding free energy by increasing or decreasing RBD size to recognize more or fewer nucleotides. Multiple RBDs result in proteins with affinity distributions that are a combination of the individual RBDs. Adding one or more RBDs to an RT enzyme can result in increased affinity, increased or decreased sequence specificity, and / or enhanced cooperativity.

[0105] After reverse transcriptase incorporates the edit into the genome, sequence redundancy exists between the newly synthesized edited sequence and the original WT sequence it is intended to replace. This results in either a 5' or 3' flap at the target site, which must be repaired by the cell. The two states exist in equilibrium, with binding energy favoring the 3' flap because more base pairs are available when the WT sequence is paired with its complement than when the edited strand is paired with its complement. This is unfavorable for efficient editing because processing (removal) of the 3' flap can remove the edited residue and return the target to the WT sequence. However, cellular flap endonucleases such as FEN1 or Dna2 can efficiently process the 5' flap. Therefore, instead of relying on the function of a 5'-flap endonuclease native to the cell, some embodiments of the present invention may increase the concentration of the flap endonuclease at the target to further favor the desired equilibrium outcome (removal of the WT sequence in the 5' flap so that the edited sequence can be stably incorporated at the target site). This can be achieved by overexpressing the 5' flap endonuclease as a free protein in the cell. Alternatively, FEN or Dna2 can be actively recruited to the target site by association with the CRISPR complex via direct protein fusion or by non-covalent recruitment, such as using a peptide tag and affinity polypeptide pair (e.g., SunTag antibody / epitope pair) or chemical interactions as described herein.

[0106] The present invention further provides methods for modifying a target nucleic acid using the proteins / polypeptides and / or fusion proteins of the present invention, polynucleotides and nucleic acid constructs encoding same, and / or expression cassettes and / or vectors comprising same. The methods may be performed in vivo (e.g., in a cell or organism) or in vitro (e.g., cell-free) systems. Thus, in some embodiments, a method for modifying a target nucleic acid in a plant cell is provided, comprising contacting the target nucleic acid with (a) a type V CRISPR-Cas effector protein or a type II CRISPR-Cas effector protein, (b) a reverse transcriptase, and (c) an extended guide nucleic acid (e.g., extended type II or type V CRISPR RNA, extended type II or type V CRISPR DNA, extended type II or type V crRNA, extended type II or type V crDNA, e.g., tag RNA, tag DNA), thereby modifying the target nucleic acid. In some embodiments, the Type V or Type II CRISPR-Cas effector protein, the reverse transcriptase, and the extended guide nucleic acid may form or be included in a complex that can interact with the target nucleic acid. In some embodiments, the methods of the present invention may further include contacting the target nucleic acid with (a) a second Type V or Type II CRISPR-Cas effector protein, (b) a second reverse transcriptase, and (c) a second extended guide nucleic acid (e.g., extended CRISPR RNA, extended CRISPR DNA, extended crRNA, extended crDNA, e.g., tag DNA, tag RNA) that targets (the spacer is substantially complementary to / binds with) a site on the first strand of the target nucleic acid, thereby modifying the target nucleic acid.In some embodiments, the methods of the invention may further comprise contacting the target nucleic acid with (a) a second Type V CRISPR-Cas effector protein or a second Type II CRISPR-Cas effector protein, (b) a second reverse transcriptase, and (c) a second extended guide nucleic acid (e.g., extended CRISPR RNA, extended CRISPR DNA, extended crRNA, extended crDNA, e.g., tag DNA, tag RNA) that targets (to which the spacer is substantially complementary / binds) a site on the second strand of the target nucleic acid, thereby modifying the target nucleic acid. In some embodiments, the methods of the invention comprise contacting the target nucleic acid at a temperature of about 20°C to 42°C (e.g., about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, or 42°C, and any value or range therein). In some embodiments, the target nucleic acid may be contacted with an additional polypeptide and / or a nucleic acid construct encoding the same to improve mismatch repair.In some embodiments, the methods of the invention include contacting a target nucleic acid with (a) a CRISPR-Cas effector protein and (b) a guide nucleic acid, wherein (i) the CRISPR-Cas effector protein is a nickase (e.g., nCas9, nCas12a), and the guide nucleic acid is about 10 to about 125 bases 5' or 3' to the site on the second strand nicked by the Type II or Type V CRISPR-Cas effector protein. pairs (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 12 7, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, or 125 base pairs, or any range or value therein) of the first of the target nucleic acids or (ii) the CRISPR-Cas effector protein is a nickase (e.g., nCas9, nCasl2a) and nicks a site on the second strand of the target nucleic acid that is about 10 to about 125 base pairs (either 5' or 3') from the site on the first strand nicked by the Type II or Type V CRISPR-Cas effector protein, thereby improving mismatch repair.

[0107] In some embodiments, the extended guide nucleic acid comprises (i) a type V CRISPR nucleic acid or a type II CRISPR nucleic acid (type II or type V CRISPR RNA, type II or type V CRISPR DNA, type II or type V crRNA, type II or type V crDNA) and / or a CRISPR nucleic acid and a tracr nucleic acid (e.g., type II or type V tracrRNA, type II or type V tracrDNA), and (ii) an extension portion comprising a primer binding site and a reverse transcriptase template (RT template). In some embodiments, the extension portion can be fused to either the 5' or 3' end of the CRISPR nucleic acid (e.g., 5' to 3' repeat-spacer-extension portion or extension portion-repeat-spacer) and / or the 5' or 3' end of the tracr nucleic acid. In some embodiments, the extension portion of the extended guide nucleic acid comprises a 5' to 3' RT template (RTT) and a primer binding site (PBS), or a 5' to 3' PBS and RTT, depending on the location of the extension portion relative to the guide CRISPR RNA. In some embodiments, the target nucleic acid is double-stranded and comprises a first strand and a second strand, and the primer binding site binds to the second strand of the target nucleic acid (the non-target, top strand). In some embodiments, the target nucleic acid is double-stranded and comprises a first strand and a second strand, and the primer binding site binds to the first strand of the target nucleic acid (e.g., the target strand, which is the same strand as the strand to which the CRISPR-Cas effector protein is recruited, the bottom strand). In some embodiments, the target nucleic acid is double-stranded and comprises a first strand and a second strand, and the primer binding site binds to the second strand of the target nucleic acid (the non-target strand, which is the opposite strand to the strand to which the CRISPR-Cas effector protein is recruited). Thus, in some embodiments, the editing reverse transcriptase (RT) adds to the target strand (the strand to which the spacer of the CRISPR RNA is complementary and to which the CRISPR-Cas effector protein is recruited), and in some embodiments, the editing reverse transcriptase (RT) adds to the non-target strand (the strand to which the spacer of the CRISPR RNA is complementary and to which the CRISPR-Cas effector protein is recruited).

[0108] In some embodiments, the method includes contacting a target nucleic acid with (a) a type V CRISPR-Cas effector protein or a type II CRISPR-Cas effector protein, (b) a reverse transcriptase, and (c) an extended guide nucleic acid (e.g., extended type II or type V CRISPR RNA, extended type II or type V CRISPR DNA, extended type II or type V crRNA, extended type II or type V crDNA), wherein the extended guide nucleic acid is (i) a type II or type V CRISPR nucleic acid (e.g., type II or type V CRISPR RNA, type II or type V CRISPR

[0003] Methods for modifying a target nucleic acid having a first strand and a second strand are provided, the method comprising: (i) a CRISPR nucleic acid and a tracr nucleic acid (e.g., type II or type V crRNA, type II or type V crDNA) and / or a CRISPR nucleic acid and a tracr nucleic acid (e.g., type II or type V tracrRNA, type II or type V tracrDNA); and (ii) an extension portion comprising a primer binding site and a reverse transcriptase template (RT template), wherein the type II or type V CRISPR nucleic acid comprises a spacer that binds to the first strand (e.g., target strand) (i.e., is complementary to a portion of consecutive nucleotides in the first strand of the target nucleic acid), and the primer binding site binds to the first strand (target strand), thereby modifying the target nucleic acid. In some embodiments, the type II CRISPR-Cas effector protein can be a Cas9 polypeptide. In some embodiments, the type V CRISPR-Cas effector protein can be a Cas12a polypeptide. In some embodiments, the Type II or Type V CRISPR-Cas effector protein, the reverse transcriptase, and the extended guide nucleic acid can form a complex or be included in a complex. In some embodiments, the contacting step can further include contacting the target nucleic acid with a 5'-3' exonuclease.

[0109] In some embodiments, the target nucleic acid may further be contacted with a 5' flap endonuclease (FEN), optionally a FEN1 and / or Dna2 polypeptide, thereby improving mismatch repair by removing the 5' flap that does not contain an edit to be incorporated into the target nucleic acid. In some embodiments, FEN and / or Dna2 may be overexpressed in the presence of the target nucleic acid. In some embodiments, FEN may be a fusion protein comprising a FEN domain fused to a V-type CRISPR-Cas effector protein or domain, thereby recruiting FEN to the target nucleic acid. In some embodiments, Dna2 may be a fusion protein comprising a Dna2 domain fused to a V-type CRISPR-Cas effector protein or domain, thereby recruiting Dna2 to the target nucleic acid.

[0110] In some embodiments, the type II or type V CRISPR-Cas effector protein may be a type II or type V CRISPR-Cas fusion protein comprising a type V CRISPR-Cas effector protein domain fused (linked) to a peptide tag (e.g., an epitope or a multimerized epitope), and the FEN may be a FEN fusion protein comprising a FEN domain fused to an affinity polypeptide that binds to the peptide tag, thereby recruiting FEN to the type II or type V CRISPR-Cas effector protein domain and the target sequence. In some embodiments, the Type II or Type V CRISPR-Cas effector protein may be a Type II or Type V CRISPR-Cas fusion protein comprising a Type II or Type V CRISPR-Cas effector protein domain fused (linked) to a peptide tag (e.g., an epitope or a multimerized epitope), and Dna2 may be a Dna2 fusion protein comprising a Type II or Type V CRISPR-Cas effector protein domain fused to (linked with) an affinity polypeptide that binds to the peptide tag, thereby recruiting Dna2 to the Type II or Type V CRISPR-Cas effector protein domain and a target sequence. In some embodiments, the Type V CRISPR-Cas effector protein may be a Type II or Type V CRISPR-Cas fusion protein comprising a Type II or Type V CRISPR-Cas effector protein domain fused (linked) to a peptide tag (e.g., an epitope or a multimerized epitope), and the FEN may be a FEN fusion protein comprising a FEN domain fused to an affinity polypeptide that binds to the peptide tag, thereby recruiting FEN to the Type II or Type V CRISPR-Cas effector protein domain and the target sequence.In some embodiments, the Type II or Type V CRISPR-Cas effector protein may be a Type II or Type V CRISPR-Cas fusion protein comprising a Type II or Type V CRISPR-Cas effector protein domain fused (linked) to a peptide tag (e.g., an epitope or a multimerized epitope), and Dna2 may be a Dna2 fusion protein comprising a Type II or Type V CRISPR-Cas effector protein domain fused (linked) to a peptide tag (e.g., an epitope or a multimerized epitope), and Dna2 is a Dna2 fusion protein comprising a Dna2 domain fused to an affinity polypeptide that binds to the peptide tag, thereby recruiting Dna2 to the Type II or Type V CRISPR-Cas effector protein domain and the target sequence. In some embodiments, the target nucleic acid may be contacted with two or more FEN fusion proteins and / or Dna2 fusion proteins.

[0111] In some embodiments, the methods of the present invention may further include contacting the target nucleic acid with a 5'-3' exonuclease, thereby improving mismatch repair by removing the edit-free 5' flap (unedited strand) incorporated into the target nucleic acid. In some embodiments, the 5'-3' exonuclease may be fused to a type II or type V CRISPR-Cas effector protein, optionally a type II or type V CRISPR-Cas fusion protein. In some embodiments, the 5'-3' exonuclease may be a fusion protein comprising a 5'-3' exonuclease fused to a peptide tag, and the type II or type V CRISPR-Cas effector protein may be a fusion protein comprising a type II or type V CRISPR-Cas effector protein domain fused to an affinity polypeptide capable of binding to the peptide tag, thereby improving mismatch repair. In some embodiments, the 5'-3' exonuclease can be a fusion protein comprising a 5'-3' exonuclease fused to an affinity polypeptide capable of binding to a peptide tag, and the type II or type V CRISPR-Cas effector protein can be a fusion protein comprising a type II or type V CRISPR-Cas effector protein domain fused to a peptide tag. In some embodiments, the 5'-3' exonuclease can be a fusion protein comprising a 5'-3' exonuclease fused to an affinity polypeptide capable of binding to an RNA recruitment motif, and the extended guide nucleic acid is linked to the RNA recruitment motif, thereby recruiting the 5'-3' exonuclease to the target nucleic acid through interaction between the affinity polypeptide and the RNA recruitment motif. The 5'-3' exonuclease can be any known or later discovered 5'-3' exonuclease functional in the organism, cell, or in vitro system of interest. In some embodiments, the 5'-3' exonuclease may include, but is not limited to, RecE exonuclease, RecJ exonuclease, T5 exonuclease, and / or T7 exonuclease.In some embodiments, a C-terminal fragment of RecE exonuclease flanked on both sides by a nuclear localization sequence (NLS) from, for example, Escherichia coli (K12 strain) may be used (SEQ ID NO: 98). In some embodiments, a RecJ exonuclease flanked on both sides by a nuclear localization sequence (NLS) from, for example, E. coli (K12 strain) may be used (SEQ ID NO: 99). In some embodiments, a T5 exonuclease flanked on both sides by a nuclear localization sequence (NLS) may be used (SEQ ID NO: 100). In some embodiments, a T7 exonuclease flanked on both sides by a nuclear localization sequence (NLS) from, for example, Escherichia phage 7 may be used (SEQ ID NO: 101).

[0112] In some embodiments, the method of the present invention can further comprise the step of reducing double-strand breaks.In some embodiments, the step of reducing double-strand breaks can be carried out by introducing a chemical inhibitor of non-homologous end joining (NHEJ) into the region of target nucleic acid, or by introducing a CRISPR guide nucleic acid or an siRNA that targets NHEJ protein to transiently knock down the expression of NHEJ protein.

[0113] In some embodiments, the Type II or Type V CRISPR-Cas effector protein may be a fusion protein, and / or the reverse transcriptase may be a fusion protein, and the Type II or Type V CRISPR-Cas fusion protein, reverse transcriptase fusion protein, and / or extended guide nucleic acid may be fused to one or more components that allow for the recruitment of the reverse transcriptase to the Type II or Type V CRISPR-Cas effector protein. In some embodiments, the one or more components are recruited via protein-protein interactions, protein-RNA interactions, and / or chemical interactions.

[0114] Thus, in some embodiments, the type V CRISPR-Cas effector protein may be a type V CRISPR-Cas effector fusion protein comprising a type V CRISPR-Cas effector protein domain fused (linked) to a peptide tag (e.g., an epitope or a multimerized epitope), the reverse transcriptase may be a reverse transcriptase fusion protein comprising a reverse transcriptase domain fused (linked) to an affinity polypeptide that binds to the peptide tag, and the type V CRISPR-Cas effector protein interacts with a guide nucleic acid that is bound to the target nucleic acid, thereby recruiting the reverse transcriptase to the type V CRISPR-Cas effector protein and the target nucleic acid. In some embodiments, the type II CRISPR-Cas effector protein is a type II CRISPR-Cas fusion protein comprising a type II CRISPR-Cas effector protein domain fused (linked) to a peptide tag (e.g., an epitope or a multimerized epitope); the FEN is a FEN fusion protein comprising a FEN domain fused to an affinity polypeptide that binds to the peptide tag; and / or the type II CRISPR-Cas effector protein is a type II CRISPR-Cas fusion protein comprising a type II CRISPR-Cas effector protein domain fused to a peptide tag; and the Dna2 polypeptide is a Dna2 fusion protein comprising a Dna2 domain fused to an affinity polypeptide that binds to the peptide tag; and optionally, the target nucleic acid is contacted with two or more FEN fusion proteins and / or two or more Dna2 fusion proteins, thereby recruiting FEN and / or Dna2 to the type II CRISPR-Cas effector protein domain and the target nucleic acid. In some embodiments, two or more reverse transcriptase fusion proteins are recruited to a Type II or Type V CRISPR-Cas effector protein, thereby contacting the target nucleic acid with the two or more reverse transcriptase fusion proteins.

[0115] Peptide tags may include, but are not limited to, a GCN4 peptide tag (e.g., Sun tag), a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or strep tag II, a V5 tag, and / or a VSV-G epitope. Any epitope that can be linked to a polypeptide and for which there is a corresponding affinity polypeptide that can be linked to another polypeptide may be used in the present invention. In some embodiments, the peptide tag can comprise one or two or more copies of the peptide tag (e.g., epitope, multimerized epitope (e.g., tandem repeats)) (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or more peptide tags. In some embodiments, the affinity polypeptide that binds to the peptide tag can be an antibody. In some embodiments, the antibody can be an scFv antibody. In some embodiments, the affinity polypeptide that binds to the peptide tag can be synthetic (e.g., evolved for affinity interaction), including, but not limited to, an affibody, anticalin, monobody, and / or DARPin (see, for example, Sha et al., Protein Research, 2009, 101:1047-1052, 2009, 102:1047-1052, 2009, 103:1057-1062, 2010, 104:1057-1062, 2010, 105:1057-1062, 2010, 106 ... Sci., 26(5):910-924(2017)), Gilbreth (Curr Opin Struc Biol, 22(4):413-420(2013)), U.S. Patent No. 9,982,053). Examples of peptide tag sequences and affinity polypeptides thereof include, but are not limited to, the amino acid sequences of SEQ ID NOs: 23-25.

[0116] In some embodiments, the extended guide nucleic acid can be linked to an RNA recruitment motif, and the reverse transcriptase can be a reverse transcriptase fusion protein, which can include a reverse transcriptase domain fused to an affinity polypeptide that binds the RNA recruitment motif, where the extended guide binds to the target nucleic acid and the RNA recruitment motif binds to the affinity polypeptide, thereby recruiting the reverse transcriptase fusion protein to the extended guide and contacting the target nucleic acid with the reverse transcriptase domain. In some embodiments, two or more reverse transcriptase fusion proteins can be recruited to the extended guide nucleic acid, thereby contacting the target nucleic acid with the two or more reverse transcriptase fusion proteins. Example RNA recruitment motifs and affinity polypeptides thereof include, but are not limited to, the sequences of SEQ ID NOs: 26-36.

[0117] In some embodiments, the RNA recruitment motif may be located on the 3' end of the extension portion of the extended guide nucleic acid (e.g., 5'-3' repeat-spacer-extension portion (RT template-primer binding site)-RNA recruitment motif). In some embodiments, the RNA recruitment motif may be embedded within the extension portion.

[0118] In some embodiments of the invention, the extended guide RNA and / or guide RNA may be linked to one or more RNA recruitment motifs (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more motifs, e.g., at least 10 to about 25 motifs), and optionally the two or more RNA recruitment motifs may be the same RNA recruitment motif or different RNA recruitment motifs. In some embodiments, the RNA recruitment motif and corresponding affinity polypeptide may include, but are not limited to, a telomerase Ku-binding motif (e.g., a Ku-binding hairpin) and a corresponding affinity polypeptide Ku (e.g., a Ku heterodimer), a telomerase Sm7-binding motif and a corresponding affinity polypeptide Sm7, an MS2 phage operator stem-loop and a corresponding affinity polypeptide MS2 coat protein (MCP), a PP7 phage operator stem-loop and a corresponding affinity polypeptide PP7 coat protein (PCP), an SfMu phage Com stem-loop and a corresponding affinity polypeptide Com RNA-binding protein, a PUF binding site (PBS) and an affinity polypeptide Pumilio / fem-3 mRNA-binding factor (PUF), and / or a synthetic RNA-aptamer and a corresponding affinity polypeptide as an aptamer ligand. In some embodiments, the RNA recruitment motif and corresponding affinity polypeptide may be an MS2 phage operator stem-loop and an affinity polypeptide MS2 coat protein (MCP). In some embodiments, the RNA recruitment motif and corresponding affinity polypeptide can be a PUF binding site (PBS) and an affinity polypeptide Pumilio / fem-3 mRNA binding factor (PUF).

[0119] In some embodiments, components for recruiting polypeptides and nucleic acids may function through chemical interactions, including, but not limited to, rapamycin-induced dimerization of FRB-FKBP, biotin-streptavidin, SNAP-tag, Halo-tag, CLIP-tag, compound-induced DmrA-DmrC heterodimers, and bifunctional ligands (e.g., two protein-binding chemicals fused together, such as dihydrofolate reductase (DHFR)).

[0120] In some embodiments of the invention, the CRISPR-Cas effector protein (e.g., the CRISPR-Cas effector protein, the first CRISPR-Cas effector protein, the second CRISPR-Cas effector protein, the third CRISPR-Cas effector protein, and / or the fourth CRISPR-Cas effector protein) may be derived from a Type I CRISPR-Cas system, a Type II CRISPR-Cas system, a Type III CRISPR-Cas system, a Type IV CRISPR-Cas system, and / or a Type V CRISPR-Cas system. In some embodiments, the CRISPR-Cas nuclease is derived from a Type II CRISPR-Cas system or a Type V CRISPR-Cas system.

[0121] In some embodiments of the invention, the CRISPR-Cas effector protein is selected from the group consisting of Cas9, C2c1, C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csx12), Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, and / or Csf5 nuclease, and optionally the CRISPR-Cas nuclease is C The nuclease may be as9, Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c.

[0122] In some embodiments, a CRISPR-Cas effector protein can be a protein that functions as a nickase (e.g., a Cas9 nickase or a Cas12a nickase). In some embodiments, a CRISPR-Cas effector protein useful in the present invention can have a mutation in its nuclease active site (e.g., a RuvC, HNH, e.g., a RuvC site in a Cas12a nuclease domain, e.g., a RuvC and / or HNH site in a Cas9 nuclease domain). A CRISPR-Cas effector protein that has a mutation in its nuclease active site and therefore no longer contains nuclease activity is commonly referred to as "dead" or "inactivated," e.g., dCas. In some embodiments, a CRISPR-Cas nuclease domain or polypeptide that has a mutation in its nuclease active site can have impaired or reduced activity compared to the same CRISPR-Cas nuclease without the mutation. In some embodiments, the CRISPR-Cas effector protein useful in the present invention can be a double-stranded nuclease. In some embodiments, the CRISPR-Cas effector protein with double-stranded nuclease activity can be a type II or type V CRISPR-Cas effector protein. In some embodiments, the type V CRISPR-Cas effector protein with double-stranded nuclease activity is a Cas12a polypeptide. In some embodiments, the type II CRISPR-Cas effector protein with double-stranded nuclease activity is a Cas9 polypeptide.

[0123] In some embodiments, the CRISPR-Cas effector protein may be a type V CRISPR-Cas effector protein. In some embodiments, the type V CRISPR-Cas effector protein may include Cas12a (Cpf1), Cas12b (C2c1), Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c effector proteins and / or domains.

[0124] In some embodiments, type V CRISPR-Cas systems may include effector proteins that utilize only type V CRISPR nucleic acids. In some embodiments, type V CRISPR-Cas systems may include effector proteins that utilize both CRISPR nucleic acids and trans-activating CRISPR (tracr) nucleic acids, similar to type II CRISPR-Cas systems. Thus, in some embodiments, type V CRISPR-Cas effector proteins useful in the present invention may function only with the corresponding CRISPR nucleic acid (e.g., Cas12a, Cas12a, Cas12i, Cas12h, Cas14b, Cas14c, C2c10, C2c9, C2c8, C2c4). In some embodiments, type V CRISPR-Cas effector proteins useful in the present invention may function with the corresponding CRISPR nucleic acid and tracr nucleic acid (e.g., Cas12b, Cas12c, Cas12e, Cas12g, Cas14a).

[0125] CRISPR nucleic acids useful in the present invention may comprise at least one repeat sequence capable of interacting with a corresponding type V CRISPR-Cas effector protein and at least one spacer sequence, wherein the at least one spacer sequence can bind to a target nucleic acid (e.g., the first or second strand of the target nucleic acid). In some embodiments, the repeat sequence of the CRISPR nucleic acid may be located 5' from the spacer sequence. In some embodiments, the CRISPR nucleic acid may comprise multiple repeat sequences, wherein the repeat sequence is linked to both the 5' and 3' ends of the spacer. In some embodiments, the CRISPR nucleic acid useful in the present invention may comprise two or more repeats and one or more spacer sequences, wherein each spacer sequence is linked to a repeat sequence at the 5' and 3' ends.

[0126] A tracr nucleic acid useful in the invention can comprise a first portion that is substantially complementary to and hybridizes with a repeat sequence of a corresponding CRISPR nucleic acid, and a second portion that interacts with a corresponding Type II or Type V CRISPR-Cas effector protein.

[0127] In some embodiments, the type V CRISPR-Cas effector protein useful in the present invention can function as a double-stranded DNA nuclease. In some embodiments, the type V CRISPR-Cas effector protein can function as a single-stranded DNA nickase, optionally with the first strand nicked. In some embodiments, the type V CRISPR-Cas effector protein can function as a single-stranded DNA nickase, optionally with the second strand nicked. In some embodiments, the type V CRISPR-Cas effector protein can be a Cas12a effector protein that functions as a nickase, optionally with the first strand (target strand) nicked. In some embodiments, the type V CRISPR-Cas effector protein can be a Cas12a effector protein that functions as a nickase, optionally with the second strand nicked.

[0128] In some embodiments, the Cas12a effector protein is LQM R The Cas12a nickase may have an arginine mutation in the NS motif. The arginine mutation in this motif may be to any amino acid, thereby providing a Cas12a nickase. In some embodiments, the mutation may be to alanine. In some embodiments, the mutation may be to alanine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine. In some embodiments, the mutation may be to alanine. In some embodiments, the mutation does not include a mutation to lysine or histidine. In some embodiments, the Cas12a effector protein can be an LbCas12a nickase comprising an R1138, optionally an R1138A mutation (see reference nucleotide sequence, SEQ ID NO:9), an R1137, optionally an R1137A mutation (see reference nucleotide sequence, SEQ ID NO:1), or an R1124, optionally an R1124A mutation (see reference nucleotide sequence, SEQ ID NO:7). In some embodiments, the Cas12a effector protein can be an AsCas12a nickase comprising an R1226, optionally an R1226A mutation (see reference nucleotide sequence, SEQ ID NO:2). In some embodiments, the Cas12a effector protein can be an FnCas12a nickase (see reference nucleotide sequence, SEQ ID NO: 6) comprising an R1218 mutation, optionally an R1218A mutation. In some embodiments, the Cas12a effector protein can be a PdCas12a nickase (see reference nucleotide sequence, SEQ ID NO: 14) comprising an R1241 mutation, optionally an R1241A mutation.

[0129] In some embodiments, type V CRISPR-Cas effector proteins useful in the invention may comprise reduced single-strand DNA cleavage activity (ssDNAse activity) (e.g., type V CRISPR-Cas effector proteins may be modified (mutated) to have reduced ssDNAse activity (e.g., about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% less ssDNAse activity than a wild-type or unmodified type V CRISPR-Cas effector protein).

[0130] In some embodiments, type V CRISPR-Cas effector proteins useful in the invention may comprise reduced self-processing RNAse activity (e.g., type V CRISPR-Cas effector proteins may be modified (mutated) to have reduced self-processing RNAse activity (e.g., about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% less self-processing RNAse activity than a wild-type or unmodified type V CRISPR-Cas effector protein). In some embodiments, the mutation to reduce self-processing RNAse activity may be a mutation of the histidine at residue position 759 with reference to the numbering of the nucleotide positions in SEQ ID NO:9, optionally a histidine to alanine mutation (H759A).

[0131] In some embodiments, a V-type CRISPR-Cas effector protein or domain useful in the present invention may contain a mutation in its nuclease active site (e.g., the RuvC site of a dV-type CRISPR-Cas effector protein or domain, e.g., a Cas12a nuclease domain). A CRISPR-Cas nuclease that has a mutation in its nuclease active site and therefore no longer contains nuclease activity is commonly referred to as "inactivated" or "dead," e.g., dCas, dCas12a. In some embodiments, a CRISPR-Cas nuclease domain or polypeptide that has a mutation in its nuclease active site may have impaired or reduced activity compared to the same CRISPR-Cas nuclease without the mutation. In some embodiments, an inactivated V-type CRISPR-Cas effector protein may function as a nickase (first-strand nickase and / or second-strand nickase).

[0132] In some embodiments, the type V CRISPR-Cas effector protein may be a type V CRISPR-Cas fusion protein, wherein the type V CRISPR-Cas fusion protein comprises a type V CRISPR-Cas effector protein domain fused to a reverse transcriptase. In some embodiments, the reverse transcriptase may be fused to the C-terminus of the type V CRISPR-Cas effector polypeptide. In some embodiments, the reverse transcriptase may be fused to the N-terminus of the type V CRISPR-Cas effector polypeptide.

[0133] In some embodiments, the type V CRISPR-Cas effector protein may be a type V CRISPR-Cas fusion protein, wherein the type V CRISPR-Cas fusion protein comprises a type V CRISPR-Cas effector protein domain fused to a nicking enzyme (e.g., Fok1, BFi1, e.g., engineered Fok1 or BFiI), and optionally the type V CRISPR-Cas effector protein domain may be an inactivated type V CRISPR-Cas domain fused to a nicking enzyme (e.g., Fok1, BFi1, e.g., engineered Fok1 or BFiI).

[0134] In some embodiments, the type II CRISPR-Cas effector protein may be a type II CRISPR-Cas fusion protein, wherein the type II CRISPR-Cas fusion protein comprises a type II CRISPR-Cas effector protein domain fused to a reverse transcriptase. In some embodiments, the reverse transcriptase may be fused to the C-terminus of the type II CRISPR-Cas effector polypeptide. In some embodiments, the reverse transcriptase may be fused to the N-terminus of the type II CRISPR-Cas effector polypeptide. In some embodiments, the type II CRISPR-Cas effector protein may be a type II CRISPR-Cas fusion protein, wherein the type II CRISPR-Cas fusion protein comprises a type II CRISPR-Cas effector protein domain fused to a nicking enzyme (e.g., Fokl, BFiI, e.g., engineered Fokl or BFiI), and optionally, the type II CRISPR-Cas effector protein domain may be an inactivated type II CRISPR-Cas domain fused to a nicking enzyme.

[0135] In some embodiments, the reverse transcriptase useful in the present invention can be a wild-type reverse transcriptase. In some embodiments, the reverse transcriptase useful in the present invention can be a synthetic reverse transcriptase (see, e.g., Heller et al., Nucleic Acids Research, 47(7), 3619-3630 (2019)).

[0136] In some embodiments, reverse transcriptases useful in the present invention can be modified to improve the transcription function of the reverse transcriptase. The transcription function of the reverse transcriptase can be improved by improving the processivity of the reverse transcriptase, for example, by increasing the ability of the reverse transcriptase to polymerize more DNA bases during a single binding event to the template (e.g., before disengaging from the template) (e.g., increasing processivity by about 5, 10, 15, 20, 25, 30, 345, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% compared to an unmodified reference reverse transcriptase).

[0137] In some embodiments, the transcription function of a reverse transcriptase may be improved by improving the template affinity of the reverse transcriptase (e.g., increasing template affinity by about 5, 10, 15, 20, 25, 30, 345, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% compared to an unmodified reference reverse transcriptase).

[0138] In some embodiments, the transcription function of a reverse transcriptase may be improved by improving the thermostability of the reverse transcriptase for improved performance at a desired temperature (e.g., increasing thermostability by about 5, 10, 15, 20, 25, 30, 345, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% compared to an unmodified reference reverse transcriptase). In some embodiments, the improved thermostability is at a temperature of about 20°C to 42°C (e.g., about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, or 42°C, and any value or range therein). In some embodiments, reverse transcriptases with improved thermostability may include, but are not limited to, the M-MuLV triple mutant D200N+L603W+T330P or the M-MuLV quintuple mutant D200N+L603W+T330P+T306K+W313F (Reference Sequence, SEQ ID NO: 53). See, e.g., Baranauskas et al., Protein Eng. Des. Sel., 25, 657-668 (2012) and Anzalone et al., Nature, 576:149-157 (2019)).

[0139] In some embodiments of the present invention, the reverse transcriptase may be fused to one or more single-stranded RNA-binding domains (RBDs). RBDs useful in the present invention may include, but are not limited to, those set forth in SEQ ID NOs: 37-52 (SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51, and / or SEQ ID NO: 52), which improve the thermostability, processivity, and template affinity of the reverse transcriptase.

[0140] In some embodiments, the activity of the reverse transcriptase may be improved for (Type V or Type II) gene editing activity to provide optimal activity when associated with a Type V or Type II CRISPR-Cas effector polypeptide (e.g., about a 5, 10, 15, 20, 25, 30, 345, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% increase in activity when associated with a Type V CRISPR-Cas effector polypeptide compared to an unmodified reference reverse transcriptase). Such mutations include those that affect or improve RT initiation, processivity, enzyme kinetics, temperature sensitivity, and / or error rate.

[0141] The polypeptides / proteins / domains of the invention (e.g., CRISPR-Cas effector proteins, e.g., type II or type V CRISPR-Cas effector proteins), reverse transcriptases, 5' flap endonucleases, and / or 5'-3' exonucleases) may be encoded by one or more polynucleotides, optionally operably linked to one or more promoters and / or other regulatory sequences (e.g., terminators, operons, and / or enhancers, etc.). In some embodiments, the polynucleotides of the invention may be comprised in one or more expression cassettes and / or vectors. In some embodiments, the at least one regulatory sequence may be, for example, a promoter, operon, terminator, or enhancer. In some embodiments, the at least one regulatory sequence may be a promoter. In some embodiments, the regulatory sequence may be an intron. In some embodiments, the at least one regulatory sequence may be, for example, a promoter operably associated with an intron or a promoter region including an intron. In some embodiments, the at least one regulatory sequence may be, for example, a ubiquitin promoter and its associated intron (e.g., Medicago truncatula and / or maize and its associated intron) (e.g., a promoter containing an intron, such as ZmUbi1, MtUb2, e.g., SEQ ID NO: 21 or 22, or an intron of SEQ ID NO: 74 or 75).

[0142] In some embodiments, the invention provides polynucleotides encoding a Type II or Type V CRISPR-Cas effector protein or domain, a polynucleotide encoding a CRISPR-Cas effector protein or domain, a polynucleotide encoding a reverse transcriptase polypeptide or domain, a polynucleotide encoding a 5'-3' exonuclease polypeptide or domain, and / or a polynucleotide encoding a flap endonuclease polypeptide or domain, operably associated with one or more promoter regions comprising or associated with an intron, optionally wherein the promoter region can be a ubiquitin promoter and intron (e.g., a promoter comprising a Medicago or maize ubiquitin promoter and intron, e.g., SEQ ID NO: 21 or 22, or the intron of SEQ ID NO: 74 or 75).

[0143] In some embodiments, the polynucleotide encoding a type II or type V CRISPR-Cas effector protein and / or the polynucleotide encoding the reverse transcriptase may be included in the same or separate expression cassettes, and when included in the same expression cassette as the polynucleotide encoding a type II or type V CRISPR-Cas effector protein and the polynucleotide encoding the reverse transcriptase, the polynucleotide encoding the type II or type V CRISPR-Cas effector protein and the polynucleotide encoding the reverse transcriptase may be operably linked to a single promoter or two or more separate promoters in any combination. In some embodiments, the polynucleotide encoding the CRISPR-Cas effector protein may be included in an expression cassette, and the polynucleotide encoding the CRISPR-Cas effector protein may be operably linked to a promoter.

[0144] In some embodiments, the extended guide nucleic acid and / or the guide nucleic acid may be comprised in an expression cassette, and optionally the expression cassette is comprised in a vector. In some embodiments, the expression cassette and / or vector comprising the extended guide nucleic acid may be the same or a different expression cassette and / or vector as those comprising the polynucleotide encoding a Type II or Type V CRISPR-Cas effector protein and / or the polynucleotide encoding a reverse transcriptase. In some embodiments, the expression cassette and / or vector comprising the guide nucleic acid may be the same or a different expression cassette and / or vector as those comprising the polynucleotide encoding a CRISPR-Cas effector protein.

[0145] In some embodiments, the polynucleotide encoding the 5' flap endonuclease and / or the polynucleotide encoding the 5'-3' exonuclease may be included in one or more expression cassettes, which may be the same or different expression cassettes. In some embodiments, the expression cassette comprising the polynucleotide encoding the 5' flap endonuclease and / or the polynucleotide encoding the 5'-3' exonuclease may be the same or a different expression cassette as the polynucleotide encoding a type II or type V CRISPR-Cas effector protein, the polynucleotide encoding a type II or type V CRISPR-Cas effector protein, and / or the polynucleotide encoding a reverse transcriptase.

[0146] In some embodiments of the invention, polynucleotides encoding CRISPR-Cas effector proteins (e.g., Type II CRISPR-Cas effector proteins, Type V CRISPR-Cas effector proteins), reverse transcriptases, flap endonucleases, 5'-3' exonucleases, and fusion proteins comprising same, as well as nucleic acid constructs, expression cassettes, and / or vectors comprising the polynucleotides, may be codon-optimized for expression in an organism (e.g., an animal (e.g., a mammal, an insect, a fish, etc.), a plant (e.g., a dicotyledonous plant, a monocotyledonous plant), a bacterium, an archaea, etc.). In some embodiments, the polynucleotides, expression cassettes, and / or vectors may be codon-optimized for expression in a plant, optionally a dicotyledonous or monocotyledonous plant. Exemplary mammals in which the present invention may be useful include, but are not limited to, primates (human and non-human (e.g., chimpanzees, baboons, monkeys, gorillas, etc.)), cats, dogs, ferrets, gerbils, hamsters, cows, pigs, horses, goats, donkeys, or sheep.

[0147] In some embodiments, a polynucleotide, nucleic acid construct, expression cassette, or vector of the invention that has been optimized for expression in an organism may be about 70% to 100% identical (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%) to a nucleic acid construct, expression cassette, or vector encoding it that has not been codon-optimized for expression in a plant.

[0148] In some embodiments, polynucleotides, nucleic acid constructs, expression cassettes, and vectors can be provided for carrying out the methods of the present invention. Thus, in some embodiments, an expression cassette is provided, which is codon-optimized for expression in an organism, and includes, from 5' to 3', (a) a polynucleotide encoding a promoter sequence; (b) a polynucleotide encoding a type V CRISPR-Cas nuclease (e.g., Cpf1 (Cas12a), dCas12a, etc.) or a type II CRISPR-Cas nuclease (e.g., Cas9, dCas9, etc.), which is codon-optimized for expression in an organism; (c) a linker sequence; and (d) a polynucleotide encoding a reverse transcriptase, which is codon-optimized for expression in an organism. In some embodiments, the organism is an animal, a plant, a fungus, an archaea, or a bacterium. In some embodiments, the organism is a plant, the polynucleotide encoding the Type V CRISPR-Cas nuclease is codon-optimized for expression in the plant, and the promoter sequence is a plant-specific promoter sequence (e.g., ZmUbi1, MtUb2, RNA polymerase II (Pol II)).

[0149] In some embodiments, polynucleotides, nucleic acid constructs, expression cassettes, and vectors may be provided for carrying out the methods of the present invention. Thus, in some embodiments, an expression cassette is provided that is codon-optimized for expression in plants and includes, from 5' to 3', (a) a polynucleotide encoding a plant-specific promoter sequence (e.g., ZmUbi1, MtUb2, RNA polymerase II (Pol II)), (b) a plant-codon-optimized polynucleotide encoding a type II or type V CRISPR-Cas effector protein (e.g., Cpf1 (Cas12a), dCas12a, etc.), (c) a linker sequence, and (d) a plant-codon-optimized polynucleotide encoding a reverse transcriptase.

[0150] In some embodiments, the polypeptides of the present invention can be fusion proteins comprising one or more polypeptides linked to each other via a linker. In some embodiments, the linker can be an amino acid or peptide linker. In some embodiments, the peptide linker can be from about 2 to about 100 amino acids (residues) in length, as described herein. In some embodiments, the peptide linker can be, for example, a GS linker.

[0151] In some embodiments, the present invention provides expression cassettes that are codon-optimized for expression in plants and include: (a) a polynucleotide encoding a plant-specific promoter sequence (e.g., ZmUbi1, MtUb2); and (b) an extended guide nucleic acid sequence that includes an extension portion at its 3' end that includes a primer binding site and an edit (e.g., a reverse transcriptase template) to be incorporated into the target nucleic acid (e.g., 5'-3'-crRNA-RTT-PBS), optionally included in the expression cassette and optionally operably linked to a Pol II promoter. In some embodiments, when the extended portion of the guide nucleic acid is attached to the CRISPR RNA at its 5' end, the extension portion includes a primer binding site at its 5' end and an edit (e.g., a reverse transcriptase template) to be incorporated into the target nucleic acid at its 3' end (5'-3'-PBS-RTT-crRNA).

[0152] In some embodiments, the expression cassettes of the present invention may be codon-optimized for expression in dicotyledonous plants or monocotyledonous plants. In some embodiments, the expression cassettes of the present invention may be used in a method for modifying a target nucleic acid in a plant or plant cell, comprising introducing one or more expression cassettes of the present invention into a plant or plant cell, thereby modifying the target nucleic acid in the plant or plant cell to produce a plant or plant cell comprising the modified target nucleic acid. In some embodiments, the method may further comprise regenerating the plant cell comprising the modified target nucleic acid to produce a plant comprising the modified target nucleic acid.

[0153] A CRISPR Cas9 polypeptide or CRISPR Cas9 domain (e.g., a Type II CRISPR Cas9 effector protein) useful in the present invention can be any known or later identified Cas9 nuclease. In some embodiments, the CRISPR Cas9 polypeptide can be, for example, a Cas9 polypeptide from a Streptococcus species (e.g., S. pyogenes, S. thermophilus 9), a Lactobacillus species, a Bifidobacterium species, a Kandleria species, a Leuconostoc species, an Oenococcus species, a Pediococcus species, a Weissella species, and / or an Olsenella species.

[0154] Cas12a is a type V clustered regularly interspaced short palindromic repeats (CRISPR)-Cas effector protein or domain. Cas12a differs from the more well-known type II CRISPR Cas9 effector protein in several respects. For example, Cas9 recognizes a G-rich protospacer adjacent motif (PAM) located 3' of its guide RNA (gRNA, sgRNA) binding site (protospacer, target nucleic acid, target DNA) (3'-NGG), whereas Cas12a recognizes a T-rich PAM located 5' of the target nucleic acid (5'-TTN, 5'-TTTN). In fact, the orientations in which Cas9 and Cas12a bind their guide RNAs are nearly reversed with respect to their N- and C-termini. Furthermore, the Cas12a effector protein is distinct from the two PAMs found in the native Cas9 system. Cas12a uses a single guide RNA (gRNA, CRISPR array, crRNA) rather than a heavy guide RNA (sgRNA (e.g., crRNA and tracrRNA)), and processes its own gRNA. Furthermore, the nuclease activity of Cas12a produces a staggered double-stranded DNA break instead of the blunt ends produced by the nuclease activity of Cas9, and Cas12a relies on a single RuvC domain to cleave both DNA strands, while Cas9 utilizes an HNH domain and a RuvC domain for cleavage.

[0155] A CRISPR Cas12a effector protein or domain useful in the present invention can be any known or later identified Cas12a nuclease (formerly known as Cpf1) (see, e.g., U.S. Patent No. 9,790,490, which is incorporated by reference for its disclosure of the Cpf1 (Cas12a) sequence). The terms "Cas12a," "Cas12a polypeptide," or "Cas12a domain" refer to an RNA-guided effector protein comprising Cas12a, or a fragment thereof comprising the guide nucleic acid binding domain of Cas12a and / or an active, inactive, or partially active DNA cleavage domain of Cas12a. In some embodiments, a Cas12a useful in the present invention can contain a mutation in the nuclease active site (e.g., the RuvC site of a Cas12a domain). A Cas12a effector protein or domain that has a mutation in its nuclease active site and therefore no longer contains nuclease activity is commonly referred to as dead or inactivated Cas12a (e.g., dCas12a).

[0156] In some embodiments, Cas12a effector polypeptides that may be optimized or otherwise modified (e.g., inactivated) in accordance with the present invention can include, but are not limited to, the amino acid sequence of any one of SEQ ID NOs: 1-20 (e.g., SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20), or a polynucleotide encoding same.

[0157] As used herein, "guide nucleic acid," "guide RNA," "gRNA," "CRISPR RNA / DNA," "crRNA," or "crDNA" refers to a nucleic acid that includes at least one spacer sequence that is complementary to (and hybridizes with) a target DNA (e.g., a protospacer) and at least one repeat sequence that corresponds to a specific CRISPR-Cas effector protein (e.g., for a type V CRISPR Cas effector protein, the repeat or a fragment or portion thereof is derived from a type V Cas12a CRISPR-Cas system, and for a type II CRISPR Cas effector protein, the repeat or a fragment or portion thereof is derived from a type II Cas9 CRISPR-Cas system). Thus, repeats of CRISPR-Cas systems useful in the present invention include, for example, Cas9, C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csx12), Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl , Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG), and / or Csf5 CRISPR-Cas effector proteins, or fragments thereof, and the repeat sequence may be linked to the 5' and / or 3' end of the spacer sequence. The design of the guide nucleic acid of the present invention may be based on Type I, Type II, Type III, Type IV, or Type V CRISPR-Cas systems. In some embodiments, the design of the guide nucleic acid of the present invention is based on Type V CRISPR-Cas systems.

[0158] In some embodiments, a Cas12a guide nucleic acid or extended guide nucleic acid may comprise, from 5' to 3', a repeat sequence (either full length or a portion thereof ("handle"), e.g., a pseudoknot-like structure) and a spacer sequence.

[0159] In some embodiments, a guide nucleic acid can include multiple repeat sequence-spacer sequences (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more repeat-spacer sequences) (e.g., repeat-spacer-repeat, e.g., repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer, etc.). The guide nucleic acids of the present invention are synthetic, artificial, and not found in nature. Guide nucleic acids can be quite long and can be used as aptamers (as in MS2 recruitment strategies) or other RNA structures that hang off a spacer. In some embodiments, as described herein, a guide nucleic acid can include a template for editing and a primer binding site. In some embodiments, a guide nucleic acid can include a region or sequence on the 5' or 3' end that is complementary to the editing template (reverse transcriptase template) and thereby recruits the editing template to the target nucleic acid (i.e., extended guide nucleic acid). In some embodiments, a guide nucleic acid may include a region or sequence on its 5' or 3' end that is complementary to a primer on a target nucleic acid (the primer binding site), thereby recruiting the primer binding site to the target nucleic acid (i.e., the extended guide nucleic acid).

[0160] As used herein, "repeat sequence" refers to, for example, any repeat sequence of a wild-type CRISPR Cas locus (e.g., Cas9 locus, Cas12a locus, C2c1 locus, etc.) or a repeat sequence of a synthetic crRNA functional in a CRISPR-Cas nuclease encoded by a nucleic acid construct of the present invention. Repeat sequences useful in the present invention can be any known or later identified repeat sequence of a CRISPR-Cas locus (e.g., Type I, Type II, Type III, Type IV, Type V, or Type VI), or can be synthetic repeats designed to function in a Type I, II, III, IV, V, or VI CRISPR-Cas system. Thus, in some embodiments, the repeat sequence can be identical or substantially identical to a repeat sequence from a wild-type type I CRISPR-Cas locus, a type II CRISPR-Cas locus, a type III CRISPR-Cas locus, a type IV CRISPR-Cas locus, a type V CRISPR-Cas locus, and / or a type VI CRISPR-Cas locus. In some embodiments, the repeat sequence useful in the present invention can be any known or later identified repeat sequence from a type V CRISPR-Cas locus, or can be a synthetic repeat designed to function in a type V CRISPR-Cas system. The repeat sequence can include a hairpin structure and / or a stem-loop structure. In some embodiments, the repeat sequence can form a pseudoknot-like structure at its 5' end (i.e., a "handle"). Thus, in some embodiments, the repeat sequence can be identical or substantially identical to a repeat sequence from a wild-type type V CRISPR-Cas locus or a wild-type type II CRISPR-Cas locus. Repeat sequences from wild-type CRISPR-Cas loci can be determined through established algorithms, such as using CRISPRfinder, provided through CRISPRdb (see Grissa et al., Nucleic Acids Res., 35(Web Server Issue):W52-7 or BMC Informatics, 8:172(2007)(doi:10.1186 / 1471-2105-8-172)).In some embodiments, the repeat sequence or a portion thereof is linked at its 3' end to the 5' end of a spacer sequence, thereby forming a repeat-spacer sequence (e.g., guide RNA, crRNA).

[0161] In some embodiments, the repeat sequence comprises, consists essentially of, or consists of at least 10 nucleotides (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 to 100 or more nucleotides, or any range or value therein, e.g., about), depending on whether the particular repeat and the guide RNA comprising the repeat is processed or unprocessed. In some embodiments, the repeat sequence comprises, consists essentially of, or consists of about 10 to about 20, about 10 to about 30, about 10 to about 45, about 10 to about 50, about 15 to about 30, about 15 to about 40, about 15 to about 45, about 15 to about 50, about 20 to about 30, about 20 to about 40, about 20 to about 50, about 30 to about 40, about 40 to about 80, about 50 to about 100, or more nucleotides.

[0162] The repeat sequence linked to the 5' end of the spacer sequence can include a portion of the repeat sequence (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, or more consecutive nucleotides of the wild-type repeat sequence). In some embodiments, the portion of the repeat sequence linked to the 5' end of the spacer sequence can be about 5 to about 10 contiguous nucleotides in length (e.g., about 5, 6, 7, 8, 9, 10 nucleotides) and have at least 90% identity (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more) to the same region (e.g., the 5' end) of a wild-type CRISPR Cas repeat nucleotide sequence. In some embodiments, the portion of the repeat sequence can include a pseudoknot-like structure at its 5' end (e.g., a "handle").

[0163] As used herein, a "spacer sequence" refers to a nucleotide sequence that is complementary to a target nucleic acid (e.g., a target DNA) (e.g., a protospacer). The spacer sequence can be fully complementary or substantially complementary (e.g., at least about 70% complementary (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or higher)) to the target nucleic acid. Thus, in some embodiments, the spacer sequence can have 1, 2, 3, 4, or 5 mismatches compared to the target nucleic acid, and the mismatches can be consecutive or non-consecutive. In some embodiments, the spacer sequence can have 70% complementarity to the target nucleic acid. In other embodiments, the spacer nucleotide sequence can have 80% complementarity to the target nucleic acid. In still other embodiments, the spacer nucleotide sequence can have 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or higher complementarity to the target nucleic acid (protospacer), etc. In some embodiments, the spacer sequence is 100% complementary to the target nucleic acid. The spacer sequence can have a length of about 15 nucleotides to about 30 nucleotides (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides, or any range or value therein). Thus, in some embodiments, the spacer sequence can have perfect complementarity or substantial complementarity over a region of the target nucleic acid (e.g., protospacer) that is at least about 15 nucleotides to about 30 nucleotides in length. In some embodiments, the spacer is about 20 nucleotides in length. In some embodiments, the spacer is about 23 nucleotides in length.

[0164] In some embodiments, the 5' region of the spacer sequence of the guide RNA may be identical to the target DNA, while the 3' region of the spacer may be substantially complementary to the target DNA (e.g., Type V CRISPR-Cas), or the 3' region of the spacer sequence of the guide RNA may be identical to the target DNA, while the 5' region of the spacer may be substantially complementary to the target DNA (e.g., Type II CRISPR-Cas), such that the overall complementarity of the spacer sequence to the target DNA may be less than 100%. Thus, for example, in a Type V CRISPR-Cas system guide, for example, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in the 5' region of a 20-nucleotide spacer sequence (i.e., the seed region) may be 100% complementary to the target DNA, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary (e.g., at least about 70% complementary) to the target DNA. In some embodiments, the first 1 to 8 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8 nucleotides, and any range therein) at the 5' end of the spacer sequence can be 100% complementary to the target DNA, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary (e.g., at least about 50% complementary (e.g., 50%, 55% complementary) to the target DNA). , 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or higher).

[0165] As a further example, in a guide for a Type II CRISPR-Cas system, for example, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides in the 3' region (i.e., seed region) of a 20-nucleotide spacer sequence can be 100% complementary to the target DNA, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary (e.g., at least about 70% complementary) to the target DNA. In some embodiments, the first 1-10 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides, and any range therein) at the 3' end of the spacer sequence can be 100% complementary to the target DNA, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary (e.g., at least about 50% complementary (e.g., at least about 50%, 55%, 60%, 75%, 80%, 95%, 10 ... %, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or higher, or any range or value therein).

[0166] In some embodiments, the seed region of the spacer can be about 8 to about 10 nucleotides in length, about 5 to about 6 nucleotides in length, or about 6 nucleotides in length.

[0167] In some embodiments, the extended guide nucleic acid may be an extended guide nucleic acid, a first extended guide nucleic acid, and / or a second extended guide nucleic acid. In some embodiments, the extended guide nucleic acid useful in the present invention may include (a) a CRISPR nucleic acid (e.g., CRISPR RNA, CRISPR DNA, crRNA, crDNA) and / or a CRISPR nucleic acid and a tracr nucleic acid, and (b) an extension portion including a primer binding site and a reverse transcriptase template (RT template), where the RT template encodes a modification to be incorporated into the target nucleic acid. In some embodiments, the CRISPR nucleic acid may be a type II or type V CRISPR nucleic acid, and / or the tracr nucleic acid may be any tracr corresponding to an appropriate type II or type V CRISPR nucleic acid. The extended guide nucleic acid may also be referred to as a target allele guide RNA (tag RNA). In some embodiments, the CRISPR nucleic acid useful in the present invention may be a type V CRISPR nucleic acid. In some embodiments, the tracr nucleic acid useful in the present invention may be a type V CRISPR tracr nucleic acid. In some embodiments, the CRISPR nucleic acid useful in the present invention may be a type II CRISPR nucleic acid. In some embodiments, a tracr nucleic acid useful in the present invention can be a type II CRISPR tracr nucleic acid. In some embodiments, the CRISPR nucleic acid and / or tracr nucleic acid can be, for example, Cas9, C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csx12), , Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG), and / or Csf5 systems.

[0168] In some embodiments, the extension portion of the extension guide may comprise an RT template and a primer binding site 5' to 3' (when the extension guide is linked to the 3' end of the CRISPR nucleic acid). In some embodiments, the extension portion of the extension guide may comprise a primer binding site and an RT template 5' to 3' (when the extension guide is linked to the 5' end of the CRISPR nucleic acid). In some embodiments, the RT template is between about 1 nucleotide and about 100 nucleotides in length (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 7, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides, and any range or value therein), for example, from about 1 nucleotide to about 10 nucleotides, from about 1 nucleotide to about 15 nucleotides, about 1 nucleotide to about 20 nucleotides, about 1 nucleotide to about 25 nucleotides, about 1 nucleotide to about 30 nucleotides, about 1 nucleotide to about 35, 36, 37, 38, 39, or 40 nucleotides, about 1 nucleotide to about 50 nucleotides, about 5 nucleotides to about 15 nucleotides, about 5 nucleotides to about 20 nucleotides, about 5 nucleotides to about 25 nucleotides, about 5 nucleotides to about 30 nucleotides, about 5 nucleotides to about 35, 36, 37, 38, 39, or 40 nucleotides, about 5 nucleotides to about 50 nucleotides, about 8 nucleotides to about 15 nucleotides, about 8 nucleotides to about 20 nucleotides, about 8 nucleotides to about 25 nucleotides, about 8 nucleotides to about 30 nucleotides, about 8 ...or 40 nucleotides, from about 8 nucleotides to about 50 nucleotides in length, from about 8 nucleotides to about 100 nucleotides, from about 10 nucleotides to about 15 nucleotides, from about 10 nucleotides to about 20 nucleotides, from about 10 nucleotides to about 25 nucleotides, from about 10 nucleotides to about 30 nucleotides, from about 10 nucleotides to about 36 nucleotides, from about 10 nucleotides to about 40 nucleotides, from about 10 nucleotides to about 50 nucleotides, from about 10 nucleotides to about 100 nucleotides in length, and any range or value therein. In some embodiments, the RT template length can be at least 8 nucleotides, optionally from about 8 nucleotides to about 100 nucleotides. In some embodiments, the RT template is 36, 37, 38, 39, or 40 nucleotides or less in length (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in length). or any value or range therein (e.g., from about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides in length to about 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in length).

[0169] As used herein, a "primer binding site" (PBS) of an extended portion of an extended guide nucleic acid (e.g., a tag RNA) refers to a contiguous nucleotide sequence that can bind to a region or "primer" on a target nucleic acid, i.e., that is complementary to a target nucleic acid primer. As an example, a CRISPR Cas effector protein (e.g., type II or type V, e.g., Cas9 or Cas12a) nicks / cuts DNA, and the 3' end of the cleaved DNA serves as a primer for the PBS portion of the extended guide nucleic acid. The PBS is designed to be complementary to the 3' end of one strand of the target nucleic acid and can be designed to bind to either the target strand or the non-target strand. The primer binding site can be fully complementary to the primer or can be substantially complementary to the primer on the target nucleic acid (e.g., at least 70% complementary (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or more)).In some embodiments, the length of the primer binding portion of the extension is between about 1 nucleotide and about 100 nucleotides in length (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, , 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides, or any value or range therein), about 3, 4, 5, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 , 20 nucleotides to about 50 nucleotides (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 nucleotides, or any range or value therein), or about 25 nucleotides to about 80 nucleotides. nucleotide (e.g., 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides in length, or any range or value therein).In some embodiments, the primer binding site is about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49 nucleotides to about 50, 51, 52, 53, 54, 55 , 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides in length, or any range or value therein. In some embodiments, the length of the primer binding site can be at least about 45, 46, 47, 48, 49, or 50 nucleotides, or more (e.g., about 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides in length, or any range or value therein).

[0170] In some embodiments, the extension portion of the extension guide may be fused to either the 5' or 3' end of a Type II or Type V CRISPR nucleic acid (e.g., 5' to 3' repeat-spacer-extension or extension-repeat-spacer) and / or the 5' or 3' end of a tracr nucleic acid. In some embodiments, when the extension portion is located 5' to the crRNA, the Type V CRISPR-Cas effector protein is modified to reduce (or eliminate) self-processing RNAse activity.

[0171] In some embodiments, the extension portion of the extended guide nucleic acid may be linked to the Type II or Type V CRISPR nucleic acid and / or Type II or Type V tracrRNA via a linker. In some embodiments, the linker is about 1 to about 100 nucleotides in length or more (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides in length, and any range therein (e.g., about 2 to about 40, about 2 to about 50, about 2 to about 60, about 4 to about 40, about 4 to about 50, about 4 to about 60, about 5 to about 40, about 5 to about 50, about 5 to about 60, about 9 to about 40, about 9 to about 50, about 9 to about 60, about 10 to about 40, about 10 to about 50, about 10 to about 60, about 40 to about 100, about 50 to about 100, or about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 nucleotides to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, It can be 6, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more nucleotides in length (e.g., about 105, 110, 115, 120, 130, 140 150 or more nucleotides in length).

[0172] As used herein, the terms "target nucleic acid," "target DNA," "target nucleotide sequence," "target region," or "target region in a genome" refer to a region of an organism's genome that is fully complementary (100% complementary) or substantially complementary (e.g., at least 70% complementary (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher)) to a spacer sequence in a guide RNA of the present invention (e.g., the spacer is substantially complementary to the target strand of the target nucleic acid). Useful target regions for CRISPR-Cas systems can be located immediately 3' (e.g., Type V CRISPR-Cas systems) or immediately 5' (e.g., Type II CRISPR-Cas systems) of the PAM sequence in the genome of an organism (e.g., a plant genome). The target region can be selected from any region of at least 15 contiguous nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides, etc.) located immediately adjacent to the PAM sequence on the target strand.

[0173] "Protospacer sequence" refers to a target double-stranded DNA, specifically a portion of a target nucleic acid / target DNA (e.g., or a target region in a genome (e.g., nuclear genome, plastid genome, mitochondrial genome), or an extragenomic sequence such as a plasmid, minichromosome, etc.) that is fully or substantially complementary to (and hybridizes with) the spacer sequence of a CRISPR repeat-spacer sequence (e.g., guide RNA, CRISPR array, crRNA). Thus, the protospacer sequence is complementary to the target strand of the target nucleic acid. In some embodiments, the target nucleic acid may have a first strand and a second strand (double-stranded DNA). In some embodiments, the term "first strand" as used herein with reference to a target nucleic acid may refer to the target strand or the bottom strand. In some embodiments, the term "second strand" as used with reference to a target nucleic acid is the strand complementary to the first strand (e.g., the top strand or non-target strand).

[0174] As understood in the art and used herein, "target strand" refers to the strand of double-stranded DNA to which the spacer is complementary and to which CRISPR-Cas effector protein is recruited, while "non-target strand" refers to the opposite strand of the target strand in double-stranded nucleic acid.In some embodiments of the present invention, the non-target strand of double-stranded nucleic acid, the opposite strand to which CRISPR-Cas effector protein is recruited, is nicked by CRISPR-Cas effector protein and edited by reverse transcriptase.In some embodiments, the target strand of double-stranded nucleic acid, the same strand to which CRISPR-Cas effector protein is recruited, is nicked by CRISPR-Cas effector protein and edited by reverse transcriptase.

[0175] In Type V CRISPR-Cas (e.g., Cas12a) and Type II CRISPR-Cas (Cas9) systems, the protospacer sequence is flanked (e.g., directly adjacent) by a protospacer adjacent motif (PAM). In Type IV CRISPR-Cas systems, the PAM is located at the 5' end on the non-target strand and at the 3' end of the target strand (see below for an example). JPEG0007785002000001.jpg36166

[0176] In Type II CRISPR-Cas systems (e.g., Cas9), the PAM is located immediately 3' of the target region. In Type I CRISPR-Cas systems, the PAM is located 5' of the target strand. Type III CRISPR-Cas systems have no known PAMs. Makarova et al. describe the nomenclature for all classes, types, and subtypes of CRISPR systems (Nature Reviews Microbiology, 13:722-736 (2015)). The guide structure and PAM are described by R. Barrangou (Genome Biol., 16:247 (2015)).

[0177] The canonical Cas12a PAM is T-rich. In some embodiments, the canonical Cas12a PAM sequence can be 5'-TTN, 5'-TTTN, or 5'-TTTV. In some embodiments, the canonical Cas9 (e.g., Streptococcus pyogenes) PAM can be 5'-NGG-3'. In some embodiments, non-canonical PAMs can be used, but may be less efficient.

[0178] Additional PAM sequences can be determined by experimental and computer methods established by those skilled in the art.For example, experimental methods include targeting sequences flanked by all possible nucleotide sequences, such as through the transformation of target plasmid DNA, and identifying sequence members that are not subjected to targeting (Esvelt et al., 2013, Nat. Methods, 10:1116-1121; Jiang et al., 2013, Nat. Biotechnol., 31:233-239).In some embodiments, computer methods can include performing a BLAST search of natural spacers to identify the original target DNA sequence in bacteriophage or plasmid, and aligning these sequences to determine the conserved sequences adjacent to the target sequence (Briner and Barrangou, 2014, Appl. Environ. Microbiol., 80:994-1001; Mojica et al., 2009, Microbiology, 155:733-740).

[0179] In some embodiments, the present invention further provides a method of modifying a target nucleic acid, comprising contacting the target nucleic acid at a first site with (a)(i) a first CRISPR-Cas effector protein and (ii) a first extended guide nucleic acid (e.g., a first extended CRISPR RNA, a first extended CRISPR DNA, a first extended crRNA, a first extended crDNA), and (b)(i) a second CRISPR-Cas effector protein, (ii) a first reverse transcriptase, and (ii) the first guide nucleic acid, thereby modifying the target nucleic acid. In some embodiments, the methods of the invention may further include contacting the target nucleic acid with (a) a third CRISPR-Cas effector protein and (b) a second guide nucleic acid, wherein the third CRISPR-Cas effector protein nicks a site on the first strand of the target nucleic acid that is about 10 to about 125 base pairs (either 5' or 3') from the second site on the second strand nicked by the second CRISPR-Cas effector protein, thereby improving mismatch repair. In some embodiments, the methods of the present invention may further comprise contacting the target nucleic acid with (a) a fourth CRISPR-Cas effector protein, (b) a second reverse transcriptase, and (c) a second extended guide nucleic acid (e.g., a second extended CRISPR RNA, a second extended CRISPR DNA, a second extended crRNA, a second extended crDNA) that targets (to which the spacer is substantially complementary / binds) a site on the first strand of the target nucleic acid, thereby modifying the target nucleic acid. The CRISPR-Cas effector proteins (e.g., first, second, third, fourth) useful in the present invention can be any combination of any Type I, Type II, Type III, Type IV, or Type V CRISPR-Cas effector proteins described herein.In some embodiments, the CRISPR-Cas effector protein is selected from the group consisting of Cas9, C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3'', Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csx12), It can be Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG), and / or Csf5.

[0180] In some embodiments, an extended guide nucleic acid useful with a first CRISPR-Cas effector protein may comprise (a) a CRISPR nucleic acid (CRISPR RNA, CRISPR DNA, crRNA, crDNA) and (b) an extended portion comprising a primer binding site and a reverse transcriptase template (RT template), where the RT template encodes a modification to be incorporated into the target nucleic acid.

[0181] In some embodiments, the CRISPR nucleic acid of the extended guide nucleic acid comprises a spacer sequence that is capable of binding (having substantial homology) to a first site on a first strand of a target nucleic acid.

[0182] In some embodiments, guide nucleic acids useful in CRISPR-Cas effector proteins comprise CRISPR nucleic acids (CRISPR RNA, CRISPR DNA, crRNA, crDNA). In some embodiments, the CRISPR nucleic acid of the first guide nucleic acid comprises a spacer sequence that binds to a second site on the first strand of the target nucleic acid upstream (3') of a first site on the first strand of the target nucleic acid.

[0183] In some embodiments, the second CRISPR-Cas effector protein can be a CRISPR-Cas fusion protein comprising a CRISPR-Cas effector protein domain fused to a reverse transcriptase.

[0184] In some embodiments, the second CRISPR-Cas effector protein may be a CRISPR-Cas fusion protein comprising a CRISPR-Cas effector protein domain fused to a peptide tag, and the reverse transcriptase may be a reverse transcriptase fusion protein comprising a reverse transcriptase domain fused to an affinity polypeptide capable of binding to the peptide tag.

[0185] In some embodiments, the first guide nucleic acid may be linked to an RNA recruitment motif and the reverse transcriptase may be a reverse transcriptase fusion protein comprising a reverse transcriptase domain fused to an affinity polypeptide capable of binding to the RNA recruitment motif.

[0186] In some embodiments, the target nucleic acid may further be contacted with a 5'-3' exonuclease, optionally fused to a first CRISPR-Cas effector protein. In some embodiments, the 5'-3' exonuclease may be a fusion protein comprising the 5'-3' exonuclease fused to a peptide tag, and the first CRISPR-Cas effector protein may be a fusion protein comprising a CRISPR-Cas effector protein domain fused to an affinity polypeptide capable of binding to the peptide tag. In some embodiments, the 5'-3' exonuclease may be a fusion protein comprising the 5'-3' exonuclease fused to an affinity polypeptide capable of binding to the peptide tag, and the first CRISPR-Cas effector protein may be a fusion protein comprising a CRISPR-Cas effector protein domain fused to the peptide tag. In some embodiments, the 5'-3' exonuclease can be a fusion protein comprising the 5'-3' exonuclease fused to an affinity polypeptide capable of binding to an RNA recruitment motif, and the extended guide nucleic acid is linked to the RNA recruitment motif.

[0187] In some embodiments, the methods of the invention may further comprise reducing double-strand breaks by introducing a chemical inhibitor of non-homologous end joining (NHEJ), by introducing a CRISPR guide nucleic acid or an siRNA targeting an NHEJ protein to transiently knock down expression of the NHEJ protein, or by introducing a polypeptide that prevents NHEJ (e.g., Gam protein).

[0188] In some embodiments, a complex is provided that includes: (a) a type II CRISPR-Cas effector protein or a type V CRISPR-Cas effector protein; (b) a reverse transcriptase; and (c) an extended guide nucleic acid (e.g., extended CRISPR RNA, extended CRISPR DNA, extended crRNA, extended crDNA, e.g., tag DNA, tag RNA).

[0189] In some embodiments, the type II or type V CRISPR-Cas effector protein of the complex may be a fusion protein comprising a type II or type V CRISPR-Cas effector protein domain fused to a peptide tag. In some embodiments, the type II or type V CRISPR-Cas effector protein of the complex may be a fusion protein comprising a type II or type V CRISPR-Cas effector protein domain fused to an affinity polypeptide capable of binding to the peptide tag. In some embodiments, the type II or type V CRISPR-Cas effector protein of the complex may be a fusion protein comprising a type II or type V CRISPR-Cas effector protein domain fused to an affinity polypeptide capable of binding to an RNA recruitment motif.

[0190] In some embodiments, the reverse transcriptase of the complex can be a fusion protein comprising a reverse transcriptase domain fused to a peptide tag. In some embodiments, the reverse transcriptase of the complex can be a fusion protein comprising an affinity polypeptide capable of binding to the reverse transcriptase domain fused to the peptide tag. In some embodiments, the reverse transcriptase of the complex can be a fusion protein comprising a reverse transcriptase domain fused to an affinity polypeptide capable of binding to an RNA mobilizing polypeptide. In some embodiments, the complex can further comprise a guide nucleic acid (e.g., extended CRISPR RNA, extended CRISPR DNA, extended crRNA, extended crDNA). In some embodiments, the complex can further comprise an extended guide nucleic acid (e.g., extended CRISPR RNA, extended CRISPR DNA, extended crRNA, extended crDNA).

[0191] In some embodiments, the complexes of the invention may be comprised in an expression cassette, which optionally is comprised in a vector.

[0192] The present invention further provides an expression cassette codon-optimized for expression in an organism, comprising, from 5' to 3': (a) a polynucleotide encoding a promoter sequence; (b) a polynucleotide encoding a Type V CRISPR-Cas nuclease (e.g., Cpf1 (Cas12a), dCas12a, etc.) or a Type II CRISPR-Cas nuclease (e.g., Cas9, dCas9, etc.), which has been codon-optimized for expression in the organism; (c) a linker sequence; and (d) a polynucleotide encoding a reverse transcriptase, which has been codon-optimized for expression in the organism, optionally wherein the organism is an animal such as a human, a plant, a fungus, an archaea, or a bacterium. Further provided is an expression cassette that is codon-optimized for expression in plants and includes, from 5' to 3', (a) a polynucleotide encoding a plant-specific promoter sequence (e.g., ZmUbi1, MtUb2, RNA polymerase II (Pol II)), (b) a plant-codon-optimized polynucleotide encoding a type V CRISPR-Cas nuclease (e.g., Cpf1 (Cas12a), dCas12a, etc.), (c) a linker sequence, and (d) a plant-codon-optimized polynucleotide encoding a reverse transcriptase. In some embodiments, the reverse transcriptase included in the expression cassette may be fused to one or more ssRNA-binding domains (RBDs). In some embodiments, the linker sequence may be an amino acid or peptide linker described herein.

[0193] The present invention further provides expression cassettes that are codon-optimized for expression in plants and include (a) a polynucleotide encoding a plant-specific promoter sequence (e.g., ZmUbi1, MtUb2), and (b) an extended RNA guide sequence that includes at its 3' end an extension that includes a primer binding site and an edit desired to be incorporated into the target nucleic acid (e.g., a reverse transcriptase template), optionally contained in the expression cassette and optionally operably linked to a Pol II promoter.

[0194] In some embodiments, plant-specific promoters useful in the expression cassettes of the present invention may be associated with an intron or are promoter regions that include an intron (e.g., ZmUbi1 including an intron, MtUb2 including an intron).

[0195] In some embodiments, the expression cassette may be codon-optimized for expression in dicotyledonous plants. In some embodiments, the expression cassette may be codon-optimized for expression in monocotyledonous plants.

[0196] In some embodiments, the invention provides methods for modifying a target nucleic acid in a plant or plant cell, comprising introducing one or more expression cassettes of the invention into the plant or plant cell, thereby modifying the target nucleic acid in the plant or plant cell, to produce a plant or plant cell comprising the modified target nucleic acid. In some embodiments, the methods of the invention further comprise regenerating a plant from the plant cell comprising the modified target nucleic acid, to produce a plant comprising the modified target nucleic acid. In some embodiments, the methods of the invention comprise contacting the target nucleic acid at a temperature between about 20°C and 42°C (e.g., about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, or 42°C, and any value or range therein.

[0197] In some embodiments, the invention provides cells comprising one or more polynucleotides, guide nucleic acids, nucleic acid constructs, expression cassettes, or vectors of the invention.

[0198] When used in combination with a guide nucleic acid, the polynucleotides / nucleic acid constructs / expression cassettes of the present invention can be used to modify a target nucleic acid. The target nucleic acid can be contacted with the polynucleotides / nucleic acid constructs / expression cassettes of the present invention before, simultaneously with, or after the step of contacting the target nucleic acid with the guide nucleic acid. In some embodiments, the polynucleotides and guide nucleic acids of the present invention can be contained in the same expression cassette or vector, and thus the target nucleic acid can be contacted with the polynucleotides and guide nucleic acid of the present invention simultaneously. In some embodiments, the polynucleotides and guide nucleic acids of the present invention can be in different expression cassettes or vectors, and thus the target nucleic acid can be contacted with the polynucleotides of the present invention before, simultaneously, or after contacting with the guide nucleic acid.

[0199] Polynucleotides of the invention can be used to modify (e.g., mutate, e.g., base edit, cleave, nick, etc.) target nucleic acids in any organism, including, but not limited to, plants, animals, bacteria, archaea, and / or fungi. Polynucleotides of the invention can be used to modify (e.g., mutate, e.g., base edit, cleave, nick, etc.) any animal or cell thereof, including, but not limited to, insects, fish, birds, amphibians, reptiles, and / or mammals. Exemplary mammals in which the invention may be useful include, but are not limited to, primates (humans and non-humans (e.g., chimpanzees, baboons, monkeys, gorillas, etc.)), cats, dogs, ferrets, gerbils, hamsters, cows, pigs, horses, goats, donkeys, or sheep.

[0200] The polynucleotides of the present invention can be used to modify (e.g., mutate, e.g., base edit, truncate, nick, etc.) any plant or plant part target nucleic acid. The nucleic acid constructs of the present invention can be used to modify any plant (or group of plants, e.g., into a genus or higher classification), including angiosperms, gymnosperms, monocotyledons, dicotyledons, C3, C4, CAM plants, bryophytes, ferns and / or fern allies, microalgae, and / or macroalgae. Plants and / or plant parts useful in the present invention can be plants and / or plant parts of any plant species / variety / cultivar. As used herein, the term "plant part" includes, but is not limited to, embryos, pollen, ovules, seeds, leaves, stems, shoots, flowers, branches, fruits, grains, ears, cobs, husks, stalks, roots, root tips, anthers, plant cells including intact plant cells in plants and / or plant parts, plant protoplasts, plant tissues, plant cell tissue cultures, plant calli, plant aggregates, etc. As used herein, "shoot" refers to the above-ground part including leaves and stems. Furthermore, as used herein, "plant cell" refers to the structural and physiological unit of a plant comprising a cell wall and may also refer to a protoplast. Plant cells can be in the form of isolated single cells, or can be cultured cells, or can be part of a higher unit, such as a plant tissue or plant organ.

[0201] Non-limiting examples of plants useful in the present invention include turfgrasses (e.g., bluegrass, bentgrass, ryegrass, fescue), feather reedgrass, broadleaf grass, torreya, arundo, switchgrass; artichoke, kohlrabi, arugula, leek, asparagus, lettuce (e.g., taro, leaf lettuce, romaine), malanga, melons (e.g., muskmelon, watermelon, Crenshaw, honeydew, cantaloupe), Brassica crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, collards, kale, Chinese cabbage, etc.), and other crops.cabbage), bok choy), cardoon, carrot, Chinese cabbage (napa), okra, onion, celery, parsley, chickpea, parsnip, chicory, pepper, potato, cucurbits (e.g., mallow, cucumber, zucchini, pumpkin, honeydew melon, watermelon, cantaloupe), radish, dry bulb onion, rutabaga, eggplant, bell pepper, esculenta, shallot, endive, garlic, spinach, leeks, pumpkin, leafy vegetables, beets (sugar beet and fodder vegetable crops, including beets, sweet potatoes, chard, horseradish, tomatoes, turnips, and spices; fruit crops, such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, quince, and figs; nuts (e.g., chestnuts, pecans, pistachios, hazelnuts, peanuts, walnuts, macadamia nuts, and almonds); citrus fruits (e.g., clementines, kumquats, oranges, grapefruits, tangerines, mandarins, lemons, and limes); etc.), blueberries, black raspberries, boysenberries, cranberries, currants, goulberries, loganberries, raspberries, strawberries, blackberries, grapes (for wine and for eating), avocados, bananas, kiwi, persimmons, pomegranates, pineapples, tropical fruits, pome fruits, melons, mangoes, papayas, and lychees, field crops such as clover, alfalfa, timothy grass, evening primrose, meadowfoam, corn (for animal feed, sweet corn, popcorn), hops , jojoba, buckwheat, safflower, quinoa, wheat, rice, barley, rye, millet, sorghum, oats, triticale, sorghum, tobacco, kapok, legumes (beans (e.g., fresh and dried), lentils, peas, soybeans), oil plants (rapeseed, canola, mustard, poppy, olive, sunflower, coconut, castor, cocoa beans, groundnut, oil palm), duckweed, Arabidopsis, fiber plants (cotton, flax, hemp, jute), cannabis (e.g., Cannabis sativa, Cannabis indica, and Cannabis ruderalis)ruderalis), camphor trees (cinnamon, camphor), or plants such as coffee, sugarcane, tea, and natural rubber plants, and / or bedding plants, e.g., flowering plants, cacti, succulents, and / or ornamental plants (e.g., roses, tulips, violets), trees, e.g., forest trees (broadleaf and evergreen trees, e.g., conifers, e.g., elm, ash, oak, maple, fir, spruce, cedar, pine, birch, cypress, eucalyptus, willow), and shrubs and other seedlings. In some embodiments, the nucleic acid constructs of the invention and / or expression cassettes and / or vectors encoding same may be used to modify corn, soybean, wheat, canola, rice, tomato, pepper, sunflower, raspberry, blackberry, black raspberry, and / or cherry.

[0202] The present invention further includes a kit or kits for carrying out the methods of the present invention. The kits of the present invention can include reagents, buffers, and equipment for mixing, measuring, sorting, labeling, etc., as well as instructions, etc. that may be suitable for modifying the target nucleic acid.

[0203] In some embodiments, the present invention provides kits comprising one or more nucleic acid constructs of the present invention and / or expression cassettes and / or vectors comprising same, along with optional instructions for their use. In some embodiments, the kits may further comprise a CRISPR-Cas guide nucleic acid (or extended guide nucleic acid) (corresponding to a CRISPR-Cas effector protein encoded by a polynucleotide of the present invention) and / or an expression cassette and / or vector comprising same. In some embodiments, the guide nucleic acid / extended guide nucleic acid may be provided on the same expression cassette and / or vector as one or more polynucleotides of the present invention. In some embodiments, the guide nucleic acid / extended guide nucleic acid may be provided on a separate expression cassette or vector from that comprising one or more of the polynucleotides of the present invention.

[0204] In some embodiments, the kit may further comprise a nucleic acid construct encoding the guide nucleic acid that includes a cloning site for cloning a nucleic acid sequence identical or complementary to the target nucleic acid sequence into the backbone of the guide nucleic acid.

[0205] In some embodiments, a nucleic acid construct of the invention may be an mRNA, which may encode one or more introns within the encoded polynucleotide. In some embodiments, expression cassettes and / or vectors comprising one or more polynucleotides of the invention may further encode one or more selectable markers useful for identifying transformants (e.g., nucleic acids encoding antibiotic resistance genes, herbicide resistance genes, etc.).

[0206] The present invention will now be described with reference to the following examples. It should be understood that these examples are not intended to limit the scope of the claimed invention, but rather are intended to be illustrative of particular embodiments. Any variations of the exemplified methods that occur to those skilled in the art are intended to be within the scope of the present invention. [Example]

[0207] Allelic RNA-encoded DNA replacement (REDRAW) utilizes a V-type Cas effector, an enzyme polymerized from a DNA:RNA hybrid from a free DNA 3' end (annealing site, AS) and an extended guide nucleic acid (i.e., the target allele guide RNA (tag RNA)). These three macromolecules work together to: i) target the CRISPR enzyme to the desired genomic site using the CRISPR effector and the crRNA portion of the tag RNA; ii) nick or cleave the DNA to generate a free 3' end; iii) provide a portion of the tag RNA that anneals to the free 3' end of the DNA; iv) provide a portion of the tag RNA that provides a template for an RNA-dependent DNA polymerase; and v) allow for the termination of reverse transcription by either enzyme collision, spontaneous termination, or encountering a stable hairpin.

[0208] We tested the REDRAW system using a non-target strand (NTS) nickase version of LbCas12a_R1138A and a RT derived from Moloney murine leukemia virus (M-MuLV). LbCas12a_R1138A was predicted to be an NTS nickase based on its alignment with the previously described AsCas12a_R1226A mutant. We demonstrated in Figure XXX that LbCas12a_R1138A is indeed a nickase. The LbCas12a used was either RNAse (+) or had a mutation that prevented RNAse activity (H759A). The LbCas12a_R1138A_H759A mutant was used to prevent self-processing of the tag RNA when generating the 5' extension or incorporating the 3' hairpin.

[0209] The tag RNAs tested included crRNAs containing either 5' or 3' extensions. Various annealing site lengths were tested to allow shorter or longer DNA:RNA hybrids to form from the nicked non-target strand. Various lengths of RNA templates were also tested. Finally, two different hairpins were also incorporated into the naturally occurring LbCas12a pseudoknot hairpin design and the decoy pseudoknot hairpin design. [Example]

[0210] LbCas12a_R1138A nickase assay A nucleic acid construct containing LbCas12a, followed by the nucleoplasmin NLS and a 6x histidine tag, was synthesized (GeneWiz) (SEQ ID NO: 57) and cloned between NcoI and XhoI in the pET28a vector to generate pWISE450 (SEQ ID NO: 58). To facilitate cloning, an additional glycine was added to the sequence between Met-1 and Ser-2. The numbering presented herein excludes this extra glycine. The R1138A mutation was then generated using the QuickChange II Site-Directed Mutagenesis Kit (Agilent) according to the manufacturer's instructions. These expression plasmids were then transformed into BL21(DE3) Star-competent E. coli cells (ThermoFisher Scientific).

[0211] BL21(DE3)Star cells were grown in Luria broth and 50 μg / ml kanamycin at 37°C until an optical density of A600 = 0.5 was achieved. Isopropyl β-d-1-thiogalactopyranoside (IPTG) was added to 0.5 mM, and protein was induced overnight at 18°C. Cells were pelleted at 5,000 × g. Purification was achieved using two columns according to the manufacturer's protocol: a HisTrap column followed by a MonoS column (GE HELTHCARE).

[0212] CRISPR RNA (crRNA) was sequenced by Synthego. The sequence was synthesized using JPEG0007785002000002.jpg11166 (guidelines are in bold font in the sequence).

[0213] The plasmid to be cut is The sequence of JPEG0007785002000003.jpg6166 was inserted into pUC19. The portion of the sequence in bold is the PAM sequence recognized by LbCas12a, and the remainder (in normal font) is the protospacer sequence. The pUC19 plasmid was transformed into XL1-Blue (Agilent) (E. coli) and then purified using a Qiagen Plasmid Spin Mini Kit.

[0214] Nuclease assays were performed by mixing a 10:10:1 ratio of LbCas12a_R1138:crRNA:plasmid, incubating at 37°C for 15 minutes in New England Biolabs buffer 2.1, heat-inactivating at 80°C for 20 minutes, and loading onto a 1% TAE-agarose gel with embedded SYBR-Safe stain (Invitrogen) to stain the DNA. As shown in Figure 4, in vitro assays, LbCas12a_R1138A is a nickase. As shown in lanes 2 and 3, a supercoiled 2.8 kB plasmid ran at an apparent size of 2.0 kB (lane 2) until a double-strand break was generated by wild-type LbCas12a (lane 3). The mutant enzyme LbCas12a_R1138A primarily produced a nicked product that ran at an apparent size of 5.0 kB. Lanes 4-6 show that increasing concentrations of mutant enzyme did not alter the ratio until general nuclease digestion of the plasmid occurred using very high concentrations of enzyme (256 nM).

[0215] Design and construction of REDRAW editor plasmids - bacterial screening The REDRAW (replacement of allelic RNA-encoded DNA) expression construct was synthesized by solid-phase synthesis and cloned into the expression vector pET28a(+) between the NcoI and XhoI restriction sites. The REDRAW expression vector contains a ColE1 origin of replication, a kanamycin resistance marker, and the REDRAW editor under the control of the T7 promoter and terminator. The REDRAW editor contains either a Cas12a nickase (R1138A) or an Rnase-dead Cas12a nickase (R1138A, H759A) fused to the Mu-LV reverse transcriptase MuLV(5M) (see, e.g., SEQ ID NO: 97) (a murine leukemia virus reverse transcriptase with five mutations: D200N+L603W+T330P+T306K+W313F) (Anzalone et al., Nature, 576(.7785):149-157 (2019)) using an XTEN or 5R linker. All REDRAW editor sequences were E. coli codon-optimized. The configurations of the REDRAW editors tested are shown in Figure 5. Two configurations provided in Figure 5 had Cas12a N-terminal to the reverse transcriptase, and two configurations had Cas12a C-terminal to the reverse transcriptase. The configuration tested was constructed using a Cas12a mutant with an additional H759A mutation to prevent processing of tag RNAs containing 5' extensions.

[0216] Design and construction of tag RNA plasmids - bacterial screening The sequences of the tag RNA (target allele guide RNA) library were designed using an algorithm that assembled Cas12a spacer and scaffold sequences together with a reverse transcriptase template and a primer binding site unique to each target. The design parameters, shown in Table 1, span a wide range of primer binding sites and reverse transcriptase template lengths. The desired changes, shown in Table 3, were designed to confer resistance to antibiotics after successful editing.

[0217] [Table 1]

[0218] Figure 6 shows the configuration of tag RNAs in the first library. Both 5' and 3' extensions containing RTT and PBS were included in the library.

[0219] The second library was designed similarly to the first, but further evaluated whether the presence of a hairpin located immediately 3' to the spacer in the 3' tag RNA extension configuration would improve REDRAW editing. The design parameters, shown in Table 2, again examined a wide range of primer binding site (PBS) and reverse transcriptase template (RTT) lengths, but focused on the RTT length region found to be functional from the first library. Both 5' and 3' extensions containing RTT and PBS were included in the library. In addition, variants containing decoy hairpins were also included in the second tag RNA library. A hairpin similar to the native LbCas12a scaffold sequence but not recognized and cleaved by the Cas12a protein was desired. Therefore, an existing hairpin with a structure similar to the LbCas12a hairpin was found in the HIV-1 RNA genome and modified by adding a UA sequence to form a pseudoknot, as shown in Figure 7.

[0220] [Table 2]

[0221] Construction of tagged RNA plasmids for bacterial screening The base plasmid for the tag RNA library was generated by solid-phase synthesis and cloning the holder fragment into pTwist Amp Medium Copy (TWIST BIOSCIENCE®). The plasmid contains a p15A origin of replication and an ampicillin resistance marker. The tag RNA is constitutively expressed from a synthetic BbaJ23119 promoter and terminated by a T7 terminator. The first tag RNA library evaluated was synthesized by an external vendor (Genewiz) and cloned into the tag RNA base vector. For the second library, oligos were synthesized using the NEB HiFi assembly kit according to the manufacturer's instructions and then cloned into the tag RNA base vector. To ensure that a wide range of PBSs, RTTs, and targets were represented in the library and that substantial bias was not present, the diversity of the library was investigated by colony PCR and Sanger sequencing of 72 clones from the library.

[0222] Design and construction of reporter plasmids A basic reporter plasmid containing the CloDF13 origin of replication, a chloramphenicol resistance marker, and a spectinomycin resistance marker (aadA) was constructed by PCR amplification of the CloDF13 origin of replication and the chloramphenicol resistance marker and ligating it with the PCR-amplified aadA resistance marker. Three reporter plasmids containing aadA mutants were then constructed by excising the wild-type aadA gene between the BamHI and BglII restriction sites and ligating synthetic gene blocks containing stop codons at residue positions Thr61, Leu115, or Asp132. All reporter plasmids were verified by Sanger sequencing after construction. Furthermore, reporter plasmids containing aadA mutants with a stop codon in the coding sequence were confirmed as sensitive to both spectinomycin and streptomycin before their use in REDRAW tag RNA screening experiments.

[0223] Targeted bacterial screening for REDRAW editing Five targets were tested in REDRAW editing experiments and are shown below in Table 3. Two genomic and three plasmid targets were used in all cases. Successful REDRAW editing of any target confers resistance to antibiotics (nalidixic acid or streptomycin), linking survival of the host organism (E. coli) to successful REDRAW editing.

[0224] [Table 3]

[0225] REDRAW tag RNA experiment - bacterial screening The host organism for all bacterial REDRAW tag RNA screening experiments was E. coli BL21(DE3). Prior to selection experiments, each REDRAW expression construct was transformed into chemically competent BL21(DE3) according to the manufacturer's instructions and plated onto LB agar plates containing kanamycin. Single colonies were then picked from the transformation plates, and batches of electrocompetent cells were generated according to a previously developed method (Sambrook and Russell (Transformation of E. coli by electroporation., Cold Spring Harbor Protocols, 2006, 1(2006):pdb-prot3933). Competent cells harboring each REDRAW expression construct were then electroporated with 10 ng of each reporter plasmid, recovered for 1 hour in SOC at 37°C and 225 rpm, and plated onto LB agar plates containing kanamycin and chloramphenicol. Single colonies from these plates were then picked from the transformation plates, and batches of electrocompetent cells were generated again (Sambrook and Russell (Transformation of E. coli by electroporation., Cold Spring Harbor Protocols, 2006, 1(2006):pdb-prot3933). Protocols, 2006, 1(2006):pdb-prot3933). Table 4 below summarizes the batches of electrocompetent cells that were generated for testing of the first tag RNA library.

[0226] [Table 4]

[0227] Selection experiments were performed by first electroporating 100 ng of tag RNA library into 50 μL of each batch of electrocompetent cells. Transformations were allowed to recover for 1 hour at 37°C with shaking at 225 rpm. After 1 hour of recovery, 1 μL of the recovery was removed, mixed with 99 μL of LB, and plated onto LB agar plates containing the appropriate antibiotic to examine transformation efficiency. The remaining volume of each transformation was then added to 29 mL of LB plus antibiotics (LB Kan / Carb for genomic selection and LB Kan / Carb / Cam for plasmid selection) and 0.5 mM IPTG. Expression cultures were grown overnight at 37°C with shaking at 225 rpm.

[0228] The next day, the OD600 of each expression culture was measured. From each expression culture, 1 OD was inoculated onto five plates (approximately 0.2 OD / plate) containing antibiotics for the REDRAW expression vector (Kan), antibiotics for the tag RNA plasmid (Carb), reporter plasmid, 0.5 mM IPTG, and additional selection antibiotics (nalidixic acid or streptomycin). Plates were incubated overnight at 37°C and observed for growth the following morning. If no colonies were observed, plates were incubated for an additional 24 hours at 37°C.

[0229] Colonies observed on the selective plates were picked and restreaked onto plates with the appropriate antibiotic, then subjected to colony PCR to amplify tag RNA for gene targeting and Sanger sequencing. Sanger sequencing was performed on the colony PCR products by Genewiz.

[0230] The evaluation of the second library was performed in the same manner as the first tag RNA library, with one modification. Instead of preparing batches of 20 electrocompetent cells, one large batch of electrocompetent BL21(DE3) cells harboring the second tag RNA library was prepared. Then, the REDRAW expression construct (100 ng) or the REDRAW expression construct plus reporter plasmid (100 ng each) was transformed into the electrocompetent cells harboring the tag RNA library. All subsequent steps were repeated in the same manner.

[0231] Evaluation of REDRAW editing using the first tagged RNA library - bacterial screening The number of colonies obtained from the first tag RNA library selection experiment is summarized in Table 5 below. No colonies were observed in any of the genome selections (selections 1 to 8). Colonies were observed for each of the plasmid selections.

[0232] [Table 5] JPEG0007785002000009.jpg117166

[0233] In selections 9, 12, 15, and 18 (aadA Thr61 target), bacterial lawns were observed. Isolated colonies from these plates were false positives. In selections 10, 11, 13, 14, 16, and 17 (aadA Leu115 target and aadA Asp132 target), a small number of colonies were observed on the plates. The colonies on these plates contained both tag RNA and the target amplified by colony PCR and were subjected to Sanger sequencing to confirm the editing and identify the tag RNA responsible for the editing. All colonies evaluated from selections 11, 14, 17, and 20 (aadA Asp132 target) were false positives. Multiple colonies from selection 10 (aadA Leu115 target) contained the designed editing and the associated tag RNA. Sequencing of the edited target is shown in Figure 8 and demonstrates a TGA->CTG edit in the non-functional aadA gene that restores antibiotic resistance.

[0234] The identified tag RNA sequences responsible for editing are associated with the editing shown in Figure 8: 5'-GTTTCAAAGATTAAATAATTTCTACTAAGTGTAGATTACGGCTCCGCAGTGGATGGCGGTAA TTTCTACTAAGTGTAGATGCGGCGCGTTGTTTCATCAAGGCGTACGGTCACCGTAACCAGCAAATCAATATCACTGTGTGGCTTCAGGCCGCCATCCACTGCGG-3' (SEQ ID NO: 87).

[0235] The protein configuration from selection 10 is as follows: SV40-nCas12a-XTEN-MMLV-RT-SV40.

[0236] Evaluation of REDRAW editing using a second tagged RNA library - genome selection results The numbers of colonies obtained from the genomic selection experiments of the second tag RNA library are summarized below in Table 6. Colonies were observed on rpsL selection plates.

[0237] [Table 6]

[0238] No colonies were observed on plates from selections 2.1–2.4 and 2.9–2.12 (gyrA genome target). A small number of colonies were observed on plates from selections 2.5–2.8 and 2.13–2.16 (rpsL genome target). Colonies on these plates were restreaked to confirm resistance to all antibiotics. Colonies from these plates were then used to generate PCR products of tag RNA and target them for Sanger sequencing. Sanger sequencing was used to confirm the editing and identify the tag RNA responsible for the editing. All colonies from selections 2.6–2.8 and 2.13–2.16 were false positives. One colony from selection 2.5 contained the engineered edit AAA to CGT, which confers streptomycin resistance (see Figure 9).

[0239] The identified tag RNA sequences associated with the editing shown in Figure 9 are as follows: 5'-TATTTCTATAAGTGTAGATTACTCGTGTATATACTCCGCACCGAGGTTGGTACGAACACCGGGAGTCTTTAACACGACCGCCACGGATCAGGATCACGGAGTGCTCCTGCAGGTTGT GACCTTCACCACCGATGTAGGAAGTCACTTCGAAACCGTTAGTCAGACGAACACGGCATACTTTACGCAGCGCGGAGTTCGGTTTACGAGGAGTGGTAGTATATACACGAGT-3' SEQ ID NO: 92.

[0240] The protein configuration from selection 2.5 is as follows: SV40-MMLV-RT-XTEN-nRVRLbCas12a(H759A)-SV40.

[0241] Evaluation of REDRAW editing using a second tagged RNA library - Plasmid selection results The numbers of colonies obtained from the plasmid selection experiments of the second tag RNA library are summarized in Table 7 below.

[0242] [Table 7] JPEG0007785002000012.jpg182166

[0243] Colonies were observed on plates for Leu115 and Asp132 selections. Selections 2.18, 2.19, 2.22, 2.25, 2.28, 2.31, 2.39, and 2.40 had colonies on the selection plates. These colonies were restreaked to confirm resistance to all antibiotics. They were then used to generate PCR products of tag RNA and targets for Sanger sequencing. Sanger sequencing was used to confirm the editing and identify the tag RNA responsible for the editing. All colonies from selections 2.18, 2.19, 2.22, 2.28, 2.39, and 2.40 were false positives. Four colonies from selection 2.25 and two colonies from selection 2.31 had the designed edits and associated tag RNAs shown in Figures 10 and 11. Four colonies from selection 2.25 had identical edits and tag RNAs. Two colonies from selection 2.31 also had identical edited and tagged RNAs.

[0244] The identified tag RNA sequences associated with the edits in Figure 10 from selection 2.25 are as follows: 5'-TAATTTCTACTAAGTGTAGATTACGGCTCCGCAGTGGATGGCGGTAAGTCTCCATAGAATGGAGGACAGCGCGGAGAATCTCGCTCTCTCCAGGGGAAGCCGAAGTTTCCAAAAGGTCGTTGATCAAAGCGCGGCGCGTTGTTTCATCAAGGCGTACGGTCACCGTAACCAGCAAATCAATATCACTGTGTGGCTTCAGGCCGCCATCCACTGCGGAT-3' SEQ ID NO: 93.

[0245] The protein configuration from selection 2.25 is as follows: SV40-nCas12a-XTEN-MMLV-RT-SV40.

[0246] The identified tag RNA sequences associated with the edits in Figure 11 from selection 2.31 are as follows: 5'-TAATTTCAACTAAGTGTAGATTACGGCTCCGCAGTGGATGGCGGTAAGTCTCCATAGAATGGAGGGCGGAGAATCTCGCTCTCCAGGGGAAGCCGAAGTTTCCAAAAGGTCGTTGATCAAAGCGCGGCGCGTTGTTTCATCAAGGCGTACGGTCACCGTAACCAGCAAATCAATATCACTGTGTGGCTTCAGGCCGCCATCCACTGCGGAT-3' SEQ ID NO:94.

[0247] The protein configuration from selection 2.31 is as follows: SV40-MMLV-RT-XTEN-nLbCas12a(H759)-SV40.

[0248] Summary of REDRAW editing observed in bacterial cells Table 8 below provides a summary of observed cases of REDRAW editing in E. coli. For each example, the protein configuration (REDRAW editor), edited target, location of tag RNA extension (5' or 3' to the Cas12a hairpin and guide), length of PBS, and length of RTT are listed.

[0249] [Table 8] [Example]

[0250] Accurate editing activity in human cells A further approach using an activated form of Cas12a in conjunction with reverse transcriptase is shown in Figure 12 and outlined below. Nuclease-active Cas12a is recruited to the site via spacer-target site interactions. Cas12a makes a double-stranded break and, optionally, degrades the non-template strand provided a 5' to 3' exonuclease. Priming occurs using the tag RNA. The primer binding site (PBS) encodes the sequence to the right of the cleavage site that is complementary to the template strand DNA. · Reverse transcriptase (MMuLV-RT(5M)) extends from the primed site or primer on the target nucleus (dashed line = extension) and encodes the desired change in the newly synthesized strand. Creation of a new edited DNA strand occurs through the division of the DNA intermediate via mismatch repair and DNA ligation.

[0251] method Extended guide RNAs were designed to target two genomic sites, DMNT1 and FANCF1, in HEK293T cells. Various combinations of primer binding site (PBS) and reverse transcriptase template (RTT) lengths were assayed. The guide RNAs encoded two-base changes in the PAM region of the target guide, corresponding to TT to AA at positions -2 and -3 (counting the TTTV PAM as positions -4 to -1). The guide extensions were fused to either the 5' or 3' end of the guide RNA.

[0252] Plasmids encoding RNAse-dead mutant LbCas12a(H758A), reverse transcriptase (MMuLV-RT(5M)), and optionally exonuclease (one of T5 exonuclease, T7 exonuclease, RecE, and RecJ), as well as extended guide RNAs, were transfected into HEK293T cells grown to 70% confluence using Lipofectamine™ 3000 according to the manufacturer's protocol. Cells were harvested after 3 days, and gene editing was quantified by next-generation sequencing.

[0253] result We observed precise editing of both targeted sites. Depending on the guide design, we observed up to 0.5% editing at the FANCF1 site (Figure 13) and up to 1.7% editing at the DMNT1 site (Figure 14). The use of exonucleases improved editing efficiency in some guide designs.

[0254] [Table 9]

[0255] [Table 10]

[0256] The effect of exonuclease transfection on precise editing activity at DMNT1 sites is shown in Figure 15 (normalized to no exonuclease treatment, pUC19=1). Exonucleases improve editing in some guide configurations.

[0257] The foregoing is illustrative of the present invention, and is not to be construed as limiting thereof. The present invention is defined by the following claims, with equivalents of the claims to be included therein.

Claims

1. 1. A method for modifying a double-stranded target nucleic acid, comprising: The target nucleic acid (a) a type V CRISPR-Cas effector protein or a type II CRISPR-Cas effector protein; (b) a reverse transcriptase, and (c) contacting with an extended guide RNA; The extended guide RNA (i) a type V CRISPR nucleic acid or a type II CRISPR nucleic acid; and (ii) an extension portion comprising a primer binding site and a reverse transcriptase template (RT template), wherein the primer binding site and the reverse transcriptase template are linked to the 5' end or 3' end of the CRISPR nucleic acid in the order of reverse transcriptase template, primer binding site; the double-stranded target nucleic acid comprises a first strand and a second strand, the Type V CRISPR-Cas effector protein or the Type II CRISPR-Cas effector protein is a double-stranded nuclease that cleaves the first strand and the second strand of the target nucleic acid to create a double-stranded break, and the primer binding site binds to the first strand, which is the same strand to which the CRISPR-Cas effector protein is recruited, thereby modifying the target nucleic acid. method.

2. The type V CRISPR-Cas effector protein or the type II CRISPR-Cas effector protein, the reverse transcriptase, and the extended guide RNA form a complex or are contained in a complex. The method of claim 1.

3. 3. The method of claim 1 or 2, i) the extended portion of the extended guide RNA is linked to a CRISPR nucleic acid via a linker; ii) the Type V CRISPR-Cas effector protein is modified to reduce or eliminate self-processing RNAse activity; or iii) a combination of i) and ii); method.

4. The method of claim 1 or claim 2, the primer binding site is between 4 and 100 nucleotides in length and / or the RT template is between 7 and 100 nucleotides in length; method.

5. the type V CRISPR-Cas effector protein or the type II CRISPR-Cas effector protein is a fusion protein, and / or the reverse transcriptase is a fusion protein, and the type V CRISPR-Cas fusion protein or the type II CRISPR-Cas effector protein, the reverse transcriptase fusion protein, and / or the extended guide RNA are fused to one or more components that recruit the reverse transcriptase to the type V CRISPR-Cas effector protein or the type II CRISPR-Cas effector protein; The method according to any one of claims 1 to 4.

6. The method of claim 5, wherein the one or more components are recruited via protein-protein interactions, protein-RNA interactions, and / or chemical interactions.

7. the V-type CRISPR-Cas effector protein is a V-type CRISPR-Cas effector fusion protein comprising a V-type CRISPR-Cas effector protein domain fused (linked) with a peptide tag, and the reverse transcriptase is a reverse transcriptase fusion protein comprising a reverse transcriptase domain fused (linked) with an affinity polypeptide that binds to the peptide tag; The method according to any one of claims 1 to 6.

8. the type II CRISPR-Cas effector protein is a type II CRISPR-Cas effector fusion protein comprising a type II CRISPR-Cas effector protein domain fused (linked) to a peptide tag, and the reverse transcriptase is a reverse transcriptase fusion protein comprising a reverse transcriptase domain fused (linked) to an affinity polypeptide that binds to the peptide tag; The method according to any one of claims 1 to 7.

9. the extended guide RNA is linked to an RNA recruitment motif, and the reverse transcriptase is a reverse transcriptase fusion protein comprising a reverse transcriptase domain fused (linked) to an affinity polypeptide that binds to the RNA recruitment motif. The method according to any one of claims 1 to 8.

10. the extended guide RNA is linked to two or more RNA recruitment motifs; 10. The method of claim 9.

11. The method of claim 10, wherein the two or more RNA recruitment motifs are the same RNA recruitment motif or different RNA recruitment motifs.

12. contacting the target nucleic acid with two or more reverse transcriptase fusion proteins; The method according to any one of claims 1 to 11.

13. 10. The method of claim 9, i) the recruitment motif is located at the 3' end of or embedded within the extension portion of the extended guide RNA; ii) the recruitment motif and corresponding affinity polypeptide are a telomerase Ku-binding motif and affinity polypeptide of Ku, a telomerase Sm7-binding motif and affinity polypeptide of Sm7, an MS2 phage operator stem-loop and affinity polypeptide MS2 coat protein (MCP), a PP7 phage operator stem-loop and affinity polypeptide PP7 coat protein (PCP), an SfMu phage Com stem-loop and affinity polypeptide Com RNA-binding protein, a PUF binding site (PBS) and affinity polypeptide Pumilio / fem-3 mRNA-binding factor (PUF), and / or a synthetic RNA-aptamer and a corresponding aptamer ligand; iii) the recruitment motif and corresponding affinity polypeptide are an MS2 phage operator stem-loop and the affinity polypeptide MS2 coat protein (MCP), and / or a PUF binding site (PBS) and the affinity polypeptide Pumilio / fem-3 mRNA binding factor (PUF); or iv) Any combination of i) to iii) above; method.

14. The target nucleic acid (a) a CRISPR-Cas effector protein, and (b) guide nucleic acid 14. The method of any one of claims 1 to 13, further comprising contacting a CRISPR-Cas effector protein with a CRISPR-Cas effector protein, wherein (i) the CRISPR-Cas effector protein nicks or cuts a site on the first strand of the target nucleic acid located 10 to 125 base pairs (either 5' or 3') from the site on the second strand nicked by the type II or type V CRISPR-Cas effector protein, or (ii) the CRISPR-Cas effector protein nicks or cuts a site on the second strand of the target nucleic acid located 10 to 125 base pairs (either 5' or 3') from the site on the first strand nicked by the type II or type V CRISPR-Cas effector protein, thereby improving mismatch repair, and wherein the CRISPR-Cas effector protein is a type I, type II, type III, type IV, or type V CRISPR-Cas effector protein.

15. The method according to any one of claims 1 to 14, further comprising contacting the target nucleic acid with a Dna2 polypeptide and / or a 5' flap endonuclease (FEN); method.

16. i) the FEN is a FEN1 polypeptide; ii) said FEN and / or Dna2 polypeptide is overexpressed in the presence of said target nucleic acid; iii) the FEN is a fusion protein comprising a FEN domain fused to a type II or type V CRISPR-Cas effector protein or domain, and / or the Dna2 polypeptide is a fusion protein comprising a Dna2 domain fused to a type II or type V CRISPR-Cas effector protein or domain; or iv) Any combination of i) to iii) above; 16. The method of claim 15.

17. 16. The method of claim 15, i) the V-type CRISPR-Cas effector protein is a V-type CRISPR-Cas fusion protein comprising a V-type CRISPR-Cas effector protein domain fused (linked) to a peptide tag, the FEN is a FEN fusion protein comprising a FEN domain fused to an affinity polypeptide that binds to the peptide tag, and / or the V-type CRISPR-Cas effector protein is a V-type CRISPR-Cas fusion protein comprising a V-type CRISPR-Cas effector protein domain fused to a peptide tag, and the Dna2 polypeptide is a Dna2 fusion protein comprising a Dna2 domain fused to an affinity polypeptide that binds to the peptide tag; or ii) the type II CRISPR-Cas effector protein is a type II CRISPR-Cas fusion protein comprising a type II CRISPR-Cas effector protein domain fused (linked) to a peptide tag, the FEN is a FEN fusion protein comprising a FEN domain fused to an affinity polypeptide that binds to the peptide tag, and / or the type II CRISPR-Cas effector protein is a type II CRISPR-Cas fusion protein comprising a type II CRISPR-Cas effector protein domain fused to a peptide tag, and the Dna2 polypeptide is a Dna2 fusion protein comprising a Dna2 domain fused to an affinity polypeptide that binds to the peptide tag; method.

18. contacting the target nucleic acid with two or more FEN fusion proteins and / or two or more Dna2 fusion proteins, thereby recruiting the FENs and / or Dna2 to the Type V CRISPR-Cas effector protein domain and the target nucleic acid; or contacting the target nucleic acid with two or more FEN fusion proteins and / or two or more Dna2 fusion proteins, thereby recruiting the FENs and / or Dna2 to the Type II CRISPR-Cas effector protein domain and the target nucleic acid; 18. The method of claim 17.

19. (a) a type V or type II CRISPR-Cas effector protein, each of which cleaves a first strand and a second strand of a target nucleic acid to create a double-strand break; (b) a reverse transcriptase; and (c) an extended guide RNA; A complex comprising: the extended guide RNA (i) a type V CRISPR nucleic acid or a type II CRISPR nucleic acid; and (ii) an extension portion comprising a primer binding site and a reverse transcriptase template (RT template), wherein the primer binding site and the reverse transcriptase template are linked to the 5' end or 3' end of the CRISPR nucleic acid in the order of reverse transcriptase template, primer binding site; the primer binding site is designed to bind to the first strand, which is the same strand to which the CRISPR-Cas effector protein is recruited; Complex.

20. 20. One or more expression cassettes encoding the type V or type II CRISPR-Cas effector protein, the reverse transcriptase, and the extended guide RNA, designed to express the complex of claim 19 in an organism, wherein the expression cassettes are codon-optimized for expression in the organism. Expression cassette.

21. 21. The expression cassette of claim 20, which is codon-optimized for expression in plants.

22. 22. A method for modifying a target nucleic acid in a plant or plant cell, comprising the step of introducing the expression cassette of claim 20 or 21 into the plant or plant cell, thereby modifying the target nucleic acid in the plant or plant cell.

Citation Information

Patent Citations

  • Edited Methods and compositions for editing nucleotide sequences

    JP2022526908A