Recruiting DNA polymerase for template-based editing

A complex of sequence-specific DNA-binding proteins and DNA-dependent DNA polymerases enhances template-directed editing in plants by introducing breaks and incorporating repair templates, addressing low editing efficiencies in eukaryotic cells.

JP7765390B2Active Publication Date: 2025-11-06PAIRWISE PLANTS SERVICES INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022541803
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-06
Filing Date
2021-01-06
Publication Date
2025-11-06
Estimated Expiration
2041-01-06

AI Technical Summary

Technical Problem

The efficiency of template-directed editing in eukaryotic cells, particularly in plants, is low due to difficulties in delivering multiple reagents and ensuring sufficient availability of repair templates, leading to low editing efficiencies of less than 10% in most cases.

Method used

A complex comprising a sequence-specific DNA-binding protein, a DNA-dependent DNA polymerase, and optionally a DNA endonuclease, which can introduce single-strand or double-strand breaks and incorporate repair templates into target nucleic acids, enhancing editing efficiency by facilitating template-directed integration.

Benefits of technology

The complex significantly improves editing efficiency in plant cells, potentially achieving editing rates comparable to human cell cultures, overcoming delivery and integration challenges in the homologous recombination pathway.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007765390000001
    Figure 0007765390000001
  • Figure 0007765390000002
    Figure 0007765390000002
  • Figure 0007765390000003
    Figure 0007765390000003
Patent Text Reader

Abstract

The present invention relates to recombinant nucleic acid constructs comprising a sequence-specific DNA-binding protein, a DNA-dependent DNA polymerase, and a DNA-encoded repair template, and optionally a DNA endonuclease, wherein the sequence-specific DNA-binding protein comprises DNA endonuclease activity, and methods of use thereof for modifying nucleic acids in cells and organisms.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] STATEMENT REGARDING THE ELECTRONIC SEQUENCE LISTING FILE In lieu of a paper copy, a Sequence Listing in ASCII text format submitted via EFS-Web, filed in accordance with 37 C.FR § 1.821, entitled 1499.14.WO_ST25.txt, 435,310 bytes in size, created on January 6, 2021, is provided. This Sequence Listing is incorporated herein by reference for its disclosure.

[0002] Priority statement This application claims the benefit, pursuant to 35 U.S.C. § 119(e), of U.S. Provisional Patent Application No. 62 / 957,542, filed January 6, 2019, the entire contents of which are incorporated herein by reference.

[0003] FIELD OF THE INVENTION The present invention relates to recombinant nucleic acid constructs comprising a sequence-specific DNA-binding protein, a DNA-dependent DNA polymerase, and a DNA-encoded repair template, and optionally a DNA endonuclease, wherein the sequence-specific DNA-binding protein comprises DNA endonuclease activity, and methods of use thereof for modifying nucleic acids in cells and organisms. [Background technology]

[0004] Precise, template-directed editing typically involves introducing a double-strand break (DSB) into the target site and providing a template bearing the desired edit. Integration of the sequence from the editing template into the target site relies on template-directed repair of the DSB via the homologous recombination pathway, which is not the dominant pathway for DNA repair in most eukaryotic cells. Furthermore, the endogenous homologous recombination pathway is a complex process with multiple steps, each with its own bottlenecks, which can be difficult to manipulate. Overall, the efficiency of template-directed homologous recombination-mediated editing is typically low in human cells and even lower in plant cells. This is due to the low efficiency of reagent delivery and the difficulty of recovering edited plants.

[0005] The highest efficiency of template-directed editing in eukaryotes other than yeast has been achieved in human cell culture, where the delivery of a cocktail of reagents (e.g., DNA endonucleases or nickases, repair templates, NHEJ inhibitors, and HDR stimulators) can be easily tailored. Specifically, precise template-directed editing in human cells has been demonstrated using a three-component complex: 1) a nickase that can be recruited to a sequence-specific site by a guide RNA; 2) a guide RNA with an extension sequence that binds to the 3' end of the nicked DNA and encodes a repair template with the desired edit; and 3) an RNA-dependent DNA polymerase (reverse transcriptase) fused to the nickase, which uses the 3' end of the nicked DNA and a primer to synthesize DNA (e.g., incorporate the edit). In certain human cell types, precise template-directed editing rates of up to 50% have been reported (Anzalone et al. Nature 576:149-157 (2019)).

[0006] Unlike in human cells, in plants, it can be difficult to deliver multiple reagents in various compositions. Also, it can be difficult to deliver a high dose of repair template, which can improve the efficiency of template-based editing by increasing the availability of repair template in cells. To date, most successful template-based editing in plants has been achieved by particle bombardment of DNA expression cassettes and repair templates. The best editing efficiency is in the range of less than 10%, with many studies being less than 1%. The highest reported efficiency is often only at specific repair loci within the genome, where there is no or insufficient understanding of the mechanisms that may lead to higher efficiency of HDR. Summary of the Invention

[0007] One aspect of the present invention provides a first complex comprising (a) a first sequence-specific DNA-binding protein capable of binding to a first site on a target nucleic acid, and (b) a first DNA-dependent DNA polymerase.

[0008] A second aspect of the present invention provides a first complex comprising: (a) a first sequence-specific DNA-binding protein comprising an endonuclease activity capable of binding to a first site on a target nucleic acid and introducing a single-stranded nick or a double-stranded break; (b) a first DNA-dependent DNA polymerase; and (c) a first DNA-encoded repair template.

[0009] A third aspect of the present invention provides a first complex comprising: (a) a first sequence-specific DNA-binding protein capable of binding to a first site on a target nucleic acid; (b) a first DNA-dependent DNA polymerase; (c) a first DNA endonuclease; and (d) a first DNA-encoded repair template.

[0010] A fourth aspect of the invention provides a second complex comprising (a) a second sequence-specific DNA binding protein capable of binding to a second site on the target nucleic acid, and (b) a DNA-encoded repair template.

[0011] A fifth aspect of the present invention provides an engineered (modified) DNA-dependent DNA polymerase fused to an affinity polypeptide capable of interacting with a peptide tag or an RNA recruiting motif.

[0012] A sixth aspect of the present invention provides an RNA molecule comprising: (a) a nucleic acid sequence that mediates interaction with a CRISPR-Cas effector protein; (b) a nucleic acid sequence that directs the CRISPR-Cas effector protein to a specific nucleic acid target site via a DNA-RNA interaction; and (c) a nucleic acid sequence that forms a stem-loop structure that can interact with an engineered DNA-dependent DNA polymerase of the invention.

[0013] A seventh aspect of the present invention provides a method of modifying a target nucleic acid, the method comprising modifying the target nucleic acid by contacting the target nucleic acid with a first complex of the present invention.

[0014] An eighth aspect of the present invention provides a method for modifying a target nucleic acid, the method comprising modifying the target nucleic acid by contacting the target nucleic acid with (a) a first sequence-specific DNA binding protein capable of binding to a first site on the target nucleic acid, (b) a first DNA-dependent DNA polymerase, (c) a first DNA endonuclease, and (d) a first DNA-encoded repair template.

[0015] A ninth aspect of the present invention provides a method for modifying a target nucleic acid, the method comprising modifying the target nucleic acid by contacting the target nucleic acid with (a) a first sequence-specific DNA-binding protein comprising nickase and / or endonuclease activity capable of binding to a first site on the target nucleic acid and introducing a single-stranded nick or a double-stranded break, (b) a first DNA-dependent DNA polymerase, and (c) a first DNA-encoded repair template.

[0016] A tenth aspect of the present invention provides a system for modifying a target nucleic acid, comprising a first complex of the present invention, a polynucleotide encoding the same, and / or an expression cassette or vector comprising the polynucleotide, wherein (a) a first sequence-specific DNA-binding protein comprising DNA endonuclease activity binds to a first site on the target nucleic acid, (b) a first DNA-dependent DNA polymerase is capable of interacting with the first sequence-specific DNA-binding protein and is recruited to the first sequence-specific DNA-binding protein and to the first site on the target nucleic acid, and (c) (i) a first DNA-encoded The repair template is linked to a first guide nucleic acid comprising a spacer sequence having substantial complementarity to a first site on the target nucleic acid, thereby guiding the first DNA-encoded repair template to the first site on the target nucleic acid, or (c)(ii) the first DNA-encoded repair template is capable of interacting with a first sequence-specific DNA-binding protein or a first DNA-dependent DNA polymerase and is recruited to the first sequence-specific DNA-binding protein or the first DNA-dependent DNA polymerase and to the first site on the target nucleic acid, thereby modifying the target nucleic acid.

[0017] An eleventh aspect of the present invention provides a system for modifying a target nucleic acid, comprising a first complex of the present invention, a polynucleotide encoding the first complex, and / or an expression cassette or vector comprising the polynucleotide, wherein (a) a first sequence-specific DNA-binding protein binds to a first site on the target nucleic acid, (b) a first DNA endonuclease is capable of interacting with the first sequence-specific DNA-binding protein and / or the guide nucleic acid and is recruited to the first sequence-specific DNA-binding protein and to the first site on the target nucleic acid, and (c) a first DNA-dependent DNA polymerase is capable of interacting with the first sequence-specific DNA-binding protein and / or the guide nucleic acid and is recruited to the first sequence-specific DNA-binding protein and to the first site on the target nucleic acid. (d)(i) the first DNA-encoded repair template is linked to a guide nucleic acid comprising a spacer sequence having substantial complementarity to the first site on the target nucleic acid, thereby guiding the first DNA-encoded repair template to the first site on the target nucleic acid; or (d)(ii) the first DNA-encoded repair template is capable of interacting with a first sequence-specific DNA-binding protein or a first DNA-dependent DNA polymerase and is recruited to the sequence-specific DNA-binding protein or the first DNA-dependent DNA polymerase and to the first site on the target nucleic acid, thereby modifying the target nucleic acid. [Sequence List Free Text]

[0018] SEQ ID NOs: 1-20 are exemplary Cas12a amino acid sequences useful in the present invention. SEQ ID NOs: 21-22 are exemplary regulatory sequences encoding a promoter and an intron. SEQ ID NOs: 23-25 ​​provide exemplary peptide tags and corresponding affinity polypeptides. SEQ ID NOs: 26-36 provide exemplary RNA recruitment motifs and corresponding affinity polypeptides. SEQ ID NOs: 37-39 provide examples of protospacer adjacent motif locations for type V CRISPR-Cas12a nucleases. SEQ ID NOs: 40-47 provide exemplary HUH-tags and corresponding recognition sequences. SEQ ID NOs: 48-58 and 88-94 provide exemplary DNA-dependent DNA polymerases from a variety of different organisms. SEQ ID NOs: 59-62 provide exemplary Cas9 sequences. SEQ ID NOs: 63-70 provide exemplary retron reverse transcriptases and retron scaffolds. SEQ ID NOs: 71-74 provide exemplary chimeric guide nucleic acid sequences. SEQ ID NO: 75 provides an exemplary Cas12a ribonucleoprotein (RNP). SEQ ID NOs: 76-87 provide the target and crRNA sequences from Example 11. DETAILED DESCRIPTION OF THE INVENTION

[0019] The present invention will now be described hereinafter with reference to the accompanying drawings and examples illustrating embodiments of the invention. This description is not intended to be a detailed catalog of all the various ways in which the invention may be practiced, nor is it intended to be a complete list of all features that may be added to the present invention. For example, features specifically described with respect to one embodiment may be incorporated into other embodiments, and features specifically described with respect to a particular embodiment may be omitted from that embodiment. Thus, the present invention contemplates that any feature or combination of features set forth herein may be excluded or omitted in some embodiments of the invention. Also, numerous modifications and additions to the various embodiments suggested herein that do not depart from the invention will be apparent to those skilled in the art in light of this disclosure. Therefore, the following description is intended to illustrate some specific embodiments of the invention, and is not intended to exhaustively specify all arrangements, combinations, and variations thereof.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. The terminology used in the description of the invention herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0021] All publications, patent applications, patents, and other references cited herein are incorporated by reference in their entirety for the teachings relevant to the sentence and / or paragraph in which the reference appears.

[0022] Unless the context dictates otherwise, it is specifically contemplated that the various features of the invention described herein can be used in any combination. Moreover, the invention also contemplates that, in some embodiments of the invention, any feature or combination of features set forth herein can be excluded or omitted. To indicate that a specification states that a composition includes components A, B, and C, it is specifically contemplated that any or any combination of A, B, or C, alone or in any combination, can be omitted and waived.

[0023] As used in describing the invention and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise.

[0024] Also, as used herein, "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted as alternatives ("or").

[0025] As used herein, the term "about" when referring to a measurable value, such as an amount or concentration, is intended to encompass a ±10%, ±5%, ±1%, ±0.5%, or ±0.1% variation of the specified value, as well as the specified value. For example, "about X," where X is a measurable value, is intended to include X and a ±10%, ±5%, ±1%, ±0.5%, or ±0.1% variation of X. Ranges provided herein for measurable values ​​can include any other ranges and / or individual values ​​therein.

[0026] As used herein, phrases such as "between X and Y" and "between about X and Y" should be interpreted to include X and Y. As used herein, phrases such as "between about X and Y" mean "between about X and about Y," and phrases such as "about X to Y" mean "about X to about Y."

[0027] The recitation of ranges of values ​​herein, unless otherwise indicated herein, is merely intended to serve as a shorthand method of referring individually to each separate value falling within that range, and each separate value is incorporated herein to the same extent as if it were individually recited herein. For example, if the range 10 to 15 is disclosed, then 11, 12, 13, and 14 are also disclosed.

[0028] As used herein, the terms "comprise," "comprises," and "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0029] As used herein, the transitional phrase "consisting essentially of" means that the claims should be construed to include the specified materials or steps recited in the claims and that do not materially affect the basic and novel characteristics of the claimed invention. Thus, the term "consisting essentially of" when used in the claims of the present invention is not intended to be construed as equivalent to "comprise."

[0030] As used herein, the terms "enhance," "enhancing," "enhancing," "enhancing," "improving," and "enhancing" (and grammatical variations thereof) describe an increase of at least about 25%, 50%, 75%, 100%, 150%, 200%, 300%, 400%, 500%, or more compared to a control.

[0031] As used herein, the terms "reduce," "reduced," "reducing," "reduction," "reduce," and "lower" (and grammatical variations thereof) describe, for example, a reduction of at least about 5%, 10%, 15%, 20%, 25%, 35%, 50%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% compared to a control. In certain embodiments, the reduction may result in no, or essentially no, detectable activity or amount (i.e., an insignificant amount, e.g., less than about 10% or 5%).

[0032] A "heterologous" or "recombinant" nucleotide sequence is a nucleotide sequence that is not naturally associated with the host cell into which it is introduced, and includes non-naturally occurring multiple copies of a naturally occurring nucleotide sequence.

[0033] A "native" or "wild-type" nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence refers to a naturally occurring or endogenous nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence. Thus, for example, a "wild-type mRNA" is an mRNA that naturally occurs in or is endogenous to a reference organism. A "homologous" nucleic acid sequence is a nucleotide sequence that is naturally associated with a host cell into which it is introduced.

[0034] As used herein, the terms "nucleic acid," "nucleic acid molecule," "nucleotide sequence," and "polynucleotide" refer to RNA or DNA, whether linear or branched, single-stranded or double-stranded, or a hybrid thereof. The terms also encompass RNA / DNA hybrids. When dsRNA is produced synthetically, less common bases, such as inosine, 5-methylcytosine, 6-methyladenine, hypoxanthine, and others, can be used for antisense, dsRNA, and ribozyme pairing. For example, polynucleotides containing C-5 propyne analogs of uridine and cytidine have been shown to bind RNA with high affinity and to be potent antisense inhibitors of gene expression. Other modifications, such as modifications to the phosphodiester backbone or the 2'-hydroxyl in the ribose sugar group of RNA, can also be used.

[0035] As used herein, the term "nucleotide sequence" refers to a heteropolymer of nucleotides or a sequence of nucleotides from the 5' to 3' end of a nucleic acid molecule, including DNA or RNA molecules, including cDNA, DNA fragments or portions, genomic DNA, synthetic (e.g., chemically synthesized) DNA, plasmid DNA, mRNA, and antisense RNA (all of which may be single-stranded or double-stranded). The terms "nucleotide sequence," "nucleic acid," "nucleic acid molecule," "nucleic acid construct," "oligonucleotide," and "polynucleotide" are also used interchangeably herein to refer to a heteropolymer of nucleotides. The nucleic acid molecules and / or nucleotide sequences provided herein are presented herein from left to right in the 5' to 3' orientation and are represented using the standard code for designating nucleotide properties as set forth in the U.S. Sequence Code, 37 CFR §§ 1.821-1.825, and World Intellectual Property Organization (WIPO) Standard ST.25. As used herein, the term "5' region" may refer to the region of a polynucleotide closest to the 5' end of the polynucleotide. Thus, for example, an element within the 5' region of a polynucleotide can be located anywhere from a first nucleotide located at the 5' end of the polynucleotide to a nucleotide located within the polynucleotide. As used herein, the term "3' region" can refer to the region of a polynucleotide that is closest to the 3' end of the polynucleotide. Thus, for example, an element within the 3' region of a polynucleotide can be located anywhere from a first nucleotide located at the 3' end of the polynucleotide to a nucleotide located within the polynucleotide.

[0036] As used herein, the term "gene" refers to a nucleic acid molecule that can be used to produce mRNA, antisense RNA, miRNA, anti-microRNA antisense oligodeoxyribonucleotides (AMOs), and the like. A gene may or may not be used to produce a functional protein or gene product. A gene can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and / or 5' and 3' untranslated regions). A gene may be "isolated," which refers to a nucleic acid that is substantially or essentially free from components normally found associated with the nucleic acid in nature. Such components include other cellular material, culture medium from the recombinant product, and / or various chemicals used to chemically synthesize the nucleic acid.

[0037] The term "mutation" refers to point mutations (e.g., missense, or nonsense, or single base pair insertions or deletions resulting in frameshifts), insertions, deletions, and / or truncations. When a mutation is a substitution of a residue in an amino acid sequence with another residue, or a deletion or insertion of one or more residues in the sequence, the mutation is typically described by identifying the position of the residue in the sequence following the original residue and identifying the newly substituted residues.

[0038] As used herein, the term "complementary" or "complementarity" refers to the natural binding of polynucleotides by base pairing under permissive salt and temperature conditions. For example, the sequence "AGT" (5' to 3') binds to the complementary sequence "TCA" (3' to 5'). Complementarity between two single-stranded molecules can be "partial," where only a portion of the nucleotides bind, or complete, where total complementarity exists between the single-stranded molecules. The degree of complementarity between nucleic acid strands has a significant impact on the efficiency and strength of hybridization between nucleic acid strands.

[0039] As used herein, "complement" can mean 100% complementarity with a comparator nucleotide sequence, or it can mean less than 100% complementarity (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% complementarity).

[0040] A "portion" or "fragment" of a nucleotide sequence of the invention is a sequence that is reduced in length (e.g., reduced by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides) compared to a reference nucleic acid or nucleotide sequence, and is identical or nearly identical (e.g., 70%, 71%, "Nucleic acid fragment" is understood to mean a nucleotide sequence comprising, consisting essentially of, and / or consisting of a nucleotide sequence of contiguous nucleotides (72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical). Such a nucleic acid fragment or portion according to the invention may, where appropriate, be comprised within a larger polynucleotide of which it is a component. By way of example, the repeat sequence of a guide nucleic acid of the invention may comprise a portion of a wild-type CRISPR-Cas repeat sequence (e.g., a wild-type CRISR-Cas repeat, e.g., a repeat from a CRISPR Cas system such as Cas9, Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c).

[0041] Different nucleic acids or proteins that share homology are referred to herein as "homologues." The term homologue includes homologous sequences from the same species and other species, as well as orthologous sequences from the same species and other species. "Homology" refers to the level of similarity between two or more nucleic acid and / or amino acid sequences, expressed in terms of percentage of positional identity (i.e., sequence similarity or identity). Homology also refers to the concept of similar functional properties between different nucleic acids or proteins. Thus, the compositions and methods of the present invention further include homologues to the nucleotide and polypeptide sequences of the present invention. As used herein, "ortholog" refers to homologous nucleotide and / or amino acid sequences in different species that arose from a common ancestral gene during speciation. A homologue of a nucleotide sequence of the present invention has substantial sequence identity (e.g., at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%) to said nucleotide sequence of the present invention.

[0042] "Sequence identity," as used herein, refers to the degree to which two optimally aligned polynucleotide or polypeptide sequences are invariant throughout the window of alignment of the components, e.g., nucleotides or amino acids. "Identity" can be readily calculated by known methods, including, but not limited to, those described in Computational Molecular Biology (Lesk, AM, ed.) Oxford University Press, New York (1988), Biocomputing: Informatics and Genome Projects (Smith, DW, ed.) Academic Press, New York (1993), Computer Analysis of Sequence Data, Part I (Griffin, AM, and Griffin, HG, eds.) Humana Press, New Jersey (1994), Sequence Analysis in Molecular Biology (von Heinje, G., ed.) Academic Press (1987), and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.) Stockton Press, New York (1991).

[0043] As used herein, the term "percent sequence identity" or "percent identity" refers to the percentage of identical nucleotides in a linear polynucleotide sequence of a reference ("query") polynucleotide molecule (or its complementary strand) compared to a test ("subject") polynucleotide molecule (or its complementary strand) when the two sequences are optimally aligned. In some embodiments, "percent identity" may refer to the percentage of identical amino acids in an amino acid sequence compared to a reference polypeptide.

[0044] As used herein, the phrase "substantially identical" or "substantial identity" in the context of two nucleic acid molecules, nucleotide sequences, or protein sequences refers to two or more sequences or subsequences that have at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% nucleotide or amino acid residue identity when compared and aligned for maximum correspondence as measured using one of the following sequence comparison algorithms or by visual inspection. In some embodiments of the present invention, substantial identity exists over a region of contiguous nucleotides of the nucleotide sequences of the present invention that is about 10 to about 20 nucleotides, about 10 to about 25 nucleotides, about 10 to about 30 nucleotides, about 15 to about 25 nucleotides, about 30 to about 40 nucleotides, about 50 to about 60 nucleotides, about 70 to about 80 nucleotides, about 90 to about 100 nucleotides, or more, and any range therein, up to the full length of the sequence. In some embodiments, nucleotide sequences may be substantially identical over at least about 20 nucleotides (e.g., about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 nucleotides). In some embodiments, a substantially identical nucleotide or protein sequence performs substantially the same function as the nucleotide sequence (or encoded protein sequence) to which it is substantially identical.

[0045] For sequence comparison, typically, one sequence serves as a reference sequence to which test sequences are compared.When using a sequence comparison algorithm, test sequences and reference sequences are input into a computer, and if necessary, subsequence coordinates are designated, and sequence algorithm program parameters are designated.The sequence comparison algorithm then calculates the percent sequence identity for the test sequence compared to the reference sequence based on the designated program parameters.

[0046] Optimal alignment of sequences for aligning a comparison window is well known to those skilled in the art and can be performed by tools such as the Smith and Waterman local homology algorithm, the Needleman and Wunsch homology alignment algorithm, the Pearson and Lipman similarity search method, and in some cases by computer implementations of these algorithms, such as GAP, BESTFIT, FASTA, and TFASTA, available as part of the GCG® Wisconsin Package® (Accelrys Inc., San Diego, CA). The "percent identity" for an aligned segment of a test sequence and a reference sequence is the number of identical components shared by the two aligned sequences divided by the total number of components in the reference sequence segment, for example, the entire reference sequence, or a defined smaller portion of the reference sequence. The percent sequence identity is expressed as the percent identity multiplied by 100. Comparison of one or more polynucleotide sequences can be to the full-length polynucleotide sequence or a portion thereof, or to a longer polynucleotide sequence. For purposes of the present invention, "percent identity" may also be determined using BLASTX version 2.0 for translated nucleotide sequences and BLASTN version 2.0 for polynucleotide sequences.

[0047] Two nucleotide sequences may also be considered to be substantially complementary if the two sequences hybridize to each other under stringent conditions. In some representative embodiments, two nucleotide sequences considered to be substantially complementary hybridize to each other under highly stringent conditions.

[0048] "Stringent hybridization conditions" and "stringent hybridization wash conditions" in the context of nucleic acid hybridization experiments such as Southern and Northern hybridization are sequence-dependent and vary under various environmental parameters. An extensive guide to nucleic acid hybridization can be found in Tijssen Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, part I, chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assays," Elsevier, New York (1993). Typically, highly stringent hybridization and wash conditions are those that achieve the thermal melting point (T) for a particular sequence at a defined ionic strength and pH. m ) is chosen to be approximately 5°C lower than

[0049] T m is the temperature (under defined ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched probe. mVery stringent conditions are selected so that the hybridization is equivalent to 1×SSC at 45°C. An example of stringent hybridization conditions for hybridization of complementary nucleotide sequences with more than 100 complementary residues on a filter in a Southern or Northern blot is 50% formamide with 1 mg heparin at 42°C, with hybridization performed overnight. An example of highly stringent washing conditions is 0.15 M NaCl at 72°C for approximately 15 minutes. An example of stringent washing conditions is a 0.2×SSC wash at 65°C for 15 minutes (see Sambrook, below, for a description of SSC buffers). Often, a high stringency wash follows a low stringency wash to remove background probe signal. An example of a moderate stringency wash, for example, for a duplex of more than 100 nucleotides, is 1×SSC at 45°C for 15 minutes. An example of a low stringency wash for a duplex of, for example, more than 100 nucleotides is 4-6×SSC at 40°C for 15 minutes. For short probes (e.g., about 10-50 nucleotides), stringent conditions typically include a salt concentration of less than about 1.0 M Na ion, typically about 0.01-1.0 M Na ion (or other salt) at pH 7.0-8.3, and a temperature typically of at least about 30°C. Stringent conditions can also be achieved by adding destabilizing agents such as formamide. Generally, a signal-to-noise ratio of 2× (or higher) than that observed for an unrelated probe in a particular hybridization assay indicates detection of specific hybridization. Nucleotide sequences that do not hybridize to each other under stringent conditions are also substantially identical if the proteins they encode are substantially identical. This can occur, for example, when copies of nucleotide sequences are generated using the maximum codon degeneracy permitted by the genetic code.

[0050] Any polynucleotide, nucleic acid construct, expression cassette, and / or vector of the present invention may be codon-optimized for expression in any species of interest. Codon optimization is well known in the art and involves modifying a nucleotide sequence for codon usage bias using a species-specific codon usage table. The codon usage table is generated based on sequence analysis of the most highly expressed genes for the species of interest. If the nucleotide sequence is to be expressed in the nucleus, the codon usage table is generated based on sequence analysis of highly expressed nuclear genes for the species of interest. Modification of the nucleotide sequence is determined by comparing the species-specific codon usage table with the codons present in the native polynucleotide sequence. As is understood in the art, codon optimization of a nucleotide sequence results in a nucleotide sequence that has less than 100% identity to a native nucleotide sequence (e.g., 50%, 60%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99%, etc.) but still encodes a polypeptide having the same function (and in some embodiments, the same structure) as that encoded by the original nucleotide sequence. Thus, in some embodiments, the polynucleotides, nucleic acid constructs, expression cassettes, and / or vectors of the invention (e.g., comprising / encoding sequence-specific DNA binding domains, DNA-dependent DNA polymerases, DNA endonucleases, etc.) may be codon-optimized for expression in an organism (e.g., a plant (e.g., a particular plant species), an animal, a bacterium, a fungus, etc.).In some embodiments, codon-optimized nucleic acid constructs, polynucleotides, expression cassettes, and / or vectors of the invention have about 70% to about 99.9% or more (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100%) identity to polynucleotides, nucleic acid constructs, expression cassettes, and / or vectors of the invention that are not codon-optimized.

[0051] In any of the embodiments described herein, the polynucleotides or nucleic acid constructs of the present invention may be operably associated with various promoters and / or other regulatory elements for expression in plants and / or plant cells. Thus, in some embodiments, the polynucleotides or nucleic acid constructs of the present invention may further comprise one or more promoters, introns, enhancers, and / or terminators operably linked to one or more nucleotide sequences. In some embodiments, a promoter may be operably associated with an intron (e.g., the Ubi1 promoter and an intron). In some embodiments, a promoter associated with an intron may be referred to as a "promoter region" (e.g., the Ubi1 promoter and an intron).

[0052] As used herein with respect to polynucleotides, "operably linked" or "operably associated" means that the elements indicated are functionally related, and usually physically related, to each other. Thus, as used herein, the terms "operably linked" or "operably associated" refer to nucleotide sequences on a single nucleic acid molecule that are functionally related. Thus, a first nucleotide sequence operably linked to a second nucleotide sequence refers to a situation in which the first nucleotide sequence is placed in a functional relationship with the second nucleotide sequence. For example, a promoter is operably associated with a nucleotide sequence if the promoter effects the transcription or expression of the nucleotide sequence. Those skilled in the art will understand that a regulatory sequence (e.g., a promoter) need not be contiguous with an operably associated nucleotide sequence, so long as it functions to direct its expression. Thus, for example, an intervening nucleic acid sequence that is transcribed but not translated can be present between the promoter and the nucleotide sequence, and the promoter can still be considered "operably linked" to the nucleotide sequence.

[0053] As used herein, the term "linked" with respect to polypeptides refers to the attachment of one polypeptide to another. A polypeptide may be linked to another polypeptide (at the N-terminus or C-terminus) directly (e.g., via a peptide bond) or via a linker.

[0054] The term "linker" is art-recognized and refers to a chemical group or molecule that links two molecules or moieties, such as a fusion protein, e.g., a DNA-binding polypeptide or domain and a peptide tag and / or a reverse transcriptase and an affinity polypeptide that binds to the peptide tag, or two domains: a DNA endonuclease polypeptide or domain and a peptide tag and / or a reverse transcriptase and an affinity polypeptide that binds to the peptide tag. A linker may consist of a single linking molecule or may include multiple linking molecules. In some embodiments, a linker can be an organic molecule, group, polymer, or chemical moiety, such as a bivalent organic moiety. In some embodiments, a linker can be an amino acid or a peptide. In some embodiments, a linker is a peptide.

[0055] In some embodiments, peptide linkers useful in the present invention are from about 2 to about 100 or more amino acids in length, e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55 , 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids in length (e.g., about 2 to about 40, about 2 to about 50, about 2 to about 60, about 4 to about 40, about 4 to about 50, about 4 to about 60, about 5 to about 40, about 5 to about 50 , about 5 to about 60, about 9 to about 40, about 9 to about 50, about 9 to about 60, about 10 to about 40, about 10 to about 50, about 10 to about 60, or about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110 The peptide linker may be 4, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, or more amino acids in length (e.g., about 105, 110, 115, 120, 130, 140, 150, or more amino acids in length). In some embodiments, the peptide linker may be a GS linker.

[0056] A "promoter" is a nucleotide sequence that controls or regulates the transcription of a nucleotide sequence (e.g., a coding sequence) operably associated with the promoter. The coding sequence controlled or regulated by a promoter can encode a polypeptide and / or functional RNA. Typically, a "promoter" refers to a nucleotide sequence that contains a binding site for RNA polymerase II and directs the initiation of transcription. In general, promoters are found 5', or upstream, to the start of the coding region of the corresponding coding sequence. A promoter may contain other elements that act as regulators of gene expression, such as a promoter region, which includes a TATA box consensus sequence and often a CAAT box consensus sequence (Breathnach and Chambon, (1981) Annu. Rev. Biochem. 50:349). In plants, the CAAT box may be replaced by an AGGA box (Messing et al., (1983) in Genetic Engineering of Plants, T. Kosuge, C. Meredith and A. Hollaender (eds.), Plenum Press, pp. 211-227). In some embodiments, the promoter region may contain at least one intron (e.g., SEQ ID NO: 21 or 22).

[0057] Examples of promoters useful in the present invention may include constitutive, inducible, temporally regulated, developmentally regulated, chemically regulated, tissue-preferred and / or tissue-specific promoters used in the preparation of recombinant nucleic acid molecules, such as "synthetic nucleic acid constructs" or "protein-RNA complexes." These various types of promoters are known in the art.

[0058] The choice of promoter may vary depending on the temporal and spatial requirements of expression, and may vary based on the host cell to be transformed.Promoters for many different organisms are well known in the art.Based on the extensive knowledge existing in the art, a promoter suitable for a particular host organism of interest can be selected.Thus, for example, a great deal is known about the promoters upstream of highly constitutively expressed genes in model organisms, and such knowledge can be easily accessed and implemented in other systems as needed.

[0059] In some embodiments, promoters functional in plants may be used in the constructs of the present invention. Non-limiting examples of promoters useful for driving expression in plants include the promoter of RubisCo small subunit gene 1 (PrbcS1), the promoter of actin gene (Pactin), the promoter of nitrate reductase gene (Pnr), and the promoter of double carbon anhydrase gene 1 (Pdca1) (see Walker et al. Plant Cell Rep. 23:727-735 (2005); Li et al. Gene 403:132-142 (2007); Li et al. Mol Biol. Rep. 37:1143-1154 (2010)). PrbcS1 and Pactin are constitutive promoters, while Pnr and Pdca1 are inducible promoters. Pnr is induced by nitrate and repressed by ammonium (Li et al. Gene 403:132-142 (2007)), and Pdca1 is induced by salt (Li et al. Mol Biol Rep. 37:1143-1154 (2010)). In some embodiments, promoters useful in the present invention are RNA polymerase II (Pol II) promoters. In some embodiments, the U6 promoter or 7SL promoter from maize (Zea mays) may be useful in the constructs of the present invention. In some embodiments, the U6c promoter and / or 7SL promoter from maize may be useful for driving expression of a guide nucleic acid. In some embodiments, the U6c promoter, U6i promoter, and / or 7SL promoter from soybean (Glycine max) may be useful in the constructs of the present invention. In some embodiments, the U6c promoter, U6i promoter, and / or 7SL promoter from soybean may be useful for driving expression of a guide nucleic acid.

[0060] Examples of constitutive promoters useful in plants include, but are not limited to, the cestrum virus promoter (cmp) (U.S. Pat. No. 7,166,770), the rice actin 1 promoter (Wang et al. (1992) Mol. Cell. Biol. 12:3399-3406, and U.S. Pat. No. 5,641,876), the CaMV35S promoter (Odell et al. (1985) Nature 313:810-812), the CaMV19S promoter (Lawton et al. (1987) Plant Mol. Biol. 9:315-324), the nos promoter (Ebert et al. (1987) Proc. Natl. Acad. Sci USA 84:5745-5749), the Adh promoter (Walker et al. (1987) Proc. Natl. Acad. Sci USA 84:5745-5749), the ribosomal protein promoter (Richardson ... (1987) Proc. Natl. Acad. Sci. USA 84:6624-6629), the sucrose synthase promoter (Yang & Russell (1990) Proc. Natl. Acad. Sci. USA 87:4144-4148), and the ubiquitin promoter. Constitutive promoters derived from ubiquitin have accumulated in many cell types. Ubiquitin promoters have been cloned from several plant species used in transgenic plants, such as sunflower (Binet et al., 1991. Plant Science 79:87-94), maize (Christensen et al., 1989. Plant Molec. Biol. 12:619-632), and Arabidopsis (Norris et al., 1993. Plant Molec. Biol. 21:895-906). The maize ubiquitin promoter (UbiP) has been exploited in transgenic monocotyledonous plant systems, and its sequence and vectors constructed for monocotyledonous plant transformation are disclosed in European Patent Application Publication No. 0 342 926. The ubiquitin promoter is suitable for expression of the nucleotide sequences of the present invention in transgenic plants, particularly monocotyledonous plants.Additionally, the promoter expression cassettes described by McElroy et al. (Mol. Gen. Genet. 231:150-160 (1991)) can be readily modified for expression of the nucleotide sequences of the present invention and are particularly suitable for use in monocotyledonous hosts.

[0061] In some embodiments, tissue-specific / tissue-preferred promoters can be used to express heterologous polynucleotides in plant cells. Tissue-specific or tissue-preferred expression patterns include, but are not limited to, green tissue-specific or green tissue-preferred, root-specific or root-preferred, stem-specific or stem-preferred, flower-specific or flower-preferred, or pollen-specific or pollen-preferred. Promoters suitable for expression in green tissues include many that regulate genes involved in photosynthesis, many of which have been cloned from both monocotyledonous and dicotyledonous plants. In one embodiment, a promoter useful in the present invention is the maize PEPC promoter from the phosphoenol carboxylase gene (Hudspeth & Grula, Plant Molec. Biol. 12:579-589 (1989)). Non-limiting examples of tissue-specific promoters include those associated with genes encoding seed storage proteins (e.g., β-conglycinin, cruciferin, napin, and phaseolin), zein or oil body proteins (e.g., oleosin), or proteins involved in fatty acid biosynthesis, including acyl carrier protein, stearoyl-ACP desaturase, and fatty acid desaturase (fad2-1)), and other nucleic acids expressed during embryogenesis (e.g., Bce4; see, e.g., Kridl et al. (1991) Seed Sci. Res. 1:209-219, and European Patent Application Publication No. 255378). Tissue-specific or tissue-preferred promoters useful for expression of the nucleotide sequences of the invention in plants, particularly maize, include, but are not limited to, those directing expression in roots, pith, leaves, or pollen. Such promoters are disclosed, for example, in WO 93 / 07278, which is incorporated herein by reference in its entirety.Other non-limiting examples of tissue-specific or tissue-preferred promoters useful in the present invention include the Watarubisco promoter disclosed in U.S. Pat. No. 6,040,504, the inesucrose synthase promoter disclosed in U.S. Pat. No. 5,604,121, the root-specific promoter described by de Framond (FEBS 290:103-106 (1991), European Patent Application Publication No. 0452269 to Ciba-Geigy), the stem-specific promoter driving expression of the maize trpA gene described in U.S. Pat. No. 5,625,136 (Ciba-Geigy), the Cestrum yellow leaf curling virus promoter disclosed in WO 01 / 73087, and pollen-specific or pollen-preferred promoters such as ProOsLPS10 and ProOsLPS11 from rice (Nguyen et al., Plant Cell Pathology 2004). Biotechnol. Reports 9(5):297-306(2015)), ZmSTK2_USP from maize (Wang et al. Genome 60(6):485-495(2017)), LAT52 and LAT59 from tomato (Twell et al. Development 109(3):705-713(1990)), Zm13 (U.S. Patent No. 10,421,972), the PLA2-δ promoter from Arabidopsis (U.S. Patent No. 7,141,424), and / or the ZmC5 promoter from maize (WO 1999 / 042587).

[0062] Further examples of plant tissue-specific / tissue-preferred promoters include, but are not limited to, root hair-specific cis-elements (RHE) (Kim et al. The Plant Cell 18:2958-2970 (2006)), root-specific promoters RCc3 (Jeong et al. Plant Physiol. 153:185-197 (2010)) and RB7 (U.S. Patent No. 5,459,252), lectin promoters (Lindstrom et al. (1990) Der. Genet. 11:160-167 and Vodkin (1983) Prog. Clin. Biol. Res. 138:87-98), maize alcohol dehydrogenase 1 promoter (Dennis et al. (1984) Nucleic Acids Res. 12:3983-4000), S-adenosyl-L-methionine synthase (SAM) (Vander Mijnsbrugge et al. al. (1996) Plant and Cell Physiology, 37(8):1108-1115), maize light-harvesting complex promoter (Bansal et al. (1992) Proc. Natl. Acad. Sci. USA 89:3654-3658), maize heat shock protein promoter (O'Dell et al. (1985) EMBO J. 5:451-458 and Rochester et al. (1986) EMBO J. 5:451-458), pea small subunit RuBP carboxylase promoter (Cashmore, "Nuclear genes encoding the small subunit of ribulose-1,5-bisphosphate carboxylase," pp. 29-39 In: Genetic Engineering of Plants (Hollaender ed., Plenum Press 1983), and Poulsen et al. al. (1986) Mol. Gen. Genet. 205:193-200), Ti plasmid mannopine synthase promoter (Langridge et al. (1989) Proc. Natl. Acad. Sci. USA 86:3219-3223), Ti plasmid nopaline synthase promoter (Langridge et al.(1989), supra), petunia chalcone isomerase promoter (van Tunen et al. (1988) EMBO J. 7:1257-1263), common bean glycine-rich protein 1 promoter (Keller et al. (1989) Genes Dev. 3:1639-1646), truncated CaMV35S promoter (O'Dell et al. (1985) Nature 313:810-812), potato patatin promoter (Wenzler et al. (1989) Plant Mol. Biol. 13:347-354), root cell promoter (Yamamoto et al. (1990) Nucleic Acids Res. 18:7449), maize zein promoter (Kriz et al. (1987) Mol. Gen. Genet. 207:90-98, Langridge et al. (1987) Mol. Gen. Genet. 207:90-98, Langridge et al. (1987) Nature 313:810-812), and the like. al. (1983) Cell 34:1015-1022, Reina et al. (1990) Nucleic Acids Res. 18:6425, Reina et al. (1990) Nucleic Acids Res. 18:7449, and Wandelt et al. (1989) Nucleic Acids Res. 17:2354), the globulin-1 promoter (Belanger et al. (1991) Genetics 129:863-872), the α-tubulin cab promoter (Sullivan et al. (1989) Mol. Gen. Genet. 215:431-440), the PEPCase promoter (Hudspeth & Grula (1989) Plant Mol. Biol. 12:579-589), the R gene complex-associated promoter (Chandler et al. (1989) Plant Cell 1:1175-1183), and the chalcone synthase promoter (Franken et al. (1991) EMBO J. 10:2605-2612).

[0063] Useful for seed-specific expression are the pea vicillin promoter (Czako et al. (1992) Mol. Gen. Genet. 235:33-40) and the seed-specific promoters disclosed in U.S. Patent No. 5,625,136. Useful promoters for expression in mature leaves include those that are switched on at the onset of senescence, such as the SAG promoter from Arabidopsis (Gan et al. (1995) Science 270:1986-1988).

[0064] Alternatively, promoters functional in chloroplasts can be used. Non-limiting examples of such promoters include the bacteriophage T3 gene 9 5'UTR and other promoters disclosed in U.S. Patent No. 7,579,516. Other promoters useful in the present invention include, but are not limited to, the S-E9 small subunit RuBP carboxylase promoter and the Kunitz trypsin inhibitor gene promoter (Kti3).

[0065] Additional regulatory elements useful in the present invention include, but are not limited to, introns, enhancers, termination sequences, and / or 5' and 3' untranslated regions.

[0066] Introns useful in the present invention may be identified in plants, isolated from the plant, and then inserted into an expression cassette to be used for plant transformation. As will be understood by those skilled in the art, introns can contain sequences required for self-excision and are incorporated in frame into the nucleic acid construct / expression cassette. Introns can be used as spacers to separate multiple protein-coding sequences within a single nucleic acid construct, or introns can be used within a single protein-coding sequence, for example, to stabilize mRNA. If used within a protein-coding sequence, the intron is inserted "in frame" with the excision site. Introns can also be associated with a promoter to enhance or modify expression. Exemplary promoter / intron combinations useful in the present invention include, but are not limited to, the maize Ubi1 promoter and intron.

[0067] Non-limiting examples of introns useful in the present invention include introns from the ADHI gene (e.g., Adh1-S introns 1, 2, and 6), the ubiquitin gene (Ubi1), the RuBisCo small subunit (rbcS) gene, the RuBisCo large subunit (rbcL) gene, the actin gene (e.g., the actin-1 intron), the pyruvate dehydrogenase kinase gene (pdk), the nitrate reductase gene (nr), the duplicated carbonic anhydrase gene 1 (Tdca1), the psbA gene, the atpA gene, or any combination thereof.

[0068] In some embodiments, the polynucleotides and / or nucleic acid constructs of the present invention may be "expression cassettes" or may be contained within an expression cassette. As used herein, "expression cassette" refers to a recombinant nucleic acid molecule comprising, for example, a nucleic acid construct of the present invention (e.g., a sequence-specific DNA-binding polypeptide or domain, a DNA-dependent DNA polymerase (e.g., an engineered DNA-dependent DNA polymerase), a DNA endonuclease polypeptide or domain, a DNA-encoded repair template, a guide nucleic acid, a first complex, a second complex, a third complex, etc.), where the nucleic acid construct is operably associated with one or more control sequences (e.g., a promoter, a terminator, etc.). Thus, some embodiments of the present invention provide, for example, an expression cassette designed to express one or more polynucleotides of the present invention. When an expression cassette comprises multiple polynucleotides, the polynucleotides may be operably linked to a single promoter that drives expression of all of the polynucleotides, or the polynucleotides may be operably linked to one or more separate promoters (e.g., three polynucleotides may be driven by one, two, or three promoters in any combination). When two or more separate promoters are used, the promoters may be the same promoter or different promoters.Thus, for example, the polynucleotide encoding the sequence-specific DNA-binding polypeptide or domain, the polynucleotide encoding the DNA endonuclease polypeptide or domain, the polynucleotide encoding the DNA-dependent DNA polymerase polypeptide or domain, the repair template encoded by DNA, and / or the guide nucleic acid (if contained in the expression cassette) may each be operably linked to a separate promoter, or may be operably linked to two or more promoters in any combination.In some embodiments, the expression cassette and / or the polynucleotide contained therein may be optimized for expression in plants.

[0069] An expression cassette comprising a nucleic acid construct of the invention may be chimeric, meaning that at least one of its components is heterologous with respect to at least one of its other components (e.g., a promoter from a host organism operably linked to a polynucleotide of interest to be expressed in the host organism, where the polynucleotide of interest is from an organism different from the host or is not normally found in association with that promoter). Alternatively, the expression cassette may be naturally occurring but obtained in a recombinant form useful for heterologous expression.

[0070] The expression cassette may optionally contain transcriptional and / or translational termination regions (i.e., termination regions) and / or enhancer regions that are functional in the selected host cell. A variety of transcription terminators and enhancers are known in the art and available for use in expression cassettes. The transcription terminator is responsible for terminating transcription and ensuring accurate mRNA polyadenylation. The termination and / or enhancer regions may be native to the transcription initiation region, may be native to the gene encoding the sequence-specific DNA-binding polypeptide, the gene encoding the DNA endonuclease polypeptide, the gene encoding the DNA-dependent DNA polymerase, etc., which may be native to the host cell or may be native to another source (e.g., the promoter, the gene encoding the sequence-specific DNA-binding polypeptide, the gene encoding the DNA endonuclease polypeptide, the gene encoding the DNA-dependent DNA polymerase, etc., may be foreign or heterologous to the host cell, or any combination thereof).

[0071] The expression cassettes of the present invention can also include a polynucleotide encoding a selectable marker, which can be used to select transformed host cells. As used herein, a "selectable marker" refers to a polynucleotide sequence that, when expressed, confers a distinct phenotype on host cells expressing the marker, thereby allowing such transformed cells to be distinguished from cells that do not possess the marker. Such a polynucleotide sequence can encode a selectable or screenable marker, depending on whether the marker confers a trait that can be selected for by chemical means, such as by using a selection agent (e.g., an antibiotic), or whether the marker is a trait that can be identified simply through observation or testing, such as by screening (e.g., fluorescence). Many examples of suitable selectable markers are known in the art and can be used in the expression cassettes described herein.

[0072] In addition to expression cassettes, the nucleic acid molecules / constructs and polynucleotide sequences described herein can be used in conjunction with vectors. The term "vector" refers to a composition for transferring, delivering, or introducing a nucleic acid(s) into a cell. A vector includes a nucleic acid construct containing a nucleotide sequence to be transferred, delivered, or introduced. Vectors used to transform host organisms are well known in the art. Non-limiting examples of general classes of vectors include viral vectors, plasmid vectors, phage vectors, phagemid vectors, cosmid vectors, fosmid vectors, bacteriophages, artificial chromosomes, minicircles, or Agrobacterium binary vectors in double- or single-stranded, linear, or circular form, which may or may not be self-transmissible or mobilizable. In some embodiments, viral vectors may include, but are not limited to, retroviral, lentiviral, adenoviral, adeno-associated viral, or herpes simplex viral vectors. Vectors as defined herein are capable of transforming prokaryotic or eukaryotic hosts either by integration into a cellular genome or by extrachromosomal presence (e.g., an autonomously replicating plasmid with an origin of replication). Additionally, included are shuttle vectors, which refer to DNA vehicles capable of replication, naturally or by design, in two different host organisms, which may be selected from actinomycetes and related species, bacteria, and eukaryotes (e.g., higher plant, mammalian, yeast, or fungal cells). In some embodiments, the nucleic acid in the vector is under the control of, and operably linked to, a promoter or other regulatory elements suitable for transcription in the host cell. The vector may also be a bifunctional expression vector that functions in multiple hosts. In the case of genomic DNA, it may contain its own promoter and / or other regulatory elements, and in the case of cDNA, it may be under the control of a promoter and / or other regulatory elements suitable for expression in the host cell.Thus, the nucleic acid constructs or polynucleotides of the invention, and / or expression cassettes comprising same, may be included within the vectors described herein and known in the art.

[0073] As used herein, "contact," "contacting," "contacted," and grammatical variations thereof refer to bringing together the components of a desired reaction (e.g., transformation, transcriptional regulation, genome editing, nicking, and / or cleavage) under conditions suitable for carrying out the desired reaction. As a non-limiting example, a target nucleic acid may be contacted with a sequence-specific DNA binding domain, a DNA endonuclease, a DNA-dependent DNA polymerase, a DNA-encoded repair template, a guide nucleic acid, and / or a nucleic acid construct / expression cassette encoding / comprising the same under conditions in which the sequence-specific DNA binding protein, the DNA endonuclease, and the DNA-dependent DNA polymerase are expressed, and the sequence-specific DNA binding protein binds to the target nucleic acid, and the DNA-dependent DNA polymerase is fused to or recruited to the sequence-specific DNA binding protein (e.g., via a peptide tag fused to the sequence-specific DNA binding protein and an affinity polypeptide (e.g., a polypeptide capable of binding to the peptide tag) fused to the DNA-dependent DNA polymerase), thereby modifying the target nucleic acid by recruiting the DNA-dependent DNA polymerase to the vicinity of the target nucleic acid.

[0074] As used herein, "modifying" or "modification" with respect to a target nucleic acid includes editing (e.g., mutation), covalent modification, nucleic acid / nucleotide base exchange / substitution, deletion, truncation, nicking, and / or transcriptional regulation of the target nucleic acid. In some embodiments, modifications may include indels of any size and / or single base changes (SNPs) of any type.

[0075] "Introducing," "introduce," "introduced" (and grammatical variations thereof) in the context of a polynucleotide of interest means presenting a nucleotide sequence of interest (e.g., a polynucleotide, a nucleic acid construct, and / or a guide nucleic acid) to a host organism or a cell of said organism (e.g., a host cell, e.g., a plant cell, an animal cell, a bacterial cell, a fungal cell) so that the nucleotide sequence has access to the interior of the cell.

[0076] The terms "transformation" or "transfection" may be used interchangeably and, as used herein, refer to the introduction of heterologous nucleic acid into a cell. Cellular transformation may be stable or transient. Thus, in some embodiments, a host cell or host organism may be stably transformed with a polynucleotide / nucleic acid molecule of the present invention. In some embodiments, a host cell or host organism may be transiently transformed with a nucleic acid construct of the present invention.

[0077] "Transient transformation" in the context of a polynucleotide means that the polynucleotide is introduced into a cell without integrating into the genome of the cell.

[0078] By "stably introducing" or "stably introduced," in the context of a polynucleotide introduced into a cell, is meant that the introduced polynucleotide is stably integrated into the genome of the cell, and the cell is stably transformed with the polynucleotide.

[0079] As used herein, "stable transformation" or "stably transformed" means that a nucleic acid molecule is introduced into a cell and integrated into the cell's genome. Thus, the integrated nucleic acid molecule can be inherited by its progeny, more particularly by its progeny for multiple successive generations. As used herein, "genome" includes the nuclear genome and the plastid genome, and thus includes, for example, integration of a nucleic acid into the chloroplast genome or mitochondrial genome. As used herein, stable transformation can also refer to a transgene that is maintained extrachromosomally, for example, as a minichromosome or plasmid.

[0080] Transient transformation can be detected, for example, by enzyme-linked immunosorbent assay (ELISA) or Western blot, which can detect the presence of peptides or polypeptides encoded by one or more transgenes introduced into an organism. Stable transformation of cells can be detected, for example, by Southern blot hybridization assay of the cell's genomic DNA with a nucleic acid sequence that specifically hybridizes with the nucleotide sequence of the transgene introduced into the organism (e.g., a plant). Stable transformation of cells can be detected, for example, by Northern blot hybridization assay of the cell's RNA with a nucleic acid sequence that specifically hybridizes with the nucleotide sequence of the transgene introduced into the host organism. Stable transformation of cells can also be detected, for example, by polymerase chain reaction (PCR) or other amplification reactions well known in the art, which use specific primer sequences that hybridize with the target sequence of the transgene, resulting in amplification of the transgene sequence, which can be detected according to standard methods. Transformation can also be detected by direct sequencing and / or hybridization protocols well known in the art.

[0081] Thus, in some embodiments, the nucleotide sequences, polynucleotides, nucleic acid constructs, and / or expression cassettes of the present invention can be transiently expressed and / or stably integrated into the genome of a host organism. Thus, in some embodiments, the nucleic acid constructs of the present invention (e.g., one or more expression cassettes encoding, for example, a sequence-specific DNA-binding polypeptide or domain, a DNA endonuclease polypeptide or domain, a DNA-dependent DNA polymerase polypeptide or domain, etc.) are transiently introduced into a cell via a guide nucleic acid, such that the DNA cannot be maintained within the cell.

[0082] The nucleic acid constructs of the present invention can be introduced into cells by any method known to those of skill in the art. In some embodiments of the present invention, transformation of cells comprises nuclear transformation. In other embodiments, transformation of cells comprises plastid transformation (e.g., chloroplast transformation). In further embodiments, recombinant nucleic acid constructs of the present invention can be introduced into cells via conventional breeding techniques.

[0083] Procedures for transforming both eukaryotes and prokaryotes are well known and routine in the art and are described throughout the literature (see, e.g., Jiang et al. 2013. Nat. Biotechnol. 31:233-239; Ran et al. Nature Protocols 8:2281-2308 (2013)).

[0084] Thus, nucleotide sequences can be introduced into a host organism or its cells by any number of methods well known in the art. The methods of the present invention introduce one or more nucleotide sequences into an organism, but do not rely on a particular method for accessing the interior of at least one cell of the organism. When multiple nucleotide sequences are to be introduced, they can be assembled as part of a single nucleic acid construct or as separate nucleic acid constructs, and can be located on the same or different nucleic acid constructs. Thus, nucleotide sequences can be introduced into a cell of interest in a single transformation event and / or in separate transformation events, or, if relevant, the nucleotide sequences can be incorporated into a plant, for example, as part of a breeding protocol.

[0085] Endogenous DSB repair by homologous recombination faces competition from the non-homologous end-joining pathway, which is difficult to engineer and prone to errors. In the present invention, template-based editing is an improved bypass process for DSBs that reduces repair efficiency. A novel combination of polypeptides and nucleic acids, as well as protein-protein fusions and non-covalent recruitment, is used to deliver high-fidelity processive or partitioning DNA polymerases and repair templates to target sites in a sequence-specific manner. The target site is cleaved or nicked by a sequence-specific DNA-binding domain containing DNA endonuclease or nickase activity, or a DNA endonuclease with endonuclease or nickase activity, provided in combination with a sequence-specific DNA-binding protein. Using the DNA-encoded repair template and the target DNA with a single-strand nick or double-strand break as primers, a DNA-dependent DNA polymerase can immediately initiate DNA synthesis to copy the desired mutation or large insertion into the target site. The present invention can be used to generate specific changes of a single base or a few bases, deletions of defined genomic sequences, or insertions of small or large fragments.

[0086] As described herein, several DNA recruitment strategies can be used to improve the delivery of repair templates to targets, including HUH-tags, DNA aptamers, bacterial retron msDNA, and / or T-DNA recruitment. One specific example for improving template availability is the use of PCV, a type of HUH-tag. For example, a PCV domain can be fused to a CRISPR-Cas effector protein with nickase or endonuclease activity, which creates a nick or cleavage in the target nucleic acid. Because the PCV recognition site sequence is contained within the repair template, the repair template can be recruited to the target site by interacting with the corresponding PCV domain. Recruitment can occur approximately simultaneously with the creation of a nick or cleavage in the target nucleic acid by the CRISPR-Cas effector protein.

[0087] DNA-dependent DNA polymerase is a key component in carrying out homologous recombination. The 3' end of the target nucleic acid containing a single-stranded nick or double-stranded break can anneal to the repair template encoded by DNA and serve as a primer for the DNA-dependent DNA polymerase, which initiates strand synthesis to copy genetic information from the repair template to the target site. In some embodiments, the DNA polymerase used in this process can have high fidelity to prevent errors and / or high processivity to ensure that long templates are copied before the DNA polymerase dissociates. In the context of the association of the DNA-dependent DNA polymerase with the CRISPR-Cas effector polypeptide / complex that binds to the target nucleic acid, it can be advantageous to have a DNA-dependent DNA polymerase with a partitioning function that maximizes the efficiency of template integration into the target. To facilitate this process, DNA-dependent DNA polymerases with high fidelity and processivity and partitioning profiles can be recruited, for example, via protein fusion or noncovalent interactions with sequence-specific DNA-binding domains and DNA endonucleases (e.g., CRISPR-Cas effector proteins). Direct fusion can be achieved via optimized linker architectures. Noncovalent recruitment strategies can include recruitment via guide nucleic acids (e.g., RNA recruitment motifs, e.g., MS2 loops) or via sequence-specific DNA-binding domains (e.g., CRISPR-Cas effector proteins) and / or DNA endonucleases (e.g., via peptide tags, e.g., antibody / epitope interactions, e.g., SunTag). Of course, the present invention is not limited by these particular recruitment techniques, and any other known or later developed protein-protein or nucleic acid-protein recruitment techniques, now known or later developed, may be used to carry out the present invention.

[0088] The present inventors have developed compositions and methods that achieve improved template-based editing. Using a combination of protein-protein fusion and non-covalent recruitment, a high-fidelity processive or distributive DNA polymerase is sequence-specifically delivered to a target site, which may be cleaved or nicked, for example, by a CRISPR endonuclease or nickase. In combination with a DNA-encoded repair template, the DNA-dependent DNA polymerase can immediately initiate DNA synthesis by using the target DNA with a single-strand nick or double-strand break as a primer to copy the desired mutation or large insertion into the target site. The present invention and its variations described herein can be used to create specific changes of a single base or a few bases, deletions of defined genomic sequences, or insertions of small or large fragments.

[0089] Thus, in some embodiments, the present invention provides a complex (e.g., a first complex) comprising: (a) a sequence-specific DNA-binding protein (e.g., a first sequence-specific DNA-binding protein) capable of binding to a site (e.g., a first site) on a target nucleic acid; and (b) a DNA-dependent DNA polymerase (e.g., a first DNA-dependent DNA polymerase). In some embodiments, the complex may comprise a DNA-encoded repair template (e.g., a first DNA-encoded repair template). In some embodiments, the complex may comprise a DNA endonuclease (e.g., a first DNA endonuclease), wherein the DNA endonuclease is capable of introducing a single-stranded nick or a double-stranded break, or the sequence-specific DNA-binding protein capable of binding to a site (e.g., a first site) on the target nucleic acid also comprises endonuclease activity capable of introducing a single-stranded nick or a double-stranded break (e.g., a CRISPR-Cas effector protein).

[0090] In some embodiments, the present invention provides a complex (e.g., a first complex) comprising: (a) a sequence-specific DNA-binding protein (e.g., a first sequence-specific DNA-binding protein) that comprises an endonuclease activity capable of binding to a site (e.g., a first site) on a target nucleic acid and introducing a single-stranded nick or a double-stranded break; (b) a first DNA-dependent DNA polymerase; and (c) a DNA-encoded repair template (e.g., a first DNA-encoded repair template).

[0091] In some embodiments, the present invention provides a complex (e.g., a first complex) comprising: (a) a sequence-specific DNA binding protein (e.g., a first sequence-specific DNA binding protein) capable of binding to a site (e.g., a first site) on a target nucleic acid; (b) a DNA-dependent DNA polymerase (e.g., a first DNA-dependent DNA polymerase); (c) a DNA endonuclease (e.g., a first DNA endonuclease); and (d) a DNA-encoded repair template (e.g., a first DNA-encoded repair template).

[0092] In some embodiments, the sequence-specific DNA-binding protein of the complex (e.g., first complex) of the present invention may be derived from a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein. In some embodiments, the sequence-specific DNA-binding protein may be derived from a CRISPR-Cas polypeptide, a zinc finger, a transcription activator-like effector, and / or an Argonaute protein.

[0093] In some embodiments, the DNA endonuclease or DNA endonuclease activity useful in the complexes (e.g., first complexes) of the invention may be, or may be derived from, an endonuclease (e.g., Fok1), a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), and / or a transcription activator-like effector nuclease (TALEN). In some embodiments, the DNA endonuclease may be a nuclease or a nickase, and the DNA endonuclease activity may be a nuclease activity or a nickase activity.

[0094] In some embodiments, the sequence-specific DNA binding protein may be fused to the DNA-dependent DNA polymerase, optionally via a linker. In some embodiments, the sequence-specific DNA binding protein may be fused to the DNA-dependent DNA polymerase at its N-terminus. In some embodiments, the sequence-specific DNA binding protein may be fused to the DNA-dependent DNA polymerase at its C-terminus.

[0095] The present invention further provides engineered (modified) DNA-dependent DNA polymerases fused to affinity polypeptides capable of interacting with peptide tags or RNA recruitment motifs. In some embodiments, the engineered DNA-dependent DNA polymerases of the present invention may comprise a DNA-dependent DNA polymerase fused to a sequence-nonspecific DNA-binding domain, which may optionally be the sequence-nonspecific dsDNA-binding protein Sso7d from Sulfolobus solfataricus. The engineered DNA-dependent DNA polymerases of the present invention may exhibit increased processivity, increased fidelity, increased affinity, increased sequence specificity, decreased sequence specificity, and / or increased cooperativity compared to the same DNA-dependent DNA polymerases that have not been engineered as described herein. In some embodiments, the engineered DNA-dependent DNA polymerase may be modified to reduce or eliminate at least one of 5' to 3'-polymerase activity, 3' to 5' exonuclease activity, 5' to 3' exonuclease activity, and / or 5' to 3' RNA-dependent DNA polymerase activity. Thus, the engineered DNA-dependent DNA polymerase may not comprise at least one of 5' to 3'-polymerase activity, 3' to 5' exonuclease activity, 5' to 3' exonuclease activity, and / or 5' to 3' RNA-dependent DNA polymerase activity.

[0096] In some embodiments, a sequence-specific DNA-binding protein (e.g., a first sequence-specific DNA-binding protein) may be fused to a peptide tag, a DNA-dependent DNA polymerase (e.g., a first DNA-dependent DNA polymerase) may be fused to an affinity polypeptide capable of binding to the peptide tag, and the DNA-dependent DNA polymerase may be recruited to the sequence-specific DNA-binding protein fused to the peptide tag (and to a target nucleic acid to which the sequence-specific DNA-binding protein may be bound). In some embodiments, a DNA-dependent DNA polymerase (e.g., a first sequence-specific DNA-binding protein) may be fused to a peptide tag, and the sequence-specific DNA-binding protein (e.g., a first sequence-specific DNA-binding protein) may be fused to an affinity polypeptide that, by binding to the peptide tag, can recruit the DNA-dependent DNA polymerase to the sequence-specific DNA-binding protein fused to the affinity polypeptide and to the target nucleic acid to which the sequence-specific DNA-binding protein is bound.

[0097] The complexes of the present invention may further comprise a guide nucleic acid (e.g., a CRISPR nucleic acid, crRNA, or crDNA). The guide nucleic acid may be used in combination with a CRISPR-Cas effector protein, which, in some embodiments, may comprise endonuclease or nickase activity. In some embodiments, the endonuclease or nickase activity of the sequence-specific DNA-binding protein may be derived from, for example, a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), and / or a transcription activator-like effector nuclease (TALEN).

[0098] In some embodiments, the guide nucleic acid may be linked to an RNA recruitment motif, and the DNA-dependent DNA polymerase may be fused to an affinity polypeptide that can bind to the RNA recruitment motif. In some embodiments, the RNA recruitment motif may be linked to the 5' end or 3' end of the CRISPR nucleic acid (e.g., recruit crRNA, recruit crDNA).

[0099] In some embodiments, a DNA-encoded repair template may be recruited to a target nucleic acid by ligating the DNA-encoded repair template to a guide nucleic acid that includes a spacer that is complementary to the target nucleic acid.

[0100] The present invention may provide an additional complex (e.g., a second complex), which comprises (a) a sequence-specific DNA-binding protein (e.g., a second sequence-specific DNA-binding protein) capable of binding to a second site on the target nucleic acid, and (b) a DNA-encoded repair template (e.g., a first or second DNA-encoded repair template). In some embodiments, the sequence-specific DNA-binding protein may be derived from a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein. In some embodiments, the complex (e.g., the second complex) may further comprise a DNA endonuclease (e.g., a second DNA endonuclease), which is capable of introducing a single-stranded nick or a double-stranded break into the target nucleic acid. In some embodiments, the DNA endonuclease may be derived from a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), or a transcription activator-like effector nuclease (TALEN). In some embodiments, the sequence-specific DNA-binding protein of the second complex capable of binding to a second site on the target nucleic acid may further comprise an endonuclease activity capable of introducing a single-stranded nick or a double-stranded break into the target nucleic acid. In some embodiments, the sequence-specific DNA-binding protein (e.g., the second sequence-specific DNA-binding protein) further comprising an endonuclease activity may be a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), or a transcription activator-like effector nuclease (TALEN).

[0101] In some embodiments, the DNA-encoded repair template may be linked to a DNA recruitment motif, and the sequence-specific DNA-binding protein may be fused to an affinity polypeptide capable of interacting with the DNA recruitment motif; in some cases, the DNA recruitment motif / affinity polypeptide comprises an HUH-tag, a DNA aptamer, msDNA of a bacterial retron, or an antibody / epitope pair (e.g., T-DNA recruitment). In some embodiments, the sequence-specific DNA-binding protein may be fused to a porcine circovirus 2 (PCV) Rep protein, and the DNA template comprises a PCV recognition site. In some embodiments, the sequence-specific DNA-binding protein may be fused at its N-terminus to the PCV Rep protein. In some embodiments, the sequence-specific DNA-binding protein may be fused at its C-terminus to the PCV Rep protein. Non-limiting examples of HUH-tags and their corresponding recognition sequences that may be useful in the present invention are listed in Table 1.

[0102] JPEG0007765390000001.jpg156170

[0103] In some embodiments, the DNA-encoded repair template may be recruited to the target nucleic acid by incorporating the DNA into a T-DNA sequence that interacts with an Agrobacterium effector protein (e.g., an Agrobacterium virulence polypeptide, optionally virD2 and / or virE2), and the sequence-specific DNA-binding protein may be recruited to the Agrobacterium effector protein, for example, to recruit the DNA-encoded repair template to the sequence-specific DNA-binding protein and to the target nucleic acid to which the sequence-specific DNA-binding protein binds. As an example, one or more epitope tags may be fused to the sequence-specific DNA-binding protein, and an antibody that recognizes the epitope tag may be fused to the Agrobacterium effector protein, thereby allowing the sequence-specific DNA-binding protein and the Agrobacterium effector protein to interact within the plant cell. Any T-DNA sequence associated with an Agrobacterium effector protein will be recruited to the target nucleic acid by the action of a sequence-specific DNA-binding protein.

[0104] In some embodiments, DNA-encoded repair templates can be recruited to target nucleic acid by the attachment of DNA aptamers to the DNA-encoded repair templates.DNA aptamers are DNA sequences that can bind to specific targets with high affinity due to their unique secondary structure.DNA aptamer-guided gene targeting has been demonstrated for endonuclease I-SceI-mediated gene targeting in humans and yeast systems.Pools of candidate DNA aptamers can be screened by capillary electrophoresis for affinity with specific CRISPR proteins (Cas9, Cpf1, etc.).The DNA aptamer with the highest affinity to selected CRISPR nuclease proteins is attached to single-stranded DNA templates, and guides the DNA templates to CRISPR protein target loci.

[0105] In some embodiments, the repair template can be expressed as msDNA from a bacterial retron scaffold attached to a guide RNA. Bacterial retrons are bacterial elements that encode reverse transcriptase, which recognize specific portions of the transcribed retron genome and use them as templates to generate multiple copies of single-stranded DNA (msDNA). The msDNA remains tethered to the RNA template. The retron RNA scaffold sequence can be added to the CRISPR guide RNA scaffold as an extension, with the portion of the retron genome replaced with the repair template desired for gene editing. Expression of the template as msDNA tethered to the guide RNA scaffold extension allows for the simultaneous delivery of multiple copies of the repair template to the cleavage site upon cleavage. This system has been demonstrated in yeast, but not in mammalian or plant systems. Exemplary bacterial retrons useful for the present invention are listed in Table 2.

[0106] JPEG0007765390000002.jpg220170JPEG0007765390000003.jpg234170

[0107] Examples of chimeric guide nucleic acid sequences (guide DNA) designed to induce template-directed editing into the human genome target FANCF01 include: The repair template (bold) embedded within a single guide nucleic acid (sg nucleic acid) (italic lowercase) following the ec67 retron scaffold: JPEG0007765390000004.jpg36170

[0108] The repair template (bold) embedded within a single guide nucleic acid (sg nucleic acid) (italic lowercase) following the ec86 retron scaffold: JPEG0007765390000005.jpg36170

[0109] The repair template (bold) embedded within a single guide nucleic acid (sg nucleic acid) (italic lowercase) following the ec107 retron scaffold: JPEG0007765390000006.jpg35170

[0110] The repair template (bold) embedded within a single guide nucleic acid (sg nucleic acid) (italic lowercase) following the mx162 retron scaffold: Examples include, but are not limited to, JPEG0007765390000007.jpg37170.

[0111] In some embodiments, the complex (eg, the second complex) may further comprise a DNA-dependent DNA polymerase (eg, a second DNA-dependent DNA polymerase).

[0112] In some embodiments, the complex (e.g., the second complex) may further comprise a guide nucleic acid, which in some cases may be linked to a DNA-encoded repair template (e.g., the first or second DNA-encoded repair template).

[0113] In some embodiments, a third complex may be provided, the third complex comprising a sequence-specific DNA binding protein (e.g., a third sequence-specific DNA binding protein) capable of binding to a site (e.g., a third site) on the target nucleic acid that is on a different strand from the first and second sites, and a DNA endonuclease (e.g., a third DNA endonuclease) (e.g., a nickase capable of generating a single-strand break). In some embodiments, contacting the target nucleic acid with the third complex may boost the efficiency of repair by improving mismatch repair.

[0114] In some embodiments, the present invention provides an RNA molecule comprising (a) a nucleic acid sequence that mediates interaction with a CRISPR-Cas effector protein, (b) a nucleic acid sequence that directs the CRISPR-Cas effector protein to a specific nucleic acid target site via DNA-RNA interaction, and (c) a nucleic acid sequence that forms a stem-loop structure (e.g., an RNA recruitment motif) that can interact with an engineered DNA-dependent DNA polymerase of the present invention. In some aspects, the present invention provides an engineered DNA-dependent DNA polymerase of the present invention complexed with an RNA molecule comprising (a) a nucleic acid sequence that mediates interaction with a CRISPR-Cas effector protein, (b) a nucleic acid sequence that directs the CRISPR-Cas effector protein to a specific nucleic acid target site via DNA-RNA interaction, and (c) a nucleic acid sequence that forms a stem-loop structure.

[0115] The present invention further provides polynucleotides that encode a complex of the invention (e.g., a first complex, a second complex, and / or a third complex) and / or encode one or more sequence-specific DNA binding proteins (e.g., a first sequence-specific DNA binding protein, a second sequence-specific DNA binding protein, and / or a third sequence-specific DNA binding protein), a DNA-dependent DNA polymerase (e.g., a first DNA-dependent DNA polymerase and / or a second DNA-dependent DNA polymerase), a DNA endonuclease (e.g., a first DNA endonuclease, a second DNA endonuclease, and / or a third DNA endonuclease), or comprise one or more DNA-encoded repair templates (e.g., a first DNA-encoded repair template and / or a second DNA-encoded repair template), or one or more guide nucleic acids (e.g., a first guide nucleic acid, a second guide nucleic acid, and / or a third guide nucleic acid, etc.). In some embodiments, polynucleotides encoding the engineered DNA-dependent DNA polymerases of the invention are provided. Further provided herein are one or more expression cassettes and / or vectors comprising one or more of the polynucleotides of the invention.

[0116] In some embodiments of the present invention, polynucleotides encoding sequence-specific DNA-binding domains, sequence-nonspecific DNA-binding proteins, DNA endonucleases, DNA-dependent DNA polymerases, and / or expression cassettes and / or vectors comprising the same may be codon-optimized for expression in cells or organisms (e.g., organisms and / or cells of animals (e.g., mammals, insects, fish, etc.), plants (e.g., dicotyledonous plants, monocotyledonous plants), bacteria, archaea, etc.). In some embodiments, expression cassettes comprising polynucleotides of the present invention / encoding complexes of the present invention / polypeptides may be codon-optimized for expression in dicotyledonous plants or monocotyledonous plants.

[0117] The present invention further provides methods of using the compositions of the present invention to modify a target nucleic acid. Accordingly, the present invention provides methods of modifying a target nucleic acid, the methods comprising contacting a target nucleic acid or a cell containing the target nucleic acid with a complex or system of the present invention, a polynucleotide encoding / comprising the same, or one or more components of the complex or system of the present invention, and / or an expression cassette and / or vector comprising the same. The methods may be performed in vivo (e.g., within a cell or an organism) or in vitro (e.g., cell-free). The polypeptides and complexes of the present invention, and the polynucleotides / expression cassettes / vectors encoding them, may be used in methods of modifying a target nucleic acid, for example, in a plant or plant cell, comprising modifying the target nucleic acid in the plant or plant cell by introducing one or more expression cassettes of the present invention into the plant or plant cell to produce a plant or plant cell comprising the modified target nucleic acid. In some embodiments, the methods may further comprise regenerating the plant cell containing the modified target nucleic acid to produce a plant comprising the modified target nucleic acid.

[0118] In some embodiments, a method for modifying a target nucleic acid is provided, the method comprising modifying the target nucleic acid by contacting the target nucleic acid with a complex of the present invention (e.g., a first complex). In some embodiments, the method may further comprise modifying the target nucleic acid by contacting the target nucleic acid with a second complex of the present invention. In some embodiments, the target nucleic acid may further be contacted with a third complex of the present invention, thereby improving the efficiency of repair of the modification of the target nucleic acid.

[0119] In some embodiments, a method for modifying a target nucleic acid is provided, the method comprising modifying the target nucleic acid by contacting the target nucleic acid with (a) a first sequence-specific DNA binding protein capable of binding to a first site on the target nucleic acid, (b) a first DNA-dependent DNA polymerase, (c) a first DNA endonuclease, and (d) a first DNA-encoded repair template. In some embodiments, the first sequence-specific DNA binding protein, the first DNA-dependent DNA polymerase, the first DNA endonuclease, and the first DNA-encoded repair template can form a complex, and the complex can interact with the target nucleic acid.

[0120] In some embodiments, a method for modifying a target nucleic acid is provided, the method comprising modifying the target nucleic acid by contacting the target nucleic acid with (a) a first sequence-specific DNA-binding protein capable of binding to a first site on the target nucleic acid, the first sequence-specific DNA-binding protein comprising a nickase or endonuclease activity capable of introducing a single-stranded nick or a double-stranded break, (b) a first DNA-dependent DNA polymerase, and (c) a first DNA-encoded repair template. In some embodiments, the first sequence-specific DNA-binding protein comprising an endonuclease activity, the first DNA-dependent DNA polymerase, and the first DNA-encoded repair template can form a complex capable of interacting with the target nucleic acid. The endonuclease activity and / or nickase activity of the first sequence-specific DNA-binding protein may be derived from, for example, a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., zinc finger nuclease), and / or a transcription activator-like effector nuclease (TALEN). In some embodiments, the first sequence-specific DNA-binding protein comprising endonuclease activity may be derived from, for example, a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., zinc finger nuclease), and / or a transcription activator-like effector nuclease (TALEN). The first sequence-specific DNA-binding protein may be derived from, for example, a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., zinc finger nuclease), a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein.

[0121] In some embodiments, the first sequence-specific DNA binding protein may be fused to the first DNA-dependent DNA polymerase, optionally via a linker. In some embodiments, the first sequence-specific DNA binding protein may be fused to the first DNA-dependent DNA polymerase at its N-terminus. In some embodiments, the first sequence-specific DNA binding protein may be fused to the first DNA-dependent DNA polymerase at its C-terminus. In some embodiments, the first sequence-specific DNA binding protein may be fused to a peptide tag, and the first DNA-dependent DNA polymerase may be fused to an affinity polypeptide capable of binding to the peptide tag, thereby recruiting the first DNA-dependent DNA polymerase to the first sequence-specific DNA binding protein fused to the peptide tag and to the target nucleic acid to which the sequence-specific DNA binding protein binds and / or can bind. In some embodiments, the first DNA-dependent DNA polymerase may be fused to a peptide tag and the first sequence-specific DNA binding protein may be fused to an affinity polypeptide capable of binding to the peptide tag, whereby the first DNA-dependent DNA polymerase is recruited to the first sequence-specific DNA binding protein fused to the affinity polypeptide and to the target nucleic acid to which the sequence-specific DNA binding protein binds and / or can bind.

[0122] In some embodiments of the present invention, the first sequence-specific DNA-binding domain and / or the first DNA endonuclease may be or may be derived from a CRISPR-Cas effector protein, and the target nucleic acid may be contacted with a guide nucleic acid (e.g., a CRISPR nucleic acid, crRNA, crDNA) (e.g., a first guide nucleic acid) that directs the CRISPR-Cas effector protein to a specific nucleic acid target site via DNA-RNA interactions. In some embodiments, a DNA-encoded repair template (e.g., a first DNA-encoded repair template) may be linked to the guide nucleic acid, thereby guiding the DNA-encoded repair template to the target nucleic acid. In some embodiments, the guide nucleic acid may be linked to an RNA recruitment motif, and a DNA-dependent DNA polymerase (e.g., a first DNA-dependent DNA polymerase) may be fused to an affinity polypeptide capable of binding to the RNA recruitment motif, thereby guiding the DNA-dependent DNA polymerase to the target nucleic acid. The RNA recruitment motif may be linked to the 5' end or to the 3' end of the guide nucleic acid (e.g., recruit crRNA, recruit crDNA).

[0123] In some embodiments, the target nucleic acid contacted with the first complex of the present invention can be contacted with a second complex of the present invention, the second complex comprising (a) a second sequence-specific DNA binding protein capable of binding to a second site on the target nucleic acid, and (b) a DNA-encoded repair template (e.g., a first DNA-encoded repair template or a second DNA-encoded repair template). In some embodiments, the target nucleic acid is further contacted with a second DNA endonuclease, or the second complex further comprises a second DNA endonuclease, which can introduce a single-stranded nick or double-stranded break into the target nucleic acid. Alternatively, or in addition, the second sequence-specific DNA binding protein of the second complex may comprise an endonuclease activity capable of introducing a single-stranded nick or double-stranded break into the target nucleic acid. In some embodiments, a second sequence-specific DNA binding protein capable of binding to a second site on the target nucleic acid, a second DNA-encoded repair template, and, optionally, a DNA endonuclease, can form a complex that interacts with the second site on the target nucleic acid.

[0124] In some embodiments, the second sequence-specific DNA binding protein may be fused to a peptide tag, and the second DNA endonuclease may be fused to an affinity polypeptide capable of binding to the peptide tag, whereby the second DNA endonuclease is recruited to the second sequence-specific DNA binding protein fused to the peptide tag and to a second site on the target nucleic acid to which the second sequence-specific DNA binding protein binds and / or can bind. In some embodiments, the second DNA endonuclease may be fused to a peptide tag, and the second sequence-specific DNA binding protein may be fused to an affinity polypeptide capable of binding to the peptide tag, whereby the second DNA endonuclease is recruited to the second sequence-specific DNA binding protein fused to the affinity polypeptide and to a second site on the target nucleic acid to which the second sequence-specific DNA binding protein binds and / or can bind.

[0125] In some embodiments, the DNA-encoded repair template of the second complex (e.g., the first DNA-encoded repair template or the second DNA-encoded repair template) may be linked to a DNA recruitment motif, and the second sequence-specific DNA-binding protein may be fused to an affinity polypeptide capable of interacting with the DNA recruitment motif; in some cases, the DNA recruitment motif / affinity polypeptide may include an HUH-tag (e.g., see Table 1), a DNA aptamer, msDNA of a bacterial retron, or T-DNA recruitment, thereby recruiting the second DNA-encoded repair template to the sequence-specific DNA-binding protein and a target nucleic acid to which the sequence-specific DNA-binding protein can bind. In some embodiments, the second sequence-specific DNA-binding protein may be fused to, for example, a porcine circovirus 2 (PCV) Rep protein, and the DNA-encoded repair template may include a PCV recognition site.

[0126] In some embodiments, the second sequence-specific DNA-binding protein may be derived from and / or may be a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., zinc finger nuclease), a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein. In some embodiments, the second DNA-binding domain and / or second DNA endonuclease may be derived from and / or may be a CRISPR-Cas effector protein, and the target nucleic acid may be contacted with a guide nucleic acid (e.g., a CRISPR nucleic acid, crRNA, crDNA) (e.g., a second guide nucleic acid) that directs the CRISPR-Cas effector protein to a specific nucleic acid target site via DNA-RNA interactions. In some embodiments, a DNA-encoded repair template (e.g., a second DNA-encoded repair template) may be linked to the guide nucleic acid, thereby guiding the DNA-encoded repair template to the target nucleic acid. In some embodiments, the second guide nucleic acid may be linked to an RNA recruitment motif, and the second DNA endonuclease may be fused to an affinity polypeptide that can bind to the RNA recruitment motif, whereby the guide nucleic acid guides the second DNA endonuclease to the target nucleic acid. The RNA recruitment motif may be linked to the 5' or 3' end of the guide nucleic acid (e.g., recruit crRNA, recruit crDNA).

[0127] In some embodiments, the target nucleic acid contacted with the second complex may further be contacted with a DNA-dependent DNA polymerase (e.g., a second DNA-dependent DNA polymerase). In some embodiments, the DNA-dependent DNA polymerase may be contained within the second complex.

[0128] The method of the present invention may further include contacting the target nucleic acid with a third complex, the third complex comprising a third sequence-specific DNA-binding protein capable of binding to a third site on the target nucleic acid on a different strand from the first site and the second site, wherein the third sequence-specific DNA-binding protein comprises nuclease or nickase activity, thereby improving the efficiency of repair of the modification of the target nucleic acid.

[0129] In some embodiments, the present invention provides a system for modifying a target nucleic acid, comprising a first complex of the present invention, a polynucleotide encoding the same, and / or an expression cassette or vector comprising the polynucleotide, wherein (a) a first sequence-specific DNA-binding protein comprising DNA endonuclease activity binds to a first site on the target nucleic acid, (b) a first DNA-dependent DNA polymerase is capable of interacting with the first sequence-specific DNA-binding protein and is recruited to the first sequence-specific DNA-binding protein and to the first site on the target nucleic acid, and (c) (i) a first DNA-encoded The repair template is linked to a first guide nucleic acid comprising a spacer sequence having substantial complementarity to a first site on the target nucleic acid, thereby guiding the first DNA-encoded repair template to the first site on the target nucleic acid, or (c)(ii) the first DNA-encoded repair template is capable of interacting with a first sequence-specific DNA-binding protein or a first DNA-dependent DNA polymerase and is recruited to the first sequence-specific DNA-binding protein or the first DNA-dependent DNA polymerase and to the first site on the target nucleic acid, thereby modifying the target nucleic acid.

[0130] In some embodiments, a system for modifying a target nucleic acid is provided, the system comprising a first complex of the present invention, a polynucleotide encoding the first complex, and / or an expression cassette or vector comprising the polynucleotide, wherein (a) a first sequence-specific DNA-binding protein binds to a first site on the target nucleic acid, (b) a first DNA endonuclease is capable of interacting with the first sequence-specific DNA-binding protein and / or the guide nucleic acid and is recruited to the first sequence-specific DNA-binding protein and to the first site on the target nucleic acid, and (c) a first DNA-dependent DNA polymerase is capable of interacting with the first sequence-specific DNA-binding protein and / or the guide nucleic acid and is recruited to the first sequence-specific DNA-binding protein and to the first site on the target nucleic acid. (d)(i) the first DNA-encoded repair template is linked to a guide nucleic acid comprising a spacer sequence having substantial complementarity to the first site on the target nucleic acid, thereby guiding the first DNA-encoded repair template to the first site on the target nucleic acid; or (d)(ii) the first DNA-encoded repair template is capable of interacting with a first sequence-specific DNA-binding protein or a first DNA-dependent DNA polymerase and is recruited by the sequence-specific DNA-binding protein or the first DNA-dependent DNA polymerase to the first site on the target nucleic acid, thereby modifying the target nucleic acid.

[0131] In some embodiments, the system of the invention for modifying a target nucleic acid may further comprise a second complex of the invention, a polynucleotide encoding the same, and / or an expression cassette and / or vector comprising the polynucleotide, wherein the second sequence-specific DNA-binding domain binds to a second site proximal to the first site on the target nucleic acid, and the second, DNA-encoded repair template is recruited (via covalent or non-covalent interactions) to the second sequence-specific DNA-binding protein, thereby modifying the target nucleic acid.

[0132] The DNA-dependent DNA polymerases useful in the present invention (e.g., the first and / or second DNA-dependent DNA polymerases) can be any DNA-dependent DNA polymerase. DNA-dependent DNA polymerases are well known in the art, and a non-limiting list can be found on the Polbase website (polbase.neb.com). In some embodiments, the DNA-dependent DNA polymerases useful in the present invention may comprise 3'-5' exonuclease activity, 5'-3' exonuclease activity, and / or 5'-3' RNA-dependent DNA polymerase activity. In some embodiments, the DNA-dependent DNA polymerase may be modified or engineered to eliminate one or more of the 3'-5' exonuclease activity, 5'-3' exonuclease activity, and 5'-3' RNA-dependent DNA polymerase activity.

[0133] In some embodiments, DNA-dependent DNA polymerases (e.g., first and / or second DNA-dependent DNA polymerases) with improved delivery and / or activity may be provided, wherein the DNA-dependent DNA polymerase comprises the Klenow fragment or a subfragment thereof. As an example, the E. coli Klenow fragment may be used, which is approximately 68 kDa in size, or 62% of the molecular weight of full-length (109 kDa) DNA polymerase I.

[0134] DNA-dependent DNA polymerases can be improved in temperature sensitivity, processivity, and template affinity through fusion to a DNA-binding domain. Thus, for example, a DNA-dependent DNA polymerase (e.g., a first and / or second DNA-dependent DNA polymerase) can be fused to a sequence-nonspecific DNA-binding protein to provide a DNA-dependent DNA polymerase with improved temperature sensitivity, processivity, and / or template affinity. In some embodiments, the sequence-nonspecific DNA-binding protein can be a sequence-nonspecific dsDNA-binding protein, such as, but not limited to, Sso7d from Sulfolobus solfataricus.

[0135] The DNA-dependent DNA polymerase (e.g., the first DNA-dependent DNA polymerase and / or the second DNA-dependent DNA polymerase) may be derived from human, yeast, bacteria, or plants. In some embodiments, DNA-dependent DNA polymerases useful in the present invention may include, but are not limited to, DNA polymerase ε (e.g., human and yeast), DNA polymerase δ, E. coli polymerase I, Phusion® DNA polymerase, Vent® DNA polymerase, Vent(exo-)® DNA polymerase, Deep Vent® DNA polymerase, Deep Vent(exo-)® DNA polymerase, 9°Nm™ DNA polymerase, Q5® DNA polymerase, Q5U® DNA polymerase, Pfu DNA polymerase, and / or Phire™ DNA polymerase. In some embodiments, the DNA-dependent DNA polymerase may be human DNA-dependent DNA polymerase ε, plant DNA-dependent DNA polymerase ε, and / or yeast DNA-dependent DNA polymerase ε (see, for example, SEQ ID NOs: 48-58).

[0136] In some embodiments, the DNA-dependent DNA polymerase useful in the present invention can exhibit high fidelity and / or high processivity. Processivity relates to the number of nucleotides incorporated in a single binding event of the polymerase to the template. In some cases, the DNA-dependent DNA polymerase can have a processivity of more than 100 kb (e.g., about 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 kb, or more, and any range or value therein). In some embodiments, the DNA-dependent DNA polymerase can exhibit a high partitioning profile. Thus, the DNA-dependent DNA polymerase can be a high fidelity DNA-dependent DNA polymerase and / or a high processivity DNA-dependent DNA polymerase. In some embodiments, the DNA-dependent DNA polymerase may be a distributive polymerase (eg, a low processivity polymerase) or a DNA-dependent DNA polymerase with a high distributive profile.

[0137] A DNA-dependent DNA polymerase useful in the present invention (eg, the first DNA-dependent DNA polymerase and / or the second DNA-dependent DNA polymerase) can be an engineered DNA-dependent DNA polymerase of the present invention.

[0138] In some embodiments, the sequence-specific DNA binding proteins (e.g., the first sequence-specific DNA binding protein, the second sequence-specific DNA binding protein, and / or the third sequence-specific DNA binding protein) may be derived from a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein. In some embodiments, the sequence-specific DNA binding protein may comprise endonuclease or nickase activity and may be a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), and / or a transcription activator-like effector nuclease (TALEN).

[0139] The DNA endonuclease (e.g., the first DNA endonuclease, the second DNA endonuclease, and / or the third DNA endonuclease) may be a nuclease and / or a nickase (capable of generating a double-stranded break or a single-stranded break, respectively, in a nucleic acid). In some embodiments, the DNA endonuclease (e.g., the first DNA endonuclease, the second DNA endonuclease, and / or the third DNA endonuclease) may be an endonuclease (e.g., Fok1 or other similar endonuclease domain), a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), and / or a transcription activator-like effector nuclease (TALEN).

[0140] In some embodiments, the sequence-specific DNA-binding domain (e.g., a first sequence-specific DNA-binding protein, a second sequence-specific DNA-binding protein, and / or a third sequence-specific DNA-binding protein) and / or the DNA endonuclease (e.g., a first DNA endonuclease, a second DNA endonuclease, and / or a third DNA endonuclease) may be a CRISPR-Cas effector protein, and in some cases, the CRISPR-Cas effector protein may be from a type I CRISPR-Cas system, a type II CRISPR-Cas system, a type III CRISPR-Cas system, a type IV CRISPR-Cas system, a type V CRISPR-Cas system, or a type VI CRISPR-Cas system. In some embodiments, the CRISPR-Cas effector protein of the present invention may be from a type II CRISPR-Cas system or a type V CRISPR-Cas system. In some embodiments, the CRISPR-Cas effector protein may be a type II CRISPR-Cas effector protein, such as a Cas9 effector protein. In some embodiments, the CRISPR-Cas effector protein may be a type V CRISPR-Cas effector protein, such as a Cas12 effector protein.

[0141] Non-limiting examples of CRISPR-Cas effector proteins include Cas9, C2c1, C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csx12), Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, and Cmr In some cases, the CRISPR-Cas effector protein may be a Cas9, Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c effector protein.

[0142] In some embodiments, CRISPR-Cas effector proteins useful in the present invention may contain a mutation within a nuclease active site (e.g., a RuvC, HNH, e.g., a RuvC site in a Cas12a nuclease domain, e.g., a RuvC and / or HNH site in a Cas9 nuclease domain). A CRISPR-Cas effector protein with a mutation within a nuclease active site may have impaired or reduced activity compared to the same CRISPR-Cas effector protein without the mutation. In some embodiments, a mutation within a nuclease active site results in a CRISPR-Cas effector protein with nickase activity (e.g., Cas9n).

[0143] CRISPR Cas9 effector proteins or CRISPR Cas9 effector domains useful in the present invention can be any known or later-identified Cas9 polypeptide. In some embodiments, the CRISPR Cas9 polypeptide can be, for example, a Cas9 polypeptide from Streptococcus spp. (e.g., S. pyogenes, S. thermophilus), Lactobacillus spp., Bifidobacterium spp., Kandleria spp., Leuconostoc spp., Oenococcus spp., Pediococcus spp., Weissella spp., and / or Olsenella spp. (see, e.g., SEQ ID NOS: 59-62).

[0144] Cas12a is a type V Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-Cas nuclease. Cas12a differs from the more well-known type II CRISPR Cas9 nuclease in several respects. For example, Cas9 recognizes a G-rich protospacer adjacent motif (PAM) (3'-NGG) located 3' of its guide RNA (gRNA, sgRNA) binding site (protospacer, target nucleic acid, target DNA), while Cas12a recognizes a T-rich PAM (5'-TTN, 5'-TTTN) located 5' of the target nucleic acid. In fact, the binding orientation of Cas9 and Cas12a to their guide RNAs is largely reversed with respect to their N- and C-termini. Furthermore, the Cas12a enzyme uses a single guide RNA (gRNA, CRISPR array, crRNA) rather than the dual guide RNA (sgRNA (e.g., crRNA and tracrRNA)) found in the native Cas9 system, and Cas12a processes its own gRNA. In addition, Cas12a nuclease activity generates staggered DNA double-strand breaks instead of the blunt ends generated by Cas9 nuclease activity, and Cas12a relies on a single RuvC domain to cleave both DNA strands, whereas Cas9 utilizes an HNH domain and a RuvC domain for cleavage.

[0145] CRISPR Cas12a effector proteins / domains useful in the present invention can be any known or later identified Cas12a polypeptide (formerly known as Cpf1) (see, e.g., U.S. Patent No. 9,790,490, incorporated herein by reference for its disclosure of the Cpf1 (Cas12a) sequence). The terms "Cas12a," "Cas12a polypeptide," or "Cas12a domain" refer to an RNA-guided nuclease comprising a Cas12a polypeptide or a fragment thereof, including the guide nucleic acid binding domain of Cas12a and / or an active, inactive, or partially active DNA cleavage domain of Cas12a. In some embodiments, a Cas12a useful in the present invention may contain a mutation within the nuclease active site (e.g., the RuvC site of a Cas12a domain). A Cas12a domain or Cas12a polypeptide that has a mutation in its nuclease active site and therefore no longer contains nuclease activity is commonly referred to as a deadCas12a (e.g., dCas12a). In some embodiments, a Cas12a domain or Cas12a polypeptide that has a mutation in its nuclease active site may be impaired in activity.

[0146] In some embodiments, peptide tags (e.g., epitopes, peptide repeat units) of the invention useful for recruiting a polypeptide to a selected location (e.g., a target nucleic acid, a site on a target nucleic acid) may comprise one copy or two or more copies of the peptide tag (epitope, multimerized epitope) (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more copies (repeat units)). In some embodiments, peptide tags useful in the present invention may include, but are not limited to, a GCN4 peptide tag (e.g., Sun-tag) (see, e.g., SEQ ID NOs: 23-24), a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or strep tag II, a V5 tag, and / or a VSV-G epitope. In some embodiments, the peptide tag may be a GCN4 peptide tag. In some embodiments, the peptide tag may include two or more copies of the peptide tag (peptide repeats, e.g., two or more tandem copies, e.g., tandem copies of GCN4).

[0147] In some embodiments, the affinity polypeptide capable of binding to a peptide tag includes, but is not limited to, an antibody, and optionally a peptide tag (e.g., a GCN4 peptide tag (see, e.g., SEQ ID NO: 25), a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or strep tag II, a V5 tag, and / or a VSV-G epitope). and / or DARPins, each of which can bind to a peptide tag (e.g., a GCN4 peptide tag, a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or strep tag II, a V5 tag, and / or a VSV-G epitope).

[0148] In some embodiments of the present invention, a guide nucleic acid (CRISPR nucleic acid, crRNA, crDNA) may be linked to one or more RNA recruitment motifs (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 motifs, or more, e.g., at least 10 to about 25 motifs), and in some cases, the two or more RNA recruitment motifs may be the same or different RNA recruitment motifs, such that the guide nucleic acid linked to one or more RNA recruitment motifs can be used to recruit one or more polypeptides fused to affinity polypeptides that can interact / bind with the RNA recruitment motifs linked to the guide.

[0149] In some embodiments, RNA recruitment motifs and affinity polypeptides (e.g., corresponding affinity polypeptides) capable of interacting with the RNA recruitment motif may include, but are not limited to, a telomerase Ku-binding motif (e.g., a Ku-binding hairpin) and a corresponding affinity polypeptide Ku (e.g., a Ku heterodimer), a telomerase Sm7-binding motif and a corresponding affinity polypeptide Sm7, an MS2 phage operator stem-loop and a corresponding affinity polypeptide MS2 coat protein (MCP), a PP7 phage operator stem-loop and a corresponding affinity polypeptide PP7 coat protein (PCP), an SfMu phage Com stem-loop and a corresponding affinity polypeptide Com RNA-binding protein, and / or a synthetic RNA aptamer and a corresponding affinity polypeptide aptamer ligand (e.g., see SEQ ID NOS: 26-36). In some embodiments, an RNA recruitment motif and its corresponding affinity polypeptide useful in the present invention may be an MS2 phage operator stem-loop and an affinity polypeptide MS2 coat protein (MCP), and / or a PUF binding site (PBS) and an affinity polypeptide pumilio / fem-3 mRNA-binding factor (PUF).

[0150] The polypeptides of the present invention described herein may be fusion proteins comprising one or more polypeptides linked to each other. In some embodiments, the fusion is via a linker. In some embodiments, the linker may be an amino acid linker or a peptide linker. In some embodiments, the peptide linker may be about 2 to about 100 amino acids (residues) in length. In some embodiments, the peptide linker may be a GS linker.

[0151] As used herein, "guide nucleic acid," "guide RNA," "gRNA," "CRISPR RNA / DNA," "crRNA," or "crDNA" refers to a nucleic acid sequence comprising target DNA and at least one repeat sequence (e.g., a repeat of a type V Cas12a CRISPR-Cas system, or a fragment or portion thereof, a repeat of a type II Cas9 CRISPR-Cas system, or a fragment thereof, a repeat of a type V C2cl CRISPR Cas system repeats or fragments thereof, such as C2c3, Cas12a (also called Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Casl, CaslB, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csx12), Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cm "gRNA" refers to a nucleic acid comprising at least one spacer sequence (e.g., a protospacer) complementary to (and hybridizing to) a repeat of a CRISPR-Cas system (or a fragment thereof) of a CRISPR-Cas system of type I, type II, type III, type IV, type V, or type VI, where the repeat sequence may be linked to the 5' and / or 3' end of the spacer sequence. The design of gRNAs of the present invention may be based on a Type I, Type II, Type III, Type IV, Type V, or Type VI CRISPR-Cas system.

[0152] In some embodiments, the Cas12a gRNA may comprise, from 5' to 3', a repeat sequence (either full-length or a portion thereof ("handle"), e.g., a pseudoknot-like structure), and a spacer sequence.

[0153] In some embodiments, a guide nucleic acid may contain multiple repeat-spacer sequences (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more repeat-spacer sequences) (e.g., repeat-spacer-repeat, e.g., repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer-repeat-spacer, etc.). Guide nucleic acids of the present invention are synthetic, artificial, and not found in nature. gRNAs can be very long and may be used as aptamers (as in the MS2 recruitment strategy) or other RNA structures with hanging spacers. In some embodiments, the guide RNAs described herein may comprise a template for editing and a primer binding site. In some embodiments, a guide RNA may comprise a region or sequence on the 5' or 3' end that is complementary to the editing template (reverse transcriptase template) and thereby recruits the editing template to the target nucleic acid.

[0154] As used herein, "repeat sequence" refers to, for example, any repeat sequence of a wild-type CRISPR Cas locus (e.g., Cas9 locus, Cas12a locus, C2c1 locus, etc.) or a repeat sequence of a synthetic crRNA functional with a CRISPR-Cas nuclease encoded by a nucleic acid construct of the present invention encoding a base editor. A repeat sequence useful in the present invention can be any known or later identified repeat sequence of a CRISPR-Cas locus (e.g., Type I, Type II, Type III, Type IV, Type V, or Type VI), or a synthetic repeat designed to function in a Type I, II, III, IV, V, or VI CRISPR-Cas system. The repeat sequence may comprise a hairpin structure and / or a stem-loop structure. In some embodiments, the repeat sequence can form a pseudoknot-like structure (i.e., a "handle") at its 5' end. Thus, in some embodiments, the repeat sequence may be identical or substantially identical to a repeat sequence from a wild-type Type I CRISPR-Cas locus, Type II CRISPR-Cas locus, Type III CRISPR-Cas locus, Type IV CRISPR-Cas locus, Type V CRISPR-Cas locus, and / or Type VI CRISPR-Cas locus. Repeat sequences from wild-type CRISPR-Cas loci can be determined by established algorithms, for example, using CRISPRfinder provided by CRISPRdb (see Grissa et al. Nucleic Acids Res. 35 (Web Server Publication): W52-7). In some embodiments, the repeat sequence, or a portion thereof, is linked to the 3'-5' end of a spacer sequence, thereby forming a repeat-spacer sequence (e.g., guide RNA, crRNA).

[0155] In some embodiments, the repeat sequence comprises, consists essentially of, or consists of at least 10 nucleotides (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50-100 or more nucleotides, or any range or value therein, e.g., about), depending on the particular repeat and whether the guide RNA comprising the repeat is processed or unprocessed. In some embodiments, the repeat sequence comprises, consists essentially of, or consists of about 10 to about 20, about 10 to about 30, about 10 to about 45, about 10 to about 50, about 15 to about 30, about 15 to about 40, about 15 to about 45, about 15 to about 50, about 20 to about 30, about 20 to about 40, about 20 to about 50, about 30 to about 40, about 40 to about 80, about 50 to about 100, or more nucleotides.

[0156] The repeat sequence linked to the 5' end of the spacer sequence may comprise a portion of the repeat sequence (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, or more consecutive nucleotides of the wild-type repeat sequence). In some embodiments, the portion of the repeat sequence linked to the 5' end of the spacer sequence can be about 5 to about 10 contiguous nucleotides (e.g., about 5, 6, 7, 8, 9, 10 nucleotides) in length and have at least 90% (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more) identity to the same region (e.g., the 5' end) of a wild-type CRISPR Cas repeat nucleotide sequence. In some embodiments, the portion of the repeat sequence can include a pseudoknot-like structure (e.g., a "handle") at its 5' end.

[0157] As used herein, a "spacer sequence" refers to a nucleotide sequence (e.g., a protospacer) that is complementary to a target nucleic acid (e.g., a target DNA). The spacer sequence can be fully complementary to the target nucleic acid or can be substantially complementary (e.g., at least about 70% (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more) complementary). Thus, in some embodiments, the spacer sequence can have 1, 2, 3, 4, or 5 mismatches compared to the target nucleic acid, and the mismatches can be contiguous or non-contiguous. In some embodiments, the spacer sequence may have 70% complementarity to the target nucleic acid. In other embodiments, the spacer nucleotide sequence may have 80% complementarity to the target nucleic acid. In still other embodiments, the spacer nucleotide sequence may have 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5%, etc., complementarity to the target nucleic acid (protospacer). In some embodiments, the spacer sequence is 100% complementary to the target nucleic acid. The spacer sequence may be about 15 to about 30 nucleotides in length (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides, or any range or value therein). Thus, in some embodiments, a spacer sequence can have perfect or substantial complementarity over a region of the target nucleic acid (e.g., a protospacer) that is at least about 15 to about 30 nucleotides in length. In some embodiments, the spacer is about 20 nucleotides in length. In some embodiments, the spacer is about 23 nucleotides in length.

[0158] In some embodiments, the 5' region of the spacer sequence of the guide RNA may be identical to the target DNA while the 3' region of the spacer may be substantially complementary to the target DNA (e.g., type V CRISPR-Cas), or the 3' region of the spacer sequence of the guide RNA may be identical to the target DNA while the 5' region of the spacer may be substantially complementary to the target DNA (e.g., type II CRISPR-Cas), such that the overall complementarity of the spacer sequence to the target DNA may be less than 100%. Thus, for example, in a guide for a type V CRISPR-Cas system, for example, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in the 5' region of a 20-nucleotide spacer sequence (i.e., the seed region) may be 100% complementary to the target DNA, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary to the target DNA (e.g., at least about 70% complementary). In some embodiments, the first 1 to 8 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8 nucleotides, and any range therein) at the 5' end of the spacer sequence may be 100% complementary to the target DNA, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary to the target DNA (e.g., at least about 50% (e.g., 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) complementary).

[0159] As a further example, in a guide for a Type II CRISPR-Cas system, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides in the 3' region of, e.g., a 20 nucleotide spacer sequence (i.e., the seed region) can be 100% complementary to the target DNA, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary to the target DNA (e.g., at least about 70% complementary). In some embodiments, the first 1 to 10 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides, and any range therein) at the 3' end of the spacer sequence can be 100% complementary to the target DNA, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary to the target DNA (e.g., at least about 50% (e.g., at least about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more, or any range or value therein) complementary).

[0160] In some embodiments, the seed region of the spacer may be about 8 to about 10 nucleotides in length, about 5 to about 6 nucleotides in length, or about 6 nucleotides in length.

[0161] As used herein, the terms "target nucleic acid," "target DNA," "target nucleotide sequence," "target region," or "target region within a genome" refer to a region of an organism's genome that is fully complementary (100% complementary) or substantially complementary (e.g., at least 70% (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) complementary) to a spacer sequence within a guide RNA of the present invention. Useful target regions for CRISPR-Cas systems may be located immediately 3' (e.g., Type V CRISPR-Cas systems) or immediately 5' (e.g., Type II CRISPR-Cas systems) to the PAM sequence in the genome of an organism (e.g., a plant genome, an animal genome, a bacterial genome, a fungal genome, etc.). The target region may be selected from any region of at least 15 contiguous nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides, etc.) located immediately adjacent to the PAM sequence.

[0162] "Protospacer sequence" refers to a target double-stranded DNA, specifically a portion of the target DNA (e.g., or a target region within a genome), that is perfectly or substantially complementary to (and hybridizes to) the spacer sequence of a CRISPR repeat-spacer sequence (e.g., guide RNA, CRISPR array, crRNA).

[0163] In Type V CRISPR-Cas (e.g., Cas12a) and Type II CRISPR-Cas (Cas9) systems, the protospacer sequence is flanked (e.g., immediately adjacent) by a protospacer adjacent motif (PAM). For Type IV CRISPR-Cas systems, the PAM is positioned at the 5' end on the non-target strand and at the 3' end of the target strand (see below for an example). JPEG0007765390000008.jpg27170

[0164] In the case of type II CRISPR-Cas (e.g., Cas9) systems, the PAM is located immediately 3' to the target region. The PAM for type I CRISPR-Cas systems is located 5' to the target strand. The PAM for type III CRISPR-Cas systems is unknown. Makarova et al. describe the nomenclature for all classes, types, and subtypes of CRISPR systems (Nature Reviews Microbiology 13:722-736 (2015)). Guide structures and PAMs are described by R. Barrangou (Genome Biol. 16:247 (2015)).

[0165] The canonical Cas12a PAM is T-rich. In some embodiments, the canonical Cas12a PAM sequence may be 5'-TTN, 5'-TTTN, or 5'-TTTV. In some embodiments, the canonical Cas9 (e.g., Streptococcus pyogenes) PAM may be 5'-NGG-3'. In some embodiments, non-canonical PAMs may be used, but may be less effective.

[0166] Additional PAM sequences can be determined by those skilled in the art using established experimental and computational approaches. Thus, for example, experimental approaches include targeting sequences flanked by all possible nucleotide sequences and identifying sequence members that are not subject to targeting, for example, by transformation of target plasmid DNA (Esvelt et al. 2013. Nat. Methods 10:1116-1121; Jiang et al. 2013. Nat. Biotechnol. 31:233-239). In some embodiments, a computational approach may involve performing a BLAST search of natural spacers to identify the original target DNA sequence within a bacteriophage or plasmid and aligning these sequences to determine conserved sequences flanking the target sequence (Briner and Barrangou. 2014. Appl. Environ. Microbiol. 80:994-1001; Mojica et al. 2009. Microbiology 155:733-740).

[0167] In some embodiments, a nucleic acid construct, expression cassette, or vector of the invention that has been optimized for expression in plants may be about 70% to 100% (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%) identical to a nucleic acid construct, expression cassette, or vector encoding the same but that has not been codon-optimized for expression in plants.

[0168] In some embodiments, the invention provides cells comprising one or more polynucleotides, guide nucleic acids, nucleic acid constructs, expression cassettes, or vectors of the invention.

[0169] When used in combination with a guide nucleic acid, the nucleic acid construct of the present invention can be used to modify a target nucleic acid. The target nucleic acid can be contacted with the nucleic acid construct of the present invention before, simultaneously with, or after contacting the target nucleic acid with the guide nucleic acid. In some embodiments, the nucleic acid construct of the present invention and the guide nucleic acid can be included in the same expression cassette or vector, so the target nucleic acid can be contacted with the nucleic acid construct of the present invention and the guide nucleic acid at the same time. In some embodiments, the nucleic acid construct of the present invention and the guide nucleic acid can be in different expression cassettes or vectors, so the target nucleic acid can be contacted with the nucleic acid construct of the present invention before, simultaneously, or after contacting the guide nucleic acid.

[0170] Target nucleic acids of any organism or cells thereof can be modified (e.g., mutated, e.g., base-edited, truncated, nicked, etc.) using the nucleic acid constructs of the invention (e.g., polypeptides and complexes (e.g., sequence-specific DNA-binding proteins, DNA-dependent DNA polymerases (e.g., engineered DNA-dependent DNA polymerases), DNA endonucleases, DNA-encoded repair templates, guide nucleic acids, etc.), as well as polynucleotides, expression cassettes, and / or vectors encoding the same).

[0171] In some embodiments, target nucleic acids of any plant or plant part can be modified (e.g., mutated, e.g., base-edited, truncated, nicked, etc.) using the nucleic acid constructs of the present invention (e.g., polypeptides and complexes (e.g., sequence-specific DNA-binding proteins, DNA-dependent DNA polymerases (e.g., engineered DNA-dependent DNA polymerases), DNA endonucleases, DNA-encoded repair templates, guide nucleic acids, etc.), and polynucleotides, expression cassettes, and / or vectors encoding the same). Any plant (or group of plants, e.g., genus or higher order classification), including angiosperms, gymnosperms, monocotyledons, dicotyledons, C3, C4, CAM plants, bryophytes, ferns and / or fern relatives, microalgae, and / or macroalgae, can be modified using the nucleic acid constructs of the present invention. Plants and / or plant parts useful in the present invention can be plants and / or plant parts of any plant species / variety / cultivar. The term "plant part" as used herein includes, but is not limited to, embryos, pollen, ovules, seeds, leaves, stems, shoots, inflorescences, branches, fruits, kernels, panicles, cobs, husks, stems, roots, root tips, anthers, and plant cells (including intact plant cells, plant protoplasts, plant tissues, plant cell tissue cultures, plant calli, plant masses, etc. in plants and / or plant parts). As used herein, "shoot" refers to the aboveground part including leaves and stems. Furthermore, as used herein, "plant cell" refers to the structural physiological unit of a plant, which includes the cell wall, and may also refer to a protoplast. Plant cells may be in the form of isolated single cells, cultured cells, or part of a more highly organized unit, such as plant tissue or a plant organ.

[0172] Non-limiting examples of plants useful in the present invention include turfgrasses (e.g., bluegrass, bentgrass, ryegrass, fescue), reed grass, broadleaf grass, Miscanthus, Miscanthus, switchgrass, vegetable crops (e.g., artichoke, kohlrabi, yellow bell pepper, leek, asparagus, lettuce (e.g., salad greens, leaf lettuce, lettuce), malanga, melons (e.g., muskmelon, watermelon, Crenshaw melon, honeydew melon, cantaloupe), Brassica crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, kohlrabi), and corn crops (e.g., cornstarch, cornstarch, cornstarch, cornstarch). kale, Chinese cabbage, bok choy), cardoni, carrots, Chinese cabbage, okra, onion, celery, parsley, chickpeas, parsnip, chicory, pepper, potato, cucurbits (e.g., mallow, cucumber, zucchini, squash, pumpkin, honeydew melon, watermelon, cantaloupe), radish, dry bulb onion, rutabaga, eggplant, burdock, esculenta, shallot, endive, garlic, spinach, leeks, squash, greens, beets (sugar beet and fodder beet), sweet potato, Swiss chard fruit crops such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, quince, figs, nuts (e.g., chestnuts, pecans, pistachios, hazelnuts, peanuts, walnuts, macadamia nuts, almonds, etc.), citrus fruits (e.g., clementines, kumquats, oranges, grapefruit, tangerines, mandarins, lemons, limes, etc.), blueberries, black raspberries, boysenberries, cranberries, cucumbers ... orchids, gooseberries, loganberries, raspberries, strawberries, blackberries, grapes (for wine and for eating), avocados, bananas, kiwi, persimmons, pomegranates, pineapples, tropical fruits, pome fruits, melons, mangoes, papayas, and lychees; agricultural crops such as clover, alfalfa, timothy, evening primrose, meadowfoam, maize (for animal feed, sweet corn, popcorn), hops, jojoba, buckwheat, safflower, quinoa, wheat, rice, barley, rye, millet, sorghum, oats, triticale, sorghum, tobacco;Kapok, legumes (beans (e.g., green beans and dried beans), lentils, peas, soybeans), oil plants (rapeseed, canola, mustard, poppy, olive, sunflower, coconut, castor oil plant, cocoa beans, peanuts, oil palm), duckweed, Arabidopsis, fiber plants (cotton, flax, hemp, jute), cannabis (e.g., Cannabis sativa, Cannabis indica, and Cannabis ruderalis) ruderalis), Lauraceae (cinnamon, camphor), or coffee, sugarcane, tea, and natural rubber plants, and / or bedding plants, e.g., flowering plants, cacti, succulents, and / or ornamental plants (e.g., roses, tulips, violets), as well as trees, e.g., forest trees (broadleaf and evergreen plants, e.g., conifers, e.g., elm, ash, oak, maple, fir, spruce, cedar, pine, birch, cypress, eucalyptus, willow), and shrubs and other seedlings. In some embodiments, the nucleic acid constructs of the invention, and / or expression cassettes and / or vectors encoding same, may be used to modify corn, soybean, wheat, canola, rice, tomato, pepper, sunflower, raspberry, blackberry, black raspberry, and / or cherry. In some embodiments, the nucleic acid constructs of the invention, and / or expression cassettes and / or vectors encoding same, can be used to modify Rubus species (e.g., blackberry, black raspberry, boysenberry, loganberry, raspberry, e.g., caneberry), Vaccinium species (e.g., cranberry), Ribes species (e.g., gooseberry, currant (e.g., red currant, black currant)), or Fragaria species (e.g., strawberry).

[0173] The present invention further includes kits for carrying out the methods of the present invention, which may include reagents, buffers, and equipment for mixing, measuring, sorting, labeling, etc., as well as instructions suitable for modifying target nucleic acids.

[0174] In some embodiments, the present invention provides kits comprising one or more nucleic acid constructs of the present invention and / or expression cassettes and / or vectors comprising same (e.g., comprising or encoding a polypeptide / complex of the present invention), optionally together with instructions for use thereof. In some embodiments, the kits may further comprise a CRISPR-Cas guide nucleic acid (corresponding to a CRISPR-Cas nuclease encoded by a polynucleotide of the present invention) and / or an expression cassette and / or vector comprising same. In some embodiments, the guide nucleic acid may be provided on the same expression cassette and / or vector as the nucleic acid construct of the present invention. In some embodiments, the guide nucleic acid may be provided on a separate expression cassette or vector from that comprising the nucleic acid construct of the present invention.

[0175] In some embodiments, the kit may further comprise a nucleic acid construct encoding the guide nucleic acid, the construct comprising a cloning site for cloning a nucleic acid sequence identical or complementary to the target nucleic acid sequence into the backbone of the guide nucleic acid.

[0176] In some embodiments, the nucleic acid constructs of the present invention, and / or expression cassettes and / or vectors comprising the same, may further encode one or more selectable markers useful for identifying transformants (e.g., nucleic acids encoding antibiotic resistance genes, herbicide resistance genes, etc.).

[0177] The present invention will now be described with reference to the following examples. It should be appreciated that these examples are not intended to limit the scope of the claims to the present invention, but rather are intended to be illustrative of particular embodiments. Any variations in the exemplified methods that occur to those skilled in the art are intended to fall within the scope of the present invention. [Example]

[0178] Example 1: Precise in vivo editing using templates We demonstrate precise template-directed editing in human cells via fusion of a DNA-dependent DNA polymerase to a CRISPR protein by co-transfecting a mixture of components into the human cell line HEK293T. The mixture of components includes a recipient plasmid containing a copy of the mutant EGFP gene driven by a CMV promoter, a single-stranded DNA repair template containing a correction sequence for the mutant EGFP flanked by 100–200 nt of homologous sequences to facilitate template binding to the target site, a second plasmid expressing a fusion protein of a CRISPR protein (e.g., eCas9, nCas9(D10A), or nCas9(H840A)), a DNA-dependent DNA polymerase of interest (e.g., Pol I from Escherichia coli) (the DNA-dependent DNA polymerase is fused to the N- or C-terminus of the CRISPR protein via a linker), and a third plasmid expressing a guide RNA targeting the mutant EGFP sequence. As a control, the second plasmid is replaced with a plasmid expressing only the corresponding CRISPR protein. The desired template-based editing events are identified by flow cytometry, as the mutant EGFP is corrected to a functional copy of EGFP, resulting in a green fluorescent phenotype.

[0179] Alternatively, the third plasmid expressing the DNA repair template and guide RNA (or guide DNA) can be replaced by a plasmid expressing retron reverse transcriptase and chimeric guide RNA (or chimeric guide DNA) along with a retron scaffold containing the repair template.

[0180] Example 2. Precise template-directed in vitro editing via CRISPR proteins and DNA-dependent DNA polymerase. Template-directed precise in vitro editing is mediated by CRISPR proteins and DNA-dependent DNA polymerase. Commercially available DNA-dependent DNA polymerases are evaluated in vitro for their potential to carry out template-directed displacement of target DNA sequences from the nick introduced by CRISPR nickase nCas9 (H840A). Non-limiting examples of DNA-dependent DNA polymerases for evaluation include Q5 high-fidelity DNA polymerase, Phusion® high-fidelity DNA polymerase, Hemo Klen Taq DNA polymerase, Bst2.0 DNA polymerase, Bsu DNA polymerase, Phi29 DNA polymerase, T7 DNA polymerase, Therminator™ DNA polymerase, Klenow fragment (3'→5' exo-), and Vent (exo-).

[0181] A 2 kb DNA fragment containing a Cas9 binding site in the center of the fragment is used as a recipient, and a single-stranded DNA repair template of about 100 nt is used to introduce mismatches into the recipient adjacent to the Cas9 target site.The mixture of recipient DNA, repair template, nCas9 (H840A) protein and guide RNA, and DNA-dependent DNA polymerase is incubated at 37 ° C or 25 ° C.The desired repair product containing mismatches can be digested with T7 endonuclease I, separated from other products, and quantified by gel electrophoresis.

[0182] Example 3. Precise template-directed editing via MS2 RNA loop recruitment of DNA-dependent DNA polymerase to target sites. We demonstrate precise template-directed editing via MS2 RNA loop recruitment of a DNA-dependent DNA polymerase to target sites by co-transfecting a mixture of components into the human cell line HEK293T. The components include a recipient plasmid containing a copy of the mutant EGFP gene driven by a CMV promoter, a single-stranded DNA repair template containing a correction sequence for the mutant EGFP flanked by 100–200 nt of homologous sequences to facilitate template binding to the target site, a second plasmid expressing a DNA-dependent DNA polymerase of interest (e.g., Pol I from Escherichia coli) with an MCP domain fused to its N-terminus via a linker, a third plasmid expressing a guide RNA targeting the mutant EGFP sequence (the guide nucleic acid scaffold has been modified to contain an MS2 stem loop that interacts with the MCP domain), and a fourth plasmid expressing a CRISPR protein (e.g., eCas9, nCas9(D10A), or nCas9(H840A)). As a control, the second plasmid is omitted from the transfection mixture. The desired editing events using the template are identified by flow cytometry, as the mutant EGFP has been corrected to a functional copy of EGFP. Alternatively, the third plasmid expressing the DNA repair template and MS2 guide RNA can be replaced with a plasmid expressing retron reverse transcriptase and a chimeric MS2 guide RNA, along with a retron scaffold containing the repair template.

[0183] Example 4. Precise template-directed editing via PUF-binding site (PBS) RNA aptamer recruitment of DNA-dependent DNA polymerase to target sites. By co-transfecting the component mixture into the human cell line HEK293T, we are able to demonstrate precise template-directed editing via PUF-binding site (PBS) RNA aptamer recruitment of DNA-dependent DNA polymerase to target sites. The component mixture includes a recipient plasmid containing a copy of the mutant EGFP gene driven by a CMV promoter, a single-stranded DNA repair template containing a correction sequence for the mutant EGFP flanked by 100–200 nt of homologous sequences to facilitate template binding to the target site, a second plasmid expressing a DNA-dependent DNA polymerase of interest (e.g., Pol I from E. coli) with a PUF domain fused to its N-terminus via a linker, a third plasmid expressing a guide RNA targeting the mutant EGFP sequence (the guide RNA scaffold has been modified to contain a PUF-binding site that interacts with the PUF domain), and a fourth plasmid expressing a CRISPR protein (e.g., eCas9, nCas9(D10A), or nCas9(H840A)). As a control, the second plasmid is omitted from the transfection mixture. Desired template-mediated editing events are identified by flow cytometry, as mutant EGFP has been corrected to a functional copy of EGFP. Alternatively, a third plasmid expressing a DNA repair template and guide RNA with a PBS can be replaced with a plasmid expressing a retron reverse transcriptase and chimeric guide RNA, along with the PBS and a retron scaffold containing the repair template.

[0184] Example 5. Precise template-directed editing via PUF-binding site (PBS) RNA aptamer recruitment of DNA-dependent DNA polymerase to target sites. We demonstrated precise template-directed editing of a DNA-dependent DNA polymerase at a target site via antibody / epitope cloning by cotransfecting a mixture of components into the human cell line HEK293T. The mixture of components includes a recipient plasmid containing a copy of a mutant EGFP gene driven by a CMV promoter, a single-stranded DNA repair template containing a correction sequence for the mutant EGFP flanked by 100–200 nt of homologous sequences to facilitate template binding to the target site, a second plasmid expressing a DNA-dependent DNA polymerase of interest (e.g., Pol I from E. coli) with an scFV domain fused to its N-terminus via a linker, a third plasmid expressing a guide RNA targeting the mutant EGFP sequence, and a fourth plasmid expressing a CRISPR protein (e.g., eCas9, nCas9(D10A), or nCas9(H840A)) with eight copies of a GCN4 tag fused to its C-terminus. As a control, the second plasmid is omitted from the transfection mixture. The desired editing events using the template are identified by flow cytometry, as the mutant EGFP is corrected to a functional copy of EGFP. Alternatively, the third plasmid expressing the DNA repair template and guide RNA can be replaced with a plasmid expressing a retron reverse transcriptase and a chimeric guide RNA, along with a retron scaffold containing the repair template.

[0185] Example 6. Precise template-directed editing via PUF-binding site (PBS) RNA aptamer recruitment of DNA-dependent DNA polymerase to target sites. By inserting an in-frame EGFP gene (approximately 700 bp) into an exon of a highly expressed gene (e.g., actin), we can demonstrate precise template-directed editing and site-specific integration of long fragments in plants via the recruitment of a DNA-dependent DNA polymerase. In this experimental design, two T-DNAs are co-transformed into plant tissue. The first T-DNA contains a tool cassette expressing a CRISPR protein and a DNA-dependent DNA polymerase in the correct architecture for efficient recruitment of the DNA-dependent DNA polymerase to the target site, and a guide cassette expressing a guide RNA that targets the last exon of actin in the configuration required for protein recruitment. The second T-DNA contains a repair template encoding the full-length EGFP and an in-frame deletion of the stop codon in the target exon. The repair template is flanked by the target site recognized by the guide RNA expressed in the first T-DNA. The desired site-specific integration of EGFP results in EGFP expression driven by the actin gene promoter, whereas random integration results in no EGFP expression due to the lack of a promoter. The frequency of site-specific integration can be quantified by microscopic observation. Additionally, the first T-DNA will express only the tool cassette, while the second T-DNA will contain a retron reverse transcriptase cassette and a chimeric guide RNA cassette encoding a repair template within the retron scaffold attached to the guide RNA scaffold.

[0186] Example 7. Recruitment and optimization of DNA-dependent DNA polymerases. As described in the previous example and more generally herein, DNA-dependent DNA polymerase can be recruited to editing site by many different methods.For example, DNA-dependent DNA polymerase can be fused to the C-terminus or N-terminus of CRISPR protein via flexible linker, for example in the architecture of base editor.Otherwise, DNA-dependent DNA polymerase can be recruited to cut or nicked target DNA through the interaction with guide RNA (for example, MS2 loop) or CRISPR protein (for example, SunTag).

[0187] The function of DNA-dependent DNA polymerases can be improved / optimized in several ways, including, but not limited to, removing 3'-5' exonuclease, 5'-3' exonuclease, and / or 5'-3' RNA-dependent DNA polymerase activity. DNA-dependent DNA polymerases may further comprise the Klenow fragment or other subfragments of the protein. The Klenow fragment or other active fragments may be useful for delivery or activity purposes. As an example, the E. coli Klenow fragment is 68 kDa, or 62% of the molecular weight of the complete (109 kDa) DNA polymerase I.

[0188] Fusing protein domains to DNA-dependent DNA polymerase enzymes can have a significant effect on the temperature sensitivity and processivity of editing systems. DNA-dependent DNA polymerase enzymes can be improved in temperature sensitivity, processivity, and template affinity through fusion to DNA-binding domains (DBDs). These DBDs can be sequence-specific, nonspecific, or sequence-selective. A broad affinity distribution in different cellular and in vitro environments can be beneficial for editing. Adding one or more DBDs to a DNA-dependent DNA polymerase enzyme can increase affinity, increase or decrease sequence specificity, and / or promote cooperativity. One particular DBD known to increase the processivity of DNA-dependent DNA polymerases is the sequence-nonspecific dsDNA-binding protein Sso7d from Sulfolobus solfataricus (Wang, 2004). dsDNA-binding proteins can be fused to the C-terminus, N-terminus, or flexible loop of the polymerase. Increased processivity can be demonstrated by inserting a larger reporter gene, such as tdTomato (approximately 1500 bp), in frame into an exon of a highly expressed gene (e.g., actin). For example, two T-DNAs can be co-transformed into plant tissue. The first T-DNA contains a tool cassette expressing a CRISPR protein and a DNA-dependent DNA polymerase::ssDBD in the correct architecture for efficient recruitment of the DNA-dependent DNA polymerase::ssDBD to the target site, and a guide cassette expressing a guide RNA targeting the last exon of actin in the configuration required for protein recruitment. The second T-DNA contains a repair template encoding the full-length tdTomato (or other reporter) and an in-frame deletion of the stop codon in the target exon. The repair template is flanked by the target site recognized by the guide RNA expressed in the first T-DNA.Site-specific integration of tdTomato (or other reporter) results in expression of tdTomato driven by the promoter of the actin gene, whereas random integration results in no expression of tdTomato due to the lack of a promoter. The frequency of site-specific integration can be quantified by microscopic observation.

[0189] Example 8. CRISPR Polypeptides The present invention utilizes a highly processive DNA-dependent DNA polymerase to rapidly initiate DNA synthesis initiated by the 3' end of a cleaved or nicked target DNA annealed to a provided repair template. Cas9 nuclease and nickase, as well as Cas12a nuclease, and nickases and other CRISPR-Cas effector polypeptides can be used to generate 3' DNA target ends. Successful integration of repair templates, especially large ones, may depend on the ability of the DNA-dependent DNA polymerase to move along the DNA template away from the cleaved or nicked site. Direct fusion to the Cas9 protein, which may remain bound to the cleaved or nicked DNA, may prevent the DNA-dependent DNA polymerase from moving. For this reason, eCas9 (containing three amino acid mutations (K848A, K1003A, R1060A)) 4Cas9 with reduced DNA binding affinity, such as a nuclease or nickase, may also be used. Alternatively, non-covalent recruitment of the polymerase to the CRISPR complex can be used to maximize the polymerase's chances of functioning without steric inhibition or translocation constraints. Several covalent and non-covalent recruitment strategies are described herein. For example, Cpf1 / Cas12a has a longer seed sequence for stable binding (17-bp vs. 9-10-bp for Cas9). This indicates a lower affinity for target DNA (Jeon et al., 2018), consistent with the lower off-target rate of editing observed with Cpf1. The lower affinity of Cpf1 for target DNA compared to Cas9 may be an advantage for polymerase fusions that require translocation of the editing tool.

[0190] [Example 9. Repair template recruitment] In human cell experiments, repair templates can be recruited by several different strategies, including but not limited to: 1) interaction of a PCV domain fused to a CRISPR protein with a PCV recognition site embedded within the repair template, and 2) msDNA encoding the repair template generated from a chimeric retron-guide RNA scaffold and tethered to the guide RNA scaffold.

[0191] [Example 10. Genome editing in plants] In plant editing, various methods of repair template delivery can be used, which can vary depending on the transformation method. For example, for Agrobacterium-mediated plant transformation, VirD2- or VirE2-mediated T-DNA recruitment can be used, or it can be msDNA, and for particle bombardment, the HUH tagging system and msDNA can be used.

[0192] Example 11: Editing in human cells Eukaryotic HEK293T (ATCC CRL-3216) cells were cultured in Dulbecco's modified Eagle's medium plus GlutaMax (ThermoFisher) (FBS) supplemented with 10% (v / v) FBS at 37°C with 5% CO2. HEK293T cells were seeded onto 48-well collagen-coated BioCoat plates (Corning). Cells were transfected at approximately 70% confluency. DNA was transfected using 1.5 μl per well of Lipofectamine 3000 (ThermoFisher Scientific) according to the manufacturer's protocol. RNPs were transfected using 1.5 μl per well of RNAiMAX (ThermoFisher Scientific) according to the manufacturer's protocol. Genomic DNA from transfected cells was obtained after 3 days, and accurate editing was detected and quantified using high-throughput Illumina amplicon sequencing.

[0193] To test DNA polymerase-mediated extension of DNA templates, HEK293T cells were first transfected with 1 μg of DNA encoding various DNA-dependent DNA polymerases under a constitutive CMV promoter, including Klentaq, Therminator, Pfu-Ssod7, Klenow, E. coli pol I, HU pol E (N-term), and yeast pol E (see, e.g., SEQ ID NOS: 48-58, 88-94). All DNA-dependent DNA polymerases were enhanced with at least one SV40 nuclear localization sequence to ensure nuclear import. After 4 hours, cells were replaced with fresh medium. Cas12a RNP complexes (see, e.g., SEQ ID NOS: 75) containing various synthetic crRNA extensions (see, e.g., SEQ ID NOS: 78, 79, 82, 83, 86, and 87) were then transfected into the cells. A DNA extension encoding a homology arm downstream of the Cas12a cleavage site and a template sequence encoding the desired edit were conjugated to crRNA via chemical synthesis (Integrated DNA Technologies). Two different homology lengths (PBS; 24 bp and 36 bp) were tested. The template containing the desired edit (RTT) was 36 base pairs in length (Table 2). The system was tested using three different spacers: PWsp137 (SEQ ID NO: 76), PWsp453 (SEQ ID NO: 80), and PWsp454 (SEQ ID NO: 84) (Table 2). For all constructs, the templates contained a precise dinucleotide change to adenine (TT to AA) at positions -2 and -3 of the spacer. The PAM sequence (TTTV) corresponds to positions -4, -3, -2, and -1. JPEG0007765390000009.jpg253170

[0194] Using DNA polymerase in combination with Cas12a RNPs containing DNA extensions on the crRNA, we detected precise editing without any by-products (Table 2). Precise editing was detected in all three spacers tested (Table 2). Because the indel rate is expected to be effective from LbCas12a RNPs (5-50% editing efficiency in 293T, see, e.g., Liu et al. Nucleic Acids Res. 47(8):4169-4180 (2019)), our low (approximately 1%) indel rate (Table 3) suggests that the two rounds of transfection in our experiments significantly reduced the efficiency of the delivery system. The fact that the precise editing rate in our experiments was similar to the indel rate (Tables 3 and 4) suggests that precise editing via DNA-dependent DNA polymerases is potentially very effective for precise editing. When background levels of precise editing are subtracted and precise editing is normalized to indel editing rates (to normalize for transfection and survival rates), it is clear that the addition of DNA polymerase and template leads to a substantial increase in precise editing compared to the no DNA polymerase control at most spacer sites and PBS lengths (Table 4).

[0195] JPEG0007765390000010.jpg109170

[0196] JPEG0007765390000011.jpg99170

[0197] JPEG0007765390000012.jpg100170

[0198] The foregoing is illustrative of the present invention, and is not to be construed as limiting thereof. The present invention is defined by the following claims, including equivalents of the claims to be included therein. The claims at the time of filing are as follows: [Claim 1] (a) a first sequence-specific DNA binding protein capable of binding to a first site on a target nucleic acid; and (b) a first DNA-dependent DNA polymerase The first complex comprises: [Claim 2] The first complex of claim 1, further comprising a first DNA-encoded repair template. [Claim 3] 3. The first complex of claim 1 or claim 2, further comprising a first DNA endonuclease, wherein the DNA endonuclease is capable of introducing a single-stranded nick or a double-stranded break, or wherein the first sequence-specific DNA binding protein capable of binding to a first site on a target nucleic acid further comprises an endonuclease activity capable of introducing a single-stranded nick or a double-stranded break. [Claim 4] (a) a first sequence-specific DNA-binding protein that is capable of binding to a first site on a target nucleic acid and that contains an endonuclease activity that is capable of introducing a single-stranded nick or a double-stranded break; (b) a first DNA-dependent DNA polymerase, and (c) a first, DNA-encoded repair template; The first complex comprises: [Claim 5] (a) a first sequence-specific DNA binding protein capable of binding to a first site on a target nucleic acid; (b) a first DNA-dependent DNA polymerase; (c) a first DNA endonuclease, and (d) a first, DNA-encoded repair template; The first complex comprises: [Claim 6] 6. The first complex of claim 1, wherein the first sequence-specific DNA-binding protein is derived from a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein. [Claim 7] 7. The first complex of claim 3, wherein the first DNA endonuclease is an endonuclease (e.g., Fok1), a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), and / or a transcription activator-like effector nuclease (TALEN). [Claim 8] The first complex according to claim 3 , wherein the first DNA endonuclease is a nuclease or a nickase. [Claim 9] 9. The first complex of any one of claims 1 to 8, further comprising a guide nucleic acid (e.g., crRNA, crDNA). [Claim 10] 10. The first complex of any one of claims 1, 2, or 6 to 9, wherein the first sequence-specific DNA-binding protein comprises an endonuclease activity or a nickase activity. [Claim 11] 11. The first complex of claim 10, wherein the endonuclease or nickase activity of the first sequence-specific DNA-binding protein is derived from a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), and / or a transcription activator-like effector nuclease (TALEN). [Claim 12] 12. The first complex of claim 1, wherein the first sequence-specific DNA-binding protein is fused (at its N-terminus or C-terminus) to the first DNA-dependent DNA polymerase, optionally via a linker. [Claim 13] 13. The first complex of claim 1, wherein the first DNA-dependent DNA polymerase exhibits high fidelity and / or high processivity. [Claim 14] 13. The first complex of claim 1, wherein the first DNA-dependent DNA polymerase exhibits a high partitioning profile. [Claim 15] 15. The first complex of claim 1, wherein the first DNA-dependent DNA polymerase comprises a 3'-5' exonuclease activity, a 5'-3' exonuclease activity, and / or a 5'-3' RNA-dependent DNA polymerase activity. [Claim 16] 16. The first complex of claim 15, wherein the first DNA-dependent DNA polymerase is modified to eliminate one or more of 3'-5' exonuclease activity, 5'-3' exonuclease activity, and 5'-3' RNA-dependent DNA polymerase activity. [Claim 17] 17. The first complex of claim 1, wherein the first DNA-dependent DNA polymerase comprises the Klenow fragment or a subfragment thereof. [Claim 18] 18. The first complex of claim 1, wherein the first DNA-dependent DNA polymerase is fused to a sequence-nonspecific DNA-binding protein. [Claim 19] The first complex according to claim 18, wherein the sequence-nonspecific dsDNA-binding protein is derived from Sso7d derived from Sulfolobus solfataricus. [Claim 20] 20. The first complex according to claim 1, wherein the first DNA-dependent DNA polymerase is a DNA-dependent DNA polymerase derived from a human, yeast, bacteria, or plant. [Claim 21] 20. The first complex of any one of claims 1 to 19, wherein the first DNA-dependent DNA polymerase is DNA polymerase ε (e.g., human and yeast), DNA polymerase δ, E. coli polymerase I, Phusion® DNA polymerase, Vent® DNA polymerase, Vent(exo-)® DNA polymerase, Deep Vent® DNA polymerase, Deep Vent(exo-)® DNA polymerase, 9°Nm™ DNA polymerase, Q5® DNA polymerase, Q5U® DNA polymerase, Pfu DNA polymerase, and / or Phire™ DNA polymerase. [Claim 22] 21. The first complex according to claim 1, wherein the first DNA polymerase is human DNA-dependent DNA polymerase ε, plant DNA-dependent DNA polymerase ε, and / or yeast DNA-dependent DNA polymerase ε. [Claim 23] 23. The first complex of claim 1, wherein the first DNA-dependent DNA polymerase is a distributive polymerase. [Claim 24] 24. The first complex of claim 1, wherein the first sequence-specific DNA-binding protein is fused to a peptide tag and the first DNA-dependent DNA polymerase is fused to an affinity polypeptide capable of binding to the peptide tag, thereby recruiting the first DNA-dependent DNA polymerase to the first sequence-specific DNA-binding protein fused to the peptide tag. [Claim 25] 24. The first complex of claim 1, wherein the first DNA-dependent DNA polymerase is fused to a peptide tag and the first sequence-specific DNA-binding protein is fused to an affinity polypeptide capable of binding to the peptide tag, thereby recruiting the first DNA-dependent DNA polymerase to the first sequence-specific DNA-binding protein fused to the affinity polypeptide. [Claim 26] 26. The first complex of claim 24 or 25, wherein the peptide tag comprises a GCN4 peptide tag (e.g., a Sun-tag), a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or strep tag II, a V5 tag, and / or a VSV-G epitope. [Claim 27] 27. The first complex according to any one of claims 24 to 26, wherein the affinity polypeptide is an antibody, optionally an scFv antibody, an affibody, an anticalin, a monobody and / or a DARPin. [Claim 28] 24. The first complex of any one of claims 9 to 23, wherein the guide nucleic acid is linked to an RNA recruitment motif and the first DNA-dependent DNA polymerase is fused to an affinity polypeptide capable of binding to the RNA recruitment motif. [Claim 29] 29. The first complex of Claim 28, wherein the RNA recruitment motif is linked to the 5' end or to the 3' end of the CRISPR nucleic acid (e.g., recruit crRNA, recruit crDNA). [Claim 30] 30. The first complex of claim 28 or 29, wherein the RNA recruitment motif and corresponding affinity polypeptide are a telomerase Ku-binding motif (e.g., a Ku-binding hairpin) and a Ku affinity polypeptide (e.g., a Ku heterodimer), a telomerase Sm7-binding motif and an Sm7 affinity polypeptide, an MS2 phage operator stem-loop and affinity polypeptide MS2 coat protein (MCP), a PP7 phage operator stem-loop and affinity polypeptide PP7 coat protein (PCP), an SfMu phage Com stem-loop and affinity polypeptide Com RNA-binding protein, a PUF-binding site (PBS) and an affinity polypeptide pumilio / fem-3 mRNA-binding factor (PUF), and / or a synthetic RNA aptamer and a corresponding aptamer ligand. [Claim 31] 30. The first complex of claim 28 or claim 29, wherein the RNA recruitment motif and corresponding affinity polypeptide are an MS2 phage operator stem-loop and affinity polypeptide MS2 coat protein (MCP), and / or a PUF binding site (PBS) and affinity polypeptide pumilio / fem-3 mRNA binding factor (PUF). [Claim 32] 32. The first complex of any one of claims 9 to 31, wherein the first DNA-encoded repair template is linked to the guide nucleic acid. [Claim 33] 33. The first complex of any one of claims 1 to 32, wherein the first DNA binding domain and / or the first DNA endonuclease is a CRISPR-Cas effector protein. [Claim 34] 34. The first complex of Claim 33, wherein the CRISPR-Cas effector protein is from a Type I CRISPR-Cas system, a Type II CRISPR-Cas system, a Type III CRISPR-Cas system, a Type IV CRISPR-Cas system, or a Type V CRISPR-Cas system. [Claim 35] 34. The first complex of Claim 33, wherein the CRISPR-Cas effector protein is derived from a type II CRISPR-Cas system or a type V CRISPR-Cas system. [Claim 36] 36. The first complex of claim 34 or claim 35, wherein the type V CRISPR-Cas effector protein is Cas12a, Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c. [Claim 37] 35. The first complex of Claims 33 to 34, wherein the CRISPR-Cas effector protein is a Cas9 effector protein or a Cas12 effector protein. [Claim 38] (a) a second sequence-specific DNA-binding protein capable of binding to a second site on the target nucleic acid; and (b) DNA-encoded repair template A second complex comprising: [Claim 39] 39. The second complex of claim 38, wherein the second sequence-specific DNA-binding protein is derived from a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein. [Claim 40] 40. The second complex of claim 39, further comprising a second DNA endonuclease, wherein the second DNA endonuclease is capable of introducing a single-stranded nick or a double-stranded break. [Claim 41] 41. The second complex of claim 40, wherein the second DNA endonuclease is derived from a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), or a transcription activator-like effector nuclease (TALEN). [Claim 42] 40. The second complex of claim 38 or claim 39, wherein the second sequence-specific DNA-binding protein capable of binding to a second site on a target nucleic acid further comprises an endonuclease activity capable of introducing a single-stranded nick or a double-stranded break. [Claim 43] 43. The second complex of claim 42, wherein the second sequence-specific DNA-binding protein further comprising endonuclease activity is a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), or a transcription activator-like effector nuclease (TALEN). [Claim 44] 44. The second complex of any one of claims 38 to 43, further comprising a second DNA-dependent DNA polymerase. [Claim 45] 44. The second complex of any one of claims 38 to 43, wherein the second, DNA-encoded repair template is linked to a DNA recruitment motif, and the second sequence-specific DNA-binding protein is fused to an affinity polypeptide capable of interacting with the DNA recruitment motif, and optionally the DNA recruitment motif / affinity polypeptide comprises an HUH-tag, a DNA aptamer, msDNA of a bacterial retron, or T-DNA recruitment. [Claim 46] 46. ​​The second complex of any one of claims 38 to 45, wherein the second sequence-specific DNA-binding protein is fused (at its N-terminus or C-terminus) to a porcine circovirus 2 (PCV) Rep protein, and the DNA template comprises a PCV recognition site. [Claim 47] 47. The second complex of any one of Claims 38 to 46, wherein the second, DNA-encoded repair template is linked to a guide nucleic acid. [Claim 48] An engineered (modified) DNA-dependent DNA polymerase fused to an affinity polypeptide capable of interacting with a peptide tag or an RNA recruitment motif. [Claim 49] 49. The engineered DNA-dependent DNA polymerase of Claim 48, wherein the DNA-dependent DNA polymerase is fused to a sequence-non-specific DNA-binding domain. [Claim 50] 50. The engineered DNA-dependent DNA polymerase of claim 49, wherein the sequence-nonspecific dsDNA binding protein is derived from Sso7d from Sulfolobus solfataricus. [Claim 51] 51. The engineered DNA-dependent DNA polymerase of any one of Claims 48-50, wherein the DNA-dependent DNA polymerase exhibits increased processivity, increased fidelity, increased affinity, increased sequence specificity, decreased sequence specificity, and / or increased cooperativity [compared to the same DNA-dependent DNA polymerase that is not engineered as described herein]. [Claim 52] 52. The engineered DNA-dependent DNA polymerase of any one of Claims 48 to 51, wherein the DNA-dependent DNA polymerase does not comprise at least one of 5' to 3'-polymerase activity, 3' to 5' exonuclease activity, 5' to 3' exonuclease activity, and / or 5' to 3' RNA-dependent DNA polymerase activity. [Claim 53] A third complex that improves mismatch repair and boosts the efficiency of repair by including a third sequence-specific DNA binding protein capable of binding to a third site on the target nucleic acid that is on a different strand from the first site and the second site, and a third DNA endonuclease. [Claim 54] 38. A polynucleotide encoding the first complex of any one of claims 1 to 37. [Claim 55] 48. A polynucleotide encoding the second complex of any one of claims 38 to 47. [Claim 56] 54. A polynucleotide encoding the third complex of claim 53. [Claim 57] 53. A polynucleotide encoding the engineered DNA-dependent DNA polymerase of any one of claims 48 to 52. [Claim 58] 58. An expression cassette or vector comprising one or more of the polynucleotides of any one of claims 54 to 57. [Claim 59] 53. An RNA molecule comprising: (a) a nucleic acid sequence that mediates interaction with a CRISPR-Cas effector protein; (b) a nucleic acid sequence that directs the CRISPR-Cas effector protein to a specific nucleic acid target site via a DNA-RNA interaction; and (c) a nucleic acid sequence that forms a stem-loop structure that can interact with the engineered DNA-dependent DNA polymerase of any one of claims 48-52. [Claim 60] 38. A method for modifying a target nucleic acid, comprising modifying the target nucleic acid by contacting the target nucleic acid with a first complex of any one of claims 4 to 37. [Claim 61] 61. The method of claim 60, further comprising modifying the target nucleic acid by contacting the target nucleic acid with a second complex of any one of claims 38 to 47. [Claim 62] 62. The method of claim 61, further comprising improving the efficiency of repair of the modification of the target nucleic acid by contacting the target nucleic acid with a third complex of claim 53. [Claim 63] 1. A method for modifying a target nucleic acid, comprising: (a) a first sequence-specific DNA binding protein capable of binding to a first site on a target nucleic acid; (b) a first DNA-dependent DNA polymerase; (c) a first DNA endonuclease, and (d) a first, DNA-encoded repair template; modifying the target nucleic acid by contacting the target nucleic acid with [Claim 64] 64. The method of Claim 63, wherein the first sequence-specific DNA-binding protein, the first DNA-dependent DNA polymerase, the first DNA endonuclease, and the first DNA-encoded repair template form a complex (that interacts with the target nucleic acid). [Claim 65] 65. The method of claim 63 or claim 64, wherein the first sequence-specific DNA-binding protein is derived from a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein. [Claim 66] 66. The method of any one of claims 63 to 65, wherein the first DNA endonuclease is an endonuclease (e.g., Fok1), a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), and / or a transcription activator-like effector nuclease (TALEN). [Claim 67] 67. The method of any one of claims 63 to 66, wherein the first DNA endonuclease is a nuclease or a nickase. [Claim 68] 1. A method for modifying a target nucleic acid, comprising: (a) a first sequence-specific DNA-binding protein that comprises a nickase activity and / or an endonuclease activity capable of binding to a first site on a target nucleic acid and introducing a single-stranded nick or a double-stranded break; (b) a first DNA-dependent DNA polymerase, and (c) a first, DNA-encoded repair template; modifying the target nucleic acid by contacting the target nucleic acid with [Claim 69] 69. The method of Claim 68, wherein the first sequence-specific DNA-binding protein comprising endonuclease activity, the first DNA-dependent DNA polymerase, and the first DNA-encoded repair template form a complex (that interacts with the target nucleic acid). [Claim 70] 70. The method of claim 68 or claim 69, wherein the endonuclease and / or nickase activity of the first sequence-specific DNA-binding protein is derived from a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), and / or a transcription activator-like effector nuclease (TALEN). [Claim 71] 71. The method of any one of Claims 63 to 70, wherein the first sequence-specific DNA-binding protein is fused to the first DNA-dependent DNA polymerase, optionally via a linker. [Claim 72] 71. The method of any one of claims 63 to 70, wherein the first sequence-specific DNA-binding protein is fused to a peptide tag and the first DNA-dependent DNA polymerase is fused to an affinity polypeptide capable of binding to the peptide tag, thereby recruiting the first DNA-dependent DNA polymerase to the first sequence-specific DNA-binding protein fused to the peptide tag. [Claim 73] 71. The method of any one of claims 63 to 70, wherein the first DNA-dependent DNA polymerase is fused to a peptide tag and the first sequence-specific DNA-binding protein is fused to an affinity polypeptide capable of binding to the peptide tag, whereby the first DNA-dependent DNA polymerase is recruited to the first sequence-specific DNA-binding protein fused to the affinity polypeptide. [Claim 74] 74. The method of claim 72 or 73, wherein the peptide tag comprises a GCN4 peptide tag (e.g., Sun-tag), a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or strep tag II, a V5 tag, and / or a VSV-G epitope. [Claim 75] 75. The method of any one of claims 72 to 74, wherein the affinity polypeptide is an antibody, optionally an scFv antibody, an affibody, an anticalin, a monobody, and / or a DARPin. [Claim 76] 76. The method of any one of Claims 63 to 75, wherein the first DNA binding domain and / or first DNA endonuclease is a CRISPR-Cas effector protein. [Claim 77] 77. The method of Claim 76, further comprising contacting the target nucleic acid with a first guide nucleic acid (e.g., crRNA, crDNA). [Claim 78] 78. The method of Claim 77, wherein the first DNA-encoded repair template is linked to the first guide nucleic acid, thereby guiding the first DNA-encoded repair template to the target nucleic acid. [Claim 79] 79. The method of Claim 77 or Claim 78, wherein the first guide nucleic acid is linked to an RNA recruitment motif and the first DNA-dependent DNA polymerase is fused to an affinity polypeptide capable of binding to the RNA recruitment motif, thereby guiding the first DNA-dependent DNA polymerase to the target nucleic acid. [Claim 80] 80. The method of Claim 79, wherein the RNA recruitment motif is linked to the 5' end or to the 3' end of the CRISPR nucleic acid (e.g., recruit crRNA, recruit crDNA). [Claim 81] 81. The method of claim 79 or 80, wherein the RNA recruitment motif and corresponding affinity polypeptide are a telomerase Ku-binding motif (e.g., a Ku-binding hairpin) and a Ku affinity polypeptide (e.g., a Ku heterodimer), a telomerase Sm7-binding motif and an Sm7 affinity polypeptide, an MS2 phage operator stem-loop and affinity polypeptide MS2 coat protein (MCP), a PP7 phage operator stem-loop and affinity polypeptide PP7 coat protein (PCP), an SfMu phage Com stem-loop and affinity polypeptide Com RNA-binding protein, a PUF-binding site (PBS) and an affinity polypeptide pumilio / fem-3 mRNA-binding factor (PUF), and / or a synthetic RNA aptamer and a corresponding aptamer ligand. [Claim 82] 81. The method of claim 79 or claim 80, wherein the RNA recruitment motif and corresponding affinity polypeptide are an MS2 phage operator stem-loop and affinity polypeptide MS2 coat protein (MCP), and / or a PUF binding site (PBS) and affinity polypeptide pumilio / fem-3 mRNA binding factor (PUF). [Claim 83] 83. The method of any one of Claims 76 to 82, wherein the CRISPR-Cas effector protein is from a Type I CRISPR-Cas system, a Type II CRISPR-Cas system, a Type III CRISPR-Cas system, a Type IV CRISPR-Cas system, or a Type V CRISPR-Cas system. [Claim 84] 83. The method of any one of Claims 76 to 82, wherein the CRISPR-Cas effector protein is from a type II CRISPR-Cas system or a type V CRISPR-Cas system. [Claim 85] 85. The method of claim 84, wherein the type V CRISPR-Cas effector protein is Cas12a, Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c1, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b, and / or Cas14c. [Claim 86] 85. The method of any one of Claims 76 to 84, wherein the CRISPR-Cas effector protein is a Cas9 effector protein or a Cas12 effector protein. [Claim 87] The target nucleic acid (a) a second sequence-specific DNA binding protein capable of binding to a second site on the target nucleic acid; and (b) DNA-encoded repair template 87. The method of any one of Claims 63 to 86, further comprising contacting with a second complex comprising: [Claim 88] 88. The method of Claim 87, wherein the target nucleic acid is further contacted with a second DNA endonuclease, wherein the second DNA endonuclease is capable of introducing a single-stranded nick or a double-stranded break, or wherein the second sequence-specific DNA binding protein comprises an endonuclease activity capable of introducing a single-stranded nick or a double-stranded break. [Claim 89] 89. The method of Claim 88, wherein the second DNA endonuclease is derived from a polynucleotide-guided endonuclease, a CRISPR-Cas endonuclease (e.g., a CRISPR-Cas effector protein), a protein-guided endonuclease (e.g., a zinc finger nuclease), or a transcription activator-like effector nuclease (TALEN), or wherein the second sequence-specific DNA-binding protein comprising endonuclease activity is a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), or a transcription activator-like effector nuclease (TALEN). [Claim 90] 90. The method of any one of Claims 87 to 89, wherein the second sequence-specific DNA binding protein capable of binding to a second site on the target nucleic acid, the second DNA-encoded repair template, and optionally the DNA endonuclease form a complex that interacts with the second site on the target nucleic acid. [Claim 91] 91. The method of any one of claims 87 to 90, wherein the second sequence-specific DNA binding protein is fused to a peptide tag and the second DNA endonuclease is fused to an affinity polypeptide capable of binding to the peptide tag, whereby the second DNA endonuclease is recruited to the second sequence-specific DNA binding protein fused to the peptide tag. [Claim 92] 92. The method of any one of claims 87 to 91, wherein the second DNA endonuclease is fused to a peptide tag and the second sequence-specific DNA binding protein is fused to an affinity polypeptide capable of binding to the peptide tag, whereby the second DNA endonuclease is recruited to the second sequence-specific DNA binding protein fused to the affinity polypeptide. [Claim 93] 93. The method of claim 87 or claim 92, wherein the peptide tag comprises a GCN4 peptide tag (e.g., Sun-tag), a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or strep tag II, a V5 tag, and / or a VSV-G epitope. [Claim 94] 94. The method of any one of claims 89 to 93, wherein the affinity polypeptide is an antibody, optionally an scFv antibody, an affibody, an anticalin, a monobody, and / or a DARPin. [Claim 95] 95. The method of any one of claims 87 to 94, wherein the second, DNA-encoded repair template is linked to a DNA recruitment motif, and the second sequence-specific DNA-binding protein is fused to an affinity polypeptide capable of interacting with the DNA recruitment motif, and optionally the DNA recruitment motif / affinity polypeptide comprises an HUH-tag, a DNA aptamer, msDNA of a bacterial retron, or T-DNA recruitment. [Claim 96] 96. The method of any one of claims 87 to 95, wherein the second sequence-specific DNA-binding protein is fused to a porcine circovirus 2 (PCV) Rep protein and the repair template encoded by the DNA contains a PCV recognition site. [Claim 97] 97. The method of any one of Claims 89 to 96, wherein the second sequence-specific DNA-binding protein is derived from a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease (e.g., a zinc finger nuclease), a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein. [Claim 98] 98. The method of any one of Claims 89 to 97, wherein the second DNA binding domain and / or second DNA endonuclease is a CRISPR-Cas effector protein. [Claim 99] 99. The method of Claim 98, further comprising contacting the target nucleic acid with a second guide nucleic acid (e.g., crRNA, crDNA). [Claim 100] 100. The method of Claim 99, wherein the second, DNA-encoded repair template is linked to the second guide nucleic acid, thereby guiding the second, DNA-encoded repair template to the target nucleic acid. [Claim 101] 101. The method of Claim 99 or Claim 100, wherein the second guide nucleic acid is linked to an RNA recruitment motif and the second DNA endonuclease is fused to an affinity polypeptide capable of binding to the RNA recruitment motif, thereby guiding the second DNA endonuclease to the target nucleic acid. [Claim 102] 102. The method of Claim 101, wherein the RNA recruitment motif is linked to the 5' end or to the 3' end of the guide nucleic acid (e.g., recruit crRNA, recruit crDNA). [Claim 103] The method of claim 101 or claim 102, wherein the RNA recruitment motif and corresponding affinity polypeptide are a telomerase Ku-binding motif (e.g., a Ku-binding hairpin) and a Ku affinity polypeptide (e.g., a Ku heterodimer), a telomerase Sm7-binding motif and an Sm7 affinity polypeptide, an MS2 phage operator stem-loop and affinity polypeptide MS2 coat protein (MCP), a PP7 phage operator stem-loop and affinity polypeptide PP7 coat protein (PCP), an SfMu phage Com stem-loop and affinity polypeptide Com RNA-binding protein, a PUF-binding site (PBS) and an affinity polypeptide pumilio / fem-3 mRNA-binding factor (PUF), and / or a synthetic RNA aptamer and a corresponding aptamer ligand. [Claim 104] 103. The method of claim 101 or claim 102, wherein the RNA recruitment motif and corresponding affinity polypeptide are an MS2 phage operator stem-loop and affinity polypeptide MS2 coat protein (MCP), and / or a PUF binding site (PBS) and affinity polypeptide pumilio / fem-3 mRNA binding factor (PUF). [Claim 105] 105. The method of any one of Claims 87 to 104, further comprising contacting the target nucleic acid with a second DNA-dependent DNA polymerase. [Claim 106] 106. The method of any one of Claims 87 to 105, comprising contacting the target nucleic acid with a third complex, wherein the third complex comprises a third sequence-specific DNA binding protein capable of binding to a third site on the target nucleic acid that is on a different strand from the first site and the second site, and wherein the third sequence-specific DNA binding protein comprises nuclease activity or nickase activity, thereby improving efficiency of repair of the modification of the target nucleic acid. [Claim 107] 107. The method of any one of claims 63 to 106, wherein the first DNA-dependent DNA polymerase and / or the second DNA-dependent DNA polymerase has been modified to eliminate one or more of 3'-5' exonuclease activity, 5'-3' exonuclease activity, and 5'-3' RNA-dependent DNA polymerase activity. [Claim 108] 108. The method of any one of claims 63 to 107, wherein the first DNA-dependent DNA polymerase and / or the second DNA-dependent DNA polymerase comprises the Klenow fragment or a subfragment thereof. [Claim 109] 109. The method of any one of claims 63 to 108, wherein the first DNA-dependent DNA polymerase and / or the second DNA-dependent DNA polymerase is fused to a sequence-non-specific DNA-binding protein. [Claim 110] The method of claim 109, wherein the sequence-nonspecific dsDNA binding protein is derived from Sso7d from Sulfolobus solfataricus. [Claim 111] 111. The method of any one of claims 63 to 110, wherein the first DNA-dependent DNA polymerase and / or the second DNA-dependent DNA polymerase is a DNA-dependent DNA polymerase of human, yeast, bacterial, or plant origin. [Claim 112] 112. The method of any one of claims 63 to 111, wherein the first DNA-dependent DNA polymerase and / or the second DNA-dependent DNA polymerase is DNA polymerase ε (e.g., human and yeast), DNA polymerase δ, E. coli polymerase I, Phusion® DNA polymerase, Vent® DNA polymerase, Vent(exo-)® DNA polymerase, Deep Vent® DNA polymerase, Deep Vent(exo-)® DNA polymerase, 9°Nm™ DNA polymerase, Q5® DNA polymerase, Q5U® DNA polymerase, Pfu DNA polymerase, and / or Phire™ DNA polymerase. [Claim 113] 112. The method of any one of claims 63 to 111, wherein the first DNA-dependent DNA polymerase and / or the second DNA-dependent DNA polymerase is human DNA-dependent DNA polymerase ε, plant DNA-dependent DNA polymerase ε, and / or yeast DNA-dependent DNA polymerase ε. [Claim 114] 114. The method of any one of claims 63 to 113, wherein the first DNA-dependent DNA polymerase and / or the second DNA-dependent DNA polymerase exhibit high fidelity and / or high processivity. [Claim 115] 114. The method of any one of claims 63 to 113, wherein the first DNA-dependent DNA polymerase and / or the second DNA-dependent DNA polymerase exhibits a high partitioning profile. [Claim 116] 116. The method of any one of claims 63 to 115, wherein the first DNA-dependent DNA polymerase and / or the second DNA-dependent DNA polymerase is an engineered DNA-dependent DNA polymerase of any one of claims 48 to 52. [Claim 117] 10. A system for modifying a target nucleic acid, comprising the first complex according to any one of claims 4, 6 to 9 or 12 to 37, a polynucleotide encoding the first complex, and / or an expression cassette or vector comprising said polynucleotide, (a) the first sequence-specific DNA binding protein comprising DNA endonuclease activity binds to a first site on the target nucleic acid; (b) the first DNA-dependent DNA polymerase is capable of interacting with the first sequence-specific DNA-binding protein and is recruited to the first sequence-specific DNA-binding protein and to the first site on the target nucleic acid; (c)(i) the first DNA-encoded repair template is linked to a first guide nucleic acid that includes a spacer sequence that has substantial complementarity to the first site on the target nucleic acid, thereby guiding the first DNA-encoded repair template to the first site on the target nucleic acid; or (c)(ii) the first DNA-encoded repair template is capable of interacting with the first sequence-specific DNA-binding protein or the first DNA-dependent DNA polymerase and is recruited to the first sequence-specific DNA-binding protein or the first DNA-dependent DNA polymerase and to the first site on the target nucleic acid, thereby modifying the target nucleic acid. [Claim 118] A system for modifying a target nucleic acid, comprising the first complex according to any one of claims 5 to 9 or 12 to 37, a polynucleotide encoding the first complex, and / or an expression cassette or vector comprising said polynucleotide, (a) the first sequence-specific DNA binding protein binds to a first site on the target nucleic acid; (b) the first DNA endonuclease is capable of interacting with the first sequence-specific DNA binding protein and / or the guide nucleic acid and is recruited to the first sequence-specific DNA binding protein and to the first site on the target nucleic acid; (c) the first DNA-dependent DNA polymerase is capable of interacting with the first sequence-specific DNA-binding protein and / or the guide nucleic acid and is recruited to the first sequence-specific DNA-binding protein and to the first site on the target nucleic acid; (d)(i) the first DNA-encoded repair template is linked to a guide nucleic acid that includes a spacer sequence that has substantial complementarity to the first site on the target nucleic acid, thereby guiding the first DNA-encoded repair template to the first site on the target nucleic acid; or (d)(ii) the first DNA-encoded repair template is capable of interacting with the first sequence-specific DNA-binding protein or the first DNA-dependent DNA polymerase and is recruited to the sequence-specific DNA-binding protein or the first DNA-dependent DNA polymerase and to the first site on the target nucleic acid, thereby modifying the target nucleic acid. [Claim 119] 119. The system for modifying a target nucleic acid of claim 117 or 118, further comprising the second complex of any one of claims 38 to 47, a polynucleotide encoding the same, and / or an expression cassette and / or vector comprising said polynucleotide, wherein said second sequence-specific DNA-binding domain binds to a second site on the target nucleic acid proximal to said first site, and wherein said second, DNA-encoded repair template is recruited to said second sequence-specific DNA-binding protein (via covalent or non-covalent interactions), thereby modifying the target nucleic acid. [Claim 120] 1. A method for modifying a target nucleic acid, comprising: The target nucleic acid Modifying the target nucleic acid by contacting it with a complex of any one of claims 117 to 119. A method comprising:

Claims

1. (a) a first sequence-specific DNA-binding protein capable of binding to a first site on a target nucleic acid, wherein the first sequence-specific DNA-binding protein is derived from a CRISPR-Cas effector protein; and (b) a first DNA-dependent DNA polymerase; and (c) an RNA:DNA hybrid guide comprising a first DNA-encoded repair template, wherein the first DNA-encoded repair template is linked to an RNA guide nucleic acid, thereby guiding the first DNA-encoded repair template to the target nucleic acid; A first complex comprising: A first complex, wherein the first complex further comprises a first DNA endonuclease, wherein the first DNA endonuclease is capable of introducing a single-stranded nick or a double-stranded break, or wherein the first sequence-specific DNA binding protein further comprises an endonuclease activity capable of introducing a single-stranded nick or a double-stranded break.

2. The first complex described in claim 1, wherein the first sequence-specific DNA binding protein comprises an endonuclease activity capable of introducing a single-stranded nick or a double-stranded break.

3. A first complex described in claim 1 or claim 2, wherein the first complex comprises the first DNA endonuclease.

4. The first complex described in claim 3, wherein the first DNA endonuclease is an endonuclease, a polynucleotide-guided endonuclease, a CRISPR-Cas effector protein, a protein-guided endonuclease, and / or a transcription activator-like effector nuclease (TALEN).

5. A first complex described in any one of claims 1, 2, or 4, wherein the first sequence-specific DNA-binding protein is fused to or recruited to the first DNA-dependent DNA polymerase.

6. the first DNA-dependent DNA polymerase i) exhibit high fidelity and / or high processivity, or a high distribution profile; ii) 3'-5' exonuclease activity, 5'-3' exonuclease activity, and / or 5'-3' RNA-dependent DNA polymerase activity, optionally wherein said first DNA-dependent DNA polymerase is modified to remove one or more of the 3'-5' exonuclease activity, 5'-3' exonuclease activity, and 5'-3' RNA-dependent DNA polymerase activity; iii) comprises the Klenow fragment or a subfragment thereof; iv) fused to a sequence-nonspecific DNA binding protein, optionally wherein the sequence-nonspecific dsDNA binding protein is derived from Sso7d of Sulfolobus solfataricus; v) a DNA-dependent DNA polymerase of human, yeast, bacterial, or plant origin, or the first DNA-dependent DNA polymerase is DNA polymerase ε, DNA polymerase δ, E. coli polymerase I, Phusion® DNA polymerase, Vent® DNA polymerase, Vent(exo-)® DNA polymerase, Deep Vent® DNA polymerase, Deep Vent(exo-)® DNA polymerase, 9°Nm™ DNA polymerase, Q5® DNA polymerase, Q5U® DNA polymerase, Pfu DNA polymerase, and / or Phire™ DNA polymerase; vi) a distributive polymerase; vii) fused to an affinity polypeptide for binding to a peptide tag, wherein the first sequence-specific DNA-binding protein is fused to the peptide tag, whereby the first DNA-dependent DNA polymerase is recruited to the first sequence-specific DNA-binding protein fused to the peptide tag; optionally, a) the peptide tag comprises a GCN4 peptide tag, a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or a strep tag II, a V5 tag, and / or a VSV-G epitope; and / or b) the affinity polypeptide is an antibody, optionally an scFv antibody, an affibody, an anticalin, a monobody, and / or a DARPin; viii) any combination of i) to vii) above; The first complex according to any one of claims 1 to 5.

7. A polynucleotide encoding the first complex of any one of claims 1 to 6.

8. 1. A method of modifying a target nucleic acid, comprising modifying the target nucleic acid by contacting the target nucleic acid with a first complex, wherein the first complex comprises: (a) a first sequence-specific DNA-binding protein capable of binding to a first site on a target nucleic acid, wherein the first sequence-specific DNA-binding protein is derived from a CRISPR-Cas effector protein; (b) a first DNA-dependent DNA polymerase; and (c) an RNA:DNA hybrid guide comprising a first DNA-encoded repair template, wherein the first DNA-encoded repair template is linked to an RNA guide nucleic acid, thereby guiding the first DNA-encoded repair template to the target nucleic acid; A method comprising: The method of claim 1, wherein the first complex further comprises a first DNA endonuclease, and the first DNA endonuclease is capable of introducing a single-stranded nick or a double-stranded break, or the first sequence-specific DNA binding protein further comprises an endonuclease activity capable of introducing a single-stranded nick or a double-stranded break.

9. 9. The method of claim 8, wherein the first sequence-specific DNA binding protein comprises an endonuclease or nickase activity, or the first complex comprises a first DNA endonuclease.

10. 10. The method of claim 9, wherein the first sequence-specific DNA-binding protein, the first DNA-dependent DNA polymerase, the first DNA endonuclease, and the first DNA-encoded repair template form a complex.

11. 9. The method of claim 8, wherein the first sequence-specific DNA-binding protein comprising endonuclease activity, the first DNA-dependent DNA polymerase, and the first DNA-encoded repair template form a complex.

12. the first sequence-specific DNA binding protein i) fused to said first DNA-dependent DNA polymerase, optionally via a linker; ii) fused to a peptide tag, wherein the first DNA-dependent DNA polymerase is fused to an affinity polypeptide capable of binding to the peptide tag, whereby the first DNA-dependent DNA polymerase is recruited to the first sequence-specific DNA-binding protein fused to the peptide tag; and optionally a) the peptide tag comprises a GCN4 peptide tag, a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or a strep tag II, a V5 tag, and / or a VSV-G epitope; and / or b) the affinity polypeptide is an antibody, optionally an scFv antibody, an affibody, an anticalin, a monobody, and / or a DARPin; or iii) fused to an affinity polypeptide capable of binding to a peptide tag, wherein said first DNA-dependent DNA polymerase is fused to said peptide tag, thereby recruiting said first sequence-specific DNA-binding protein to said first DNA-dependent DNA polymerase fused to said peptide tag; and optionally a) the peptide tag comprises a GCN4 peptide tag, a c-Myc affinity tag, an HA affinity tag, a His affinity tag, an S affinity tag, a methionine-His affinity tag, an RGD-His affinity tag, a FLAG octapeptide, a strep tag or a strep tag II, a V5 tag, and / or a VSV-G epitope; and / or b) the affinity polypeptide is an antibody, optionally an scFv antibody, an affibody, an anticalin, a monobody, and / or a DARPin; The method of claim 10.

13. The method of any one of claims 9 to 12, wherein the first DNA endonuclease is a CRISPR-Cas effector protein.

14. the RNA guide nucleic acid is linked to an RNA recruitment motif, and the first DNA-dependent DNA polymerase is fused to an affinity polypeptide capable of binding to the RNA recruitment motif, thereby guiding the first DNA-dependent DNA polymerase to the target nucleic acid; and optionally, The method of any one of claims 8 to 13, wherein the RNA recruitment motif is linked to the 5' end or to the 3' end of the RNA guide nucleic acid.

15. the RNA recruitment motif and the corresponding affinity polypeptide are a telomerase Ku binding motif and Ku affinity polypeptide, a telomerase Sm7 binding motif and Sm7 affinity polypeptide, an MS2 phage operator stem-loop and affinity polypeptide MS2 coat protein (MCP), a PP7 phage operator stem-loop and affinity polypeptide PP7 coat protein (PCP), an SfMu phage Com stem-loop and affinity polypeptide Com RNA binding protein, a PUF binding site (PBS) and an affinity polypeptide pumilio / fem-3 mRNA binding factor (PUF), and / or a synthetic RNA aptamer and a corresponding aptamer ligand; and / or 15. The method of claim 14, wherein the RNA recruitment motif and corresponding affinity polypeptide are an MS2 phage operator stem-loop and affinity polypeptide MS2 coat protein (MCP), and / or a PUF binding site (PBS) and affinity polypeptide pumilio / fem-3 mRNA binding factor (PUF).

16. i) the CRISPR-Cas effector protein is derived from a type II CRISPR-Cas system or a type V CRISPR-Cas system; ii) the CRISPR-Cas effector protein is a Cas9 effector protein or a Cas12 effector protein; iii) the method further comprises: (a) a second sequence-specific DNA binding protein capable of binding to a second site on the target nucleic acid; and (b) a second, DNA-encoded repair template; and optionally further comprising contacting the second complex comprising: [i] the target nucleic acid is further contacted with a second DNA endonuclease, wherein the second DNA endonuclease is capable of introducing a single-stranded nick or a double-stranded break, or the second sequence-specific DNA binding protein comprises an endonuclease activity capable of introducing a single-stranded nick or a double-stranded break; and / or [ii] further comprising contacting the target nucleic acid with a second DNA-dependent DNA polymerase; iv) the method further comprises contacting the target nucleic acid with a third complex, the third complex comprising a third sequence-specific DNA-binding protein capable of binding to a third site on the target nucleic acid on a different strand from the first site and the second site, the third sequence-specific DNA-binding protein comprising nuclease or nickase activity, thereby improving the efficiency of repair of the modification of the target nucleic acid; or v) Any combination of i) to iv) above 16. The method according to any one of claims 8 to 15, wherein

Citation Information

Patent Citations

  • Improved nucleic acid modifying enzyme

    JP2003534796A

  • Evolved cas9 protein for gene editing

    JP2018537963A