Methods and compositions for promoting targeted genome modification using HUH endonucleases
By tethering a Cas12a nuclease to a HUH endonuclease with a novel linker, the method enhances the efficiency and precision of targeted sequence integration in eukaryotic genomes, addressing the limitations of existing CRISPR nucleases in site-specific incorporation and stability.
Patent Information
- Application Number
- JP2022506632
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-08-02
- Filing Date
- 2020-07-31
- Publication Date
- 2025-11-06
- Estimated Expiration
- 2040-07-31
AI Technical Summary
Current CRISPR nucleases, such as Cas9 and CasX, face challenges in achieving efficient and precise targeted integration of desired sequences into genomes, particularly in eukaryotic cells, due to limitations in site-specific incorporation and stability.
The use of a novel linker to tether a Cas12a nuclease to a HUH endonuclease, forming a recombinant polypeptide that enhances the site-specific incorporation of a template sequence into a target DNA molecule by leveraging the HUH endonuclease's ability to form covalent bonds with single-stranded DNA.
This approach significantly improves the efficiency and precision of targeted integration, enabling effective editing and integration of exogenous sequences into eukaryotic genomes, including plants, through mechanisms like non-homologous end-joining and homologous recombination.
Smart Images

Figure 0007765377000010 
Figure 0007765377000011 
Figure 0007765377000012
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS AND INCORPORATION OF SEQUENCE LISTINGS This application claims the benefit of U.S. Provisional Application No. 62 / 882,266, filed August 2, 2019, which is incorporated herein by reference in its entirety. The Sequence Listing contained in the file entitled "P34742WO00_SL.TXT" is 136,904 bytes (measured in MS-Windows), was created on July 31, 2020, was filed electronically hereby, and is incorporated herein by reference in its entirety.
[0002] Field The present disclosure relates to compositions and methods involving the use of RNA-guided nucleases linked to HUH endonuclease to improve targeted integration of desired sequences into genomes.
[0003] Incorporation of sequence listings The sequence listing contained in file entitled 46-21-63706_US0001_SEQ is 133 kilobytes (measured in MS-Windows), was created on July 17, 2020, contains 28 sequences, is filed electronically hereby, and is incorporated herein by reference in its entirety. [Background technology]
[0004] background CRISPR (clustered regularly interspaced short palindromic repeats) nucleases (e.g., Cas12a, CasX, Cas9) are proteins guided by guide RNA to target nucleic acid molecules, and these nucleases can cleave single or double strands of the target nucleic acid molecule. HUH (histidine-hydrophobic amino acid-histidine) endonucleases are nucleases containing an HUH tag that can form a covalent bond with a specific single-stranded DNA sequence.
[0005] This disclosure demonstrates that a Cas12a nuclease can be tethered to a HUH endonuclease using novel linker amino acids to improve site-specific incorporation of a template sequence into a target DNA molecule. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] International Publication No. 2016 / 205711 [Patent Document 2] International Publication No. 2019 / 084148 [Patent Document 3] U.S. Patent No. 5,550,318 [Patent Document 4] U.S. Patent No. 5,538,880 [Patent Document 5] U.S. Patent No. 6,160,208 [Patent Document 6] U.S. Patent No. 6,399,861 [Patent Document 7] U.S. Patent No. 6,153,812 [Patent Document 8] U.S. Patent No. 5,159,135 [Patent Document 9] U.S. Patent No. 5,824,877 [Patent Document 10] U.S. Patent No. 5,591,616 [Patent Document 11] U.S. Patent No. 6,384,301 [Patent Document 12] U.S. Patent No. 5,750,871 [Patent Document 13] U.S. Patent No. 5,463,174 [Patent Document 14] U.S. Patent No. 5,188,958 [Patent Document 15] U.S. Patent No. 5,049,386 [Patent Document 16] U.S. Patent No. 4,946,787 [Patent Document 17] U.S. Patent No. 4,897,355 [Patent Document 18] International Publication No. 91 / 17424 [Patent Document 19] International Publication No. 91 / 16024 [Patent Document 20] International Publication No. 2014 / 093622 [Patent Document 21] U.S. Patent No. 6,194,636 [Patent Document 22] U.S. Patent No. 6,232,526 [Patent Document 23] U.S. Patent Application Publication No. 2004 / 0216189 [Patent Document 24] U.S. Patent No. 6,437,217 [Patent Document 25] U.S. Patent No. 5,641,876 [Patent Document 26] U.S. Patent No. 6,426,446 [Patent Document 27] U.S. Patent No. 6,429,362 [Patent Document 28] U.S. Patent No. 6,232,526 [Patent Document 29] U.S. Patent No. 6,177,611 [Patent Document 30] U.S. Patent No. 5,322,938 [Patent Document 31] U.S. Patent No. 5,352,605 [Patent Document 32] U.S. Patent No. 5,359,142 [Patent Document 33] U.S. Patent No. 5,530,196 [Patent Document 34] U.S. Patent No. 6,433,252 [Patent Document 35] U.S. Patent No. 6,429,357 [Patent Document 36] U.S. Patent No. 5,837,848 [Patent Document 37] U.S. Patent No. 6,294,714 [Patent Document 38] U.S. Patent No. 6,140,078 [Patent Document 39] U.S. Patent No. 6,252,138 [Patent Document 40] U.S. Patent No. 6,175,060 [Patent Document 41] U.S. Patent No. 6,635,806 [Patent Document 42] U.S. Patent Application No. 09 / 757,089 [Patent Document 43] U.S. Patent No. 6,051,753 [Patent Document 44] U.S. Patent No. 5,378,619 [Patent Document 45] U.S. Patent No. 5,850,019 [Patent Document 46] U.S. Patent No. 5,106,739 [Non-patent literature]
[0007] [Non-Patent Document 1] “The American Heritage(R)Science Dictionary” (Editors of the American Heritage Dictionaries, 2011, Houghton Mifflin Harcourt, Boston and New York) [Non-patent document 2] "McGraw-Hill Dictionary of Scientific and Technical Terms" (6th edition, 2002, McGraw-Hill, New York), [Non-patent document 3] "Oxford Dictionary of Biology" (6th edition, 2008, Oxford University Press, Oxford and New York) [Non-patent document 4] Green and Sambrook, Molecular Cloning: A Laboratory Manual, 4th ed. (2012) [Non-Patent Document 5] Current Protocols in Molecular Biology (eds. FM Ausubel et al., (1987)) [Non-patent document 6] Plant Breeding Methodology (NF Jensen, Wiley-Interscience (1988)) [Non-Patent Document 7] the series Methods In Enzymology (Academic Press, Inc.): PCR 2: A Practical Approach (edited by MJ MacPherson, BD Hames and GR Taylor (1995)) [Non-patent document 8] Harlow and Lane (eds.) (1988) Antibodies, A Laboratory Manual [Non-Patent Document 9] Animal Cell Culture (edited by RI Freshney (1987)) [Non-Patent Document 10] Recombinant Protein Purification: Principles And Methods, 18-1142-75, GE Healthcare Life Sciences [Non-Patent Document 11] Edited by CN Stewart, A. Touraev V. Citovsky, and T. Tzfira (2011) Plant Transformation Technologies (Wiley-Blackwell) [Non-Patent Document 12] RH Smith (2013) Plant Tissue Culture: Techniques and Experiments (Academic Press, Inc.) [Non-Patent Document 13] Compendium of Transgenic Crop Plants (2009) Blackwell Publishing [Non-Patent Document 14] Kam et al. (2004) / . Am. Chem. Soc., 126 (22):6850-6851 [Non-Patent Document 15] Liu et al. (2009) Nano Lett, 9(3): 1007-1010 [Non-Patent Document 16] Khodakovskaya et al. (2009) ACS Nano, 3(10):3221-3227 [Non-Patent Document 17] PCR Primer: A Laboratory Manual, edited by Dieffenbach and Dveksler, Cold Spring Harbor Laboratory Press, 1995. [Non-Patent Document 18] Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY) [Non-Patent Document 19] Chenna R. et al., “Multiple sequence alignment with the Clustal series of programs”, Nucleic Acids Research 31 : 3497-3500 (2003) [Non-Patent Document 20] Thompson JD et al., “Clustal W: Improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice,” Nucleic Acids Research 22: 4673-4680 (1994) [Non-Patent Document 21] Larkin MA et al. “Clustal W and Clustal X version 2.0”, Bioinformatics 23: 2947-48 (2007) [Non-Patent Document 22] Altschul, SF, Gish, W., Miller, W., Myers, EW and Lipman, DJ (1990) "Basic local alignment search tool" J. Mol. Biol. 215:403-410 (1990) [Non-Patent Document 23] Sambrook, J. and Russell, W., Molecular Cloning: A Laboratory Manual, 3rd edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (2001) [Non-Patent Document 24] Zhang and Madden, Genome Res., 1997, 7, 649-656 [Non-Patent Document 25] Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489) [Non-Patent Document 26] Schramm and Hernandez, 2002, Genes & Development, 16:2593-2620 [Non-Patent Document 27] Lawton et al., Plant Molecular Biology (1987) 9: 315-324 [Non-patent document 28] Odell et al., Nature (1985) 313: 810-812 [Non-Patent Document 29] Yang and Russell, Proceedings of the National Academy of Sciences, USA (1990) 87: 4144-4148 [Non-Patent Document 30] Chandler et al., Plant Cell (1989) 1: 1175-1183 [Non-Patent Document 31] Depicker et al., Journal of Molecular and Applied Genetics (1982) 1: 561-573 [Non-Patent Document 32] Nakamura et al., 2000, Nucl. Acids Res. 28:292 [Non-Patent Document 33] Campbell and Gowri, 1990, Plant Physiol., 92: 1-11 [Non-Patent Document 34] Murray et al. 1989. Nucleic Acids Res. 17:477-98 [Non-Patent Document 35] Vega-Rocha et al., 2007, Biochemistry 46: 6201 [Non-Patent Document 36] Timchenko et al., 1999, J. of Virol. 73: 10173 [Non-Patent Document 37] Dostal et al., 2011, Nucleic Acids Res 39: 2658 [Brief explanation of the drawings]
[0008] [Figure 1]Figure 1 includes panels 1A and 1B. Figure 1A shows a graphical representation of an RNA-guided nuclease (Cas12a; equivalent to Cas12a) and a histidine-hydrophobic-histidine endonuclease (HUH EN) tethered to a single-stranded DNA (ssDNA) template containing an HUH recognition sequence (ori) for enhanced cleavage (scissors) of double-stranded (ds) chromosomal DNA. Figure 1B shows a graphical representation of a C-terminal cys-free LbCas12a:HUH fusion expression construct and an N-terminal HUH:cys-free LbCas12a fusion expression construct. [Figure 2] Figure 2 shows quantification of targeted indels by next-generation sequencing at three target sites. Cys-free LbCas12a:HUH fusion and cys-free LbCas12a had comparable activity. Bars represent the average indel rate of four technical replicates. Both the experimental (cys-free LbCas12a:HUH / crRNA / template) and positive control (cys-free LbCas12a / crRNA / template) showed a significantly higher percentage of indels compared to the negative control (crRNA / template) at p=0.05. "t" refers to template. [Figure 3] Figure 3A shows a schematic of the test and control systems for testing fusion proteins in plants. The control system lacks the on sequence required for ssDNA template binding to HUH endonuclease. Figure 3B shows the results of two independent particle bombardment experiments using the test and control systems shown in Figure 3A. Chromosomal cleavage activity of Cas12a was determined by quantifying targeted indels within the seven-nucleotide target site. Plants that generated at least 20% targeted indel reads were considered mutants. [Figure 4]Figure 4 shows the genomic region flanking the GmTS1 target site and a matched exogenous template added to the ribonucleoprotein complex. The distinguishable middle signature motif of the template is shaded. Upon breakage of the double-stranded chromosome, the template can integrate by non-homologous end-joining (NHEJ) or homologous recombination (HR). After sequencing, the presence / absence of the middle signature motif indicates targeted integration. The nature of the integration, e.g., NHEJ, HR, or a combination thereof, is indicated by the sequences at the 5' and 3' chromosome-template junctions, as shown. Not shown in this illustration is the ori recognition sequence that was present at the 5' end of the template used in the test sample. The ori recognition sequence was absent in the negative control. Black triangles indicate PCR primers. [Figure 5] FIG. 5 shows the results of in planta testing of cys-free LbCas12a:HUH fusion proteins for targeted integration in R0 plants. [Figure 6] Figure 6 shows the expected editing results when RNP complexes containing N- or C-terminal HUH fusion proteins, crRNA targeting the GmTS1 genomic site, and templates containing either single-stranded DNA templates (sstemplates) or double-stranded DNA oligos (dsOligos) are simultaneously delivered into soybean protoplasts. [Figure 7] Figure 7 shows chromosome breakage by N-terminal (HUH:cys-free LbCas12a) and C-terminal (cys-free LbCas12a:HUH) fusion derivatives of LbCas12a across 16 treatments (Tr) containing various combinations of enzyme, cognate crRNA, ss template with or without ori sequence, and dsOligo with DNA sequences unrelated to the target region. See Table 7 for the reagent combinations in each protoplast treatment. Chromosome breakage was quantified by the targeted indel rate. Bars represent the mean of four biological replicates, and error bars are standard deviations. Statistical significance between several significant treatments is indicated by horizontal lines above the bars. [Figure 8]Figure 8 shows targeted integration of templates using N-terminal (HUHcys-free LbCas12a) and C-terminal (cys-free LbCas12a:HUH) fusion derivatives of LbCas12a. Bars represent the mean of four biological replicates, and error bars represent standard deviations. Total integration by NHEJ (total of targeted single-copy and multiple-copy integrations; black bars), single-copy integration by NHEJ (dark gray bars), and template editing by HR (light gray bars) were all significantly better when the template was tethered to an RNP complex. See Table 7 for the reagent combinations in each treatment. Statistical significance between some key treatments is indicated by horizontal lines above the bars. [Figure 9] Figure 9 shows targeted integration of exogenous untethered dsDNA oligonucleotides (90 bp) by NHEJ using N-terminal (HUH:cys-free LbCas12a) and C-terminal (cys-free LbCas12a:HUH) fusion derivatives of LbCas12a. Bars represent the mean of four biological replicates, and error bars represent standard deviations. "dsDNA oligo, total" indicates the sum of targeted single-copy and multiple-copy integration by NHEJ (black bars). "dsDNA oligo, 1x" indicates single-copy integration by NHEJ (dark gray bars). See Table 7 for the reagent combinations in each treatment. Statistical significance between several significant treatments is indicated by horizontal lines above the bars. Summary of the Invention
[0009] overview In one aspect, the present disclosure provides (a) a recombinant polypeptide comprising: (i) an amino acid sequence encoding a Cas12a nuclease; (ii) an amino acid sequence encoding a linker; and (iii) an amino acid sequence encoding a HUH nuclease; and (b) a ribonucleoprotein comprising at least one guide nucleic acid.
[0010] In one aspect, the disclosure provides a recombinant nucleic acid comprising: (a) a first nucleic acid sequence encoding a Cas12a nuclease; (b) a second nucleic acid sequence encoding a linker; and (c) a third nucleic acid sequence encoding a HUH nuclease.
[0011] In one aspect, the disclosure provides a method of generating an edit in a target DNA molecule, the method comprising contacting the target DNA molecule with a ribonucleoprotein, the ribonucleoprotein comprising: (a) a recombinant polypeptide comprising: (i) an amino acid sequence encoding a Cas12a nuclease; (ii) an amino acid sequence encoding a linker; and (iii) an amino acid sequence encoding a HUH nuclease; (b) at least one guide nucleic acid; and (c) at least one template nucleic acid molecule, wherein the ribonucleoprotein generates at least one edit in the target DNA molecule.
[0012] In one aspect, the disclosure provides a method for generating an edit in a target DNA molecule, the method comprising providing to a cell: (a) a recombinant polypeptide, or one or more nucleic acid molecules encoding the recombinant polypeptide, comprising: (i) an amino acid sequence encoding a Cas12a nuclease; (ii) an amino acid sequence encoding a linker; and (iii) an amino acid sequence encoding a HUH nuclease; (b) at least one guide nucleic acid, or at least one nucleic acid molecule encoding at least one guide nucleic acid; and (c) at least one template nucleic acid molecule, or at least one nucleic acid molecule encoding at least one template nucleic acid molecule; wherein the recombinant polypeptide, the at least one guide nucleic acid, and the at least one template nucleic acid molecule form a ribonucleoprotein, and the ribonucleoprotein generates at least one edit in the target DNA molecule in the cell. DETAILED DESCRIPTION OF THE INVENTION
[0013] Detailed Description Unless otherwise defined, all technical and scientific terms used have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Where a term is provided in the singular, the inventors also intend aspects of the disclosure to be described by the plural of that term. Where there are discrepancies in terms and definitions used in references incorporated by reference, the terms used in this application shall have the definitions given herein. Other technical terms used have their ordinary meaning in the art in which they are used, as exemplified by various art-specific dictionaries, such as "The American Heritage® Science Dictionary" (Editors of the American Heritage Dictionaries, 2011, Houghton Mifflin Harcourt, Boston and New York), "McGraw-Hill Dictionary of Scientific and Technical Terms" (6th ed., 2002, McGraw-Hill, New York), or "Oxford Dictionary of Biology" (6th ed., 2008, Oxford University Press, Oxford and New York). The inventors do not intend to be limited by mechanism or mode of action. That reference is provided for illustrative purposes only.
[0014] The practice of the present disclosure will involve, unless otherwise indicated, conventional techniques of biochemistry, chemistry, molecular biology, microbiology, cell biology, plant biology, genomics, biotechnology and genetics, which are within the skill of the art. See, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); Current Protocols In Molecular Biology (eds. F.M. Ausubel et al., (1987)); Plant Breeding Methodology (N.F. Jensen, Wiley-Interscience (1988)); the series Methods In Enzymology (Academic Press, Inc.): PCR 2: A Practical Approach (eds. M.J. MacPherson, B.D. Hames and G.R. Taylor, (1995)); Harlow and Lane, (1988) Antibodies, A Laboratory Manual; Animal Cell Culture (eds. R.I. Freshney, (1987)); Recombinant Protein Purification: Principles And Methods, 18-1142-75, GE Healthcare Life Sciences; C.N. Stewart, A. Touraev, V. Citovsky, T. Tzfira, (2011) Plant Transformation Technologies (Wiley-Blackwell); and R.H. Smith (2013) Plant Tissue Culture: Techniques and Experiments (Academic Press, Inc.).
[0015] All references cited herein, including all patents, published patent applications, and non-patent publications, are hereby incorporated by reference in their entirety.
[0016] Where a grouping of options is presented, any and all combinations of the members comprising the grouping of options are specifically contemplated. For example, if an item is selected from the group consisting of A, B, C, and D, we specifically contemplate each option individually (e.g., A alone, B alone, etc.), as well as combinations of A, B and D; A and C; B and C, etc.
[0017] As used herein, the singular and singular terms "a," "an," and "the," for example, include plural referents unless the content clearly dictates otherwise.
[0018] Any of the compositions, nucleic acid molecules, polypeptides, cells, plants, etc. provided herein are specifically contemplated for use with any of the methods provided herein.
[0019] In one aspect, the present disclosure provides a ribonucleoprotein comprising: (a) a recombinant polypeptide comprising: (i) an amino acid sequence encoding a Cas12a nuclease; (ii) an amino acid sequence encoding a linker; and (iii) an amino acid sequence encoding a HUH endonuclease; and (b) at least one guide nucleic acid. In one aspect, the ribonucleoprotein further comprises at least one template nucleic acid molecule.
[0020] In one aspect, the present disclosure provides a ribonucleoprotein comprising: (a) a recombinant polypeptide comprising: (i) an amino acid sequence encoding a CasX nuclease; (ii) an amino acid sequence encoding a linker; and (iii) an amino acid sequence encoding a HUH endonuclease; and (b) at least one guide nucleic acid. In one aspect, the ribonucleoprotein further comprises at least one template nucleic acid molecule.
[0021] In one aspect, the present disclosure provides a recombinant nucleic acid comprising: (a) a first nucleic acid sequence encoding a Cas12a nuclease; (b) a second nucleic acid sequence encoding a linker; and (c) a third nucleic acid sequence encoding a HUH endonuclease. In one aspect, the present disclosure provides a recombinant nucleic acid comprising: (a) a first nucleic acid sequence encoding a CasX nuclease; (b) a second nucleic acid sequence encoding a linker; and (c) a third nucleic acid sequence encoding a HUH endonuclease. In one aspect, the recombinant nucleic acid further comprises: (d) a fourth nucleic acid sequence encoding at least one guide nucleic acid. In one aspect, the recombinant nucleic acid further comprises: (d) a fourth nucleic acid sequence encoding at least one template nucleic acid molecule. In another aspect, the recombinant nucleic acid further comprises: (d) a fourth nucleic acid sequence encoding at least one guide nucleic acid; and (e) a fifth nucleic acid sequence encoding at least one template nucleic acid molecule.
[0022] CRISPR (clustered regularly interspaced short palindromic repeats) nucleases (e.g., Cas9, CasX, Cas12a (also called Cpf1), CasY) are proteins found in bacteria that are guided by a guide RNA ("gRNA") to a target nucleic acid molecule; the endonuclease can then cleave one or both strands of the target nucleic acid molecule. Although CRISPR nucleases originate from bacteria, many CRISPR nucleases have been shown to function in eukaryotic cells.
[0023] Without being limited to a particular scientific theory, CRISPR nucleases form a complex with guide RNA (gRNA) and hybridize to a complementary target site, thereby guiding the CRISPR nuclease to the target site. In Class II CRISPR-Cas systems, a CRISPR array containing a spacer is transcribed upon encountering recognized invasive DNA and processed into a small interfering CRISPR RNA (crRNA). The crRNA contains a repeat sequence and a spacer sequence that is complementary to a specific protospacer sequence in the invading pathogen. The spacer sequence can be designed to be complementary to a target sequence in the eukaryotic genome.
[0024] CRISPR nucleases associate with their respective crRNAs in their active forms. CasX, like the class II endonuclease Cas9, requires another non-coding RNA component called a trans-activating crRNA (tracrRNA) to be functionally active. The nucleic acid molecules provided herein can combine crRNA and tracrRNA into a single nucleic acid molecule, referred to herein as a "single guide RNA" (sgRNA). Cas12a does not require tracrRNA to guide it to the target site; crRNA alone is sufficient for Cas12a. The gRNA guides the active CRISPR nuclease complex to the target site, allowing the CRISPR nuclease to cleave the target site.
[0025] When RNA-guided CRISPR nuclease and guide RNA form a complex, the whole system is called "ribonucleoprotein". The ribonucleoprotein provided herein can also include additional nucleic acids, such as, but not limited to, template nucleic acid molecules. The ribonucleoprotein provided herein can also include additional proteins, such as linkers and HUH endonucleases.
[0026] A prerequisite for cleavage of a target site by a CRISPR ribonucleoprotein is the presence of a conserved protospacer adjacent motif (PAM) near the target site. Depending on the CRISPR nuclease, cleavage can occur within a certain number of nucleotides from the PAM site (e.g., 18-23 nucleotides for Cas12a). The PAM site is only required for type I and type II CRISPR-associated proteins, and different CRISPR endonucleases recognize different PAM sites. Cas12a can recognize at least the following PAM sites: TTTN and YTN; and CasX can also recognize at least the following PAM sites: TTCN, TTCA, and TTC (where T is thymine; C is cytosine; A is adenine; Y is thymine or cytosine; and N is thymine, cytosine, guanine, or adenine).
[0027] Cas12a is an RNA-guided nuclease in the class II, type V CRISPR / Cas system. Cas12a nuclease produces staggered cuts when cleaving target nucleic acid molecules.
[0028] In one aspect, the Cas12a nuclease provided herein is a Lachnospiraceae bacterial Cas12a (LbCas12a) nuclease. In another aspect, the Cas12a nuclease provided herein is a Francisella novicida Cas12a (FnCas12a) nuclease. In some embodiments, the amino acid sequence of the Cas12a nuclease has been engineered to remove a cysteine.
[0029] In one aspect, the Cas12a nuclease or a nucleic acid encoding the Cas12a nuclease comprises: Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus Lactobacillus genus, Eubacterium genus, Corynebacter genus, Carnobacterium genus, Rhodobacter genus, Listeria genus, Paludibacter genus, Clostridium genus, Lachnospiraceae genus, Clostridiaridium genus, Leptotrichia ichia genus, Francisella genus, Legionella genus, Alicyclobacillus genus, Methanomethyophilus genus, Porphyromonas genus, Prevotella genus, Bacteroidetes genus, Helcococcus genus, Letospira genus, Desulfovibrio genus, Desulfovibrio genus, The genus Desulfonatronum, the genus Opitutaceae, the genus Tuberibacillus, the genus Bacillus, the genus Brevibacilus, the genus Methylobacterium, the genus Acidaminococcus, the genus Peregrinibacteria, the genus Butyrivibrio, the genus Parcubacteria,It is derived from a bacterial genus selected from the group consisting of Smithella, Candidatus, Moraxella, and Leptospira.
[0030] In one embodiment, the Cas12a nuclease provided herein comprises an amino acid sequence at least 80% identical or similar to the amino acid sequence disclosed in SEQ ID NO: 1, which encodes cys-free LbCas12a. As used herein, "cys-free LbCas12a" refers to an LbCas12a protein variant in which all nine cysteines present in the native LbCas12a sequence (WO 2016 / 205711-1150) have been mutated. In one aspect, the cys-free LbCas12a contains the following nine amino acid substitutions compared to the wt LbCas12a protein sequence: C10L, C175L, C565S, C632L, C805A, C912V, C965S, C1090P, and C1116L. Cysteine residues in the protein can form disulfide bridges, providing strong, reversible bonds between cysteines. To control and direct the binding of Cas12a in a targeted manner, natural cysteines are removed to control the formation of these crosslinks. Without wishing to be bound by any particular theory, removing cysteines from the protein backbone allows for the targeted insertion of new cysteine residues to control the placement of these reversible disulfide linkages. This can be between protein domains or on particles such as gold particles for biolistic delivery. Adding a tag containing several cysteine residues to cys-free LbCas12a allows it to specifically bind to metal beads (especially gold) in a uniform manner.
[0031] In one aspect, a Cas12a nuclease provided herein comprises an amino acid sequence at least 80% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26. In another aspect, a Cas12a nuclease provided herein comprises an amino acid sequence at least 85% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26. In another aspect, a Cas12a nuclease provided herein comprises an amino acid sequence at least 90% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26. In another aspect, a Cas12a nuclease provided herein comprises an amino acid sequence at least 95% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26. In another aspect, a Cas12a nuclease provided herein comprises an amino acid sequence at least 96% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26. In another aspect, a Cas12a nuclease provided herein comprises an amino acid sequence at least 97% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26. In another aspect, a Cas12a nuclease provided herein comprises an amino acid sequence at least 98% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26. In another aspect, a Cas12a nuclease provided herein comprises an amino acid sequence at least 99% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26. In another aspect, a Cas12a nuclease provided herein comprises an amino acid sequence 100% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26.
[0032] In one aspect, the Cas12a nuclease is encoded by a polynucleotide comprising a sequence at least 80% identical to a polynucleotide selected from the group consisting of SEQ ID NOs: 4 and 16.
[0033] In another aspect, the Cas12a nuclease is encoded by a polynucleotide comprising a sequence at least 85% identical to a polynucleotide selected from the group consisting of SEQ ID NOs: 4 and 16. In another aspect, the Cas12a nuclease is encoded by a polynucleotide comprising a sequence at least 90% identical to a polynucleotide selected from the group consisting of SEQ ID NOs: 4 and 16. In another aspect, the Cas12a nuclease is encoded by a polynucleotide comprising a sequence at least 95% identical to a polynucleotide selected from the group consisting of SEQ ID NOs: 4 and 16. In another aspect, the Cas12a nuclease is encoded by a polynucleotide comprising a sequence at least 96% identical to a polynucleotide selected from the group consisting of SEQ ID NOs: 4 and 16. In another aspect, the Cas12a nuclease is encoded by a polynucleotide comprising a sequence at least 97% identical to a polynucleotide selected from the group consisting of SEQ ID NOs: 4 and 16. In another aspect, the Cas12a nuclease is encoded by a polynucleotide comprising a sequence at least 98% identical to a polynucleotide selected from the group consisting of SEQ ID NOs: 4 and 16. In another aspect, the Cas12a nuclease is encoded by a polynucleotide comprising a sequence at least 99% identical to a polynucleotide selected from the group consisting of SEQ ID NOs: 4 and 16. In another aspect, the Cas12a nuclease is encoded by a polynucleotide comprising a sequence 100% identical to a polynucleotide selected from the group consisting of SEQ ID NOs: 4 and 16.
[0034] CasX is a class II CRISPR-Cas nuclease identified in the bacterial phyla Deltaproteobacteria and Planctomycetes. Like Cas12a, CasX nucleases produce staggered cuts when cleaving target nucleic acid molecules. However, unlike Cas12a, CasX nucleases require a crRNA and a tracrRNA, or a single guide RNA, to target and cleave the target nucleic acid.
[0035] In one aspect, the CasX nuclease provided herein is a CasX nuclease from the phylum Deltaproteobacteria. In another aspect, the CasX nuclease provided herein is a CasX nuclease from the phylum Planctomycetes. Additional suitable CasX nucleases are those described in WO 2019 / 084148, which is incorporated herein by reference in its entirety.
[0036] In one aspect, a Cas12a nuclease or CasX nuclease provided herein can be expressed from a recombinant vector in vivo. In one aspect, a Cas12a nuclease or CasX nuclease provided herein can be expressed from a recombinant vector in vitro. In one aspect, a Cas12a nuclease or CasX nuclease provided herein can be expressed from a recombinant vector ex vivo. In one aspect, a Cas12a nuclease or CasX nuclease provided herein can be expressed from a nucleic acid molecule in vivo. In one aspect, a Cas12a nuclease or CasX nuclease provided herein can be expressed from a nucleic acid molecule in vitro. In one aspect, a Cas12a nuclease or CasX nuclease provided herein can be expressed from a nucleic acid molecule ex vivo.
[0037] As used herein, a "guide nucleic acid" refers to a nucleic acid that forms a complex with a CRISPR nuclease (e.g., but not limited to, Cas12a, CasX) and then guides the complex to a specific sequence in a target nucleic acid molecule, wherein the guide nucleic acid and the target nucleic acid molecule share a complementary sequence. In one aspect, a ribonucleoprotein provided herein comprises at least one guide nucleic acid.
[0038] In one embodiment, the guide nucleic acid comprises DNA. In another embodiment, the guide nucleic acid comprises RNA. In one embodiment, the guide nucleic acid comprises DNA, RNA, or a combination thereof. In one embodiment, the guide nucleic acid is single-stranded. In another embodiment, the guide nucleic acid is at least partially double-stranded.
[0039] When the guide nucleic acid comprises RNA, it can be referred to as a "guide RNA." In another embodiment, the guide nucleic acid comprises DNA and RNA. In another embodiment, the guide nucleic acid is single-stranded. In another embodiment, the guide nucleic acid is double-stranded. In a further embodiment, the guide nucleic acid is partially double-stranded.
[0040] In another embodiment, the guide nucleic acid comprises at least 10 nucleotides. In another embodiment, the guide nucleic acid comprises at least 11 nucleotides. In another embodiment, the guide nucleic acid comprises at least 12 nucleotides. In another embodiment, the guide nucleic acid comprises at least 13 nucleotides. In another embodiment, the guide nucleic acid comprises at least 14 nucleotides. In another embodiment, the guide nucleic acid comprises at least 15 nucleotides. In another embodiment, the guide nucleic acid comprises at least 16 nucleotides. In another embodiment, the guide nucleic acid comprises at least 17 nucleotides. In another embodiment, the guide nucleic acid comprises at least 18 nucleotides. In another embodiment, the guide nucleic acid comprises at least 19 nucleotides. In another embodiment, the guide nucleic acid comprises at least 20 nucleotides. In another embodiment, the guide nucleic acid comprises at least 21 nucleotides. In another embodiment, the guide nucleic acid comprises at least 22 nucleotides. In another embodiment, the guide nucleic acid comprises at least 23 nucleotides. In another embodiment, the guide nucleic acid comprises at least 24 nucleotides. In another embodiment, the guide nucleic acid comprises at least 25 nucleotides. In another embodiment, the guide nucleic acid comprises at least 26 nucleotides. In another embodiment, the guide nucleic acid comprises at least 27 nucleotides. In another embodiment, the guide nucleic acid comprises at least 28 nucleotides. In another embodiment, the guide nucleic acid comprises at least 30 nucleotides. In another embodiment, the guide nucleic acid comprises at least 35 nucleotides. In another embodiment, the guide nucleic acid comprises at least 40 nucleotides. In another embodiment, the guide nucleic acid comprises at least 45 nucleotides. In another embodiment, the guide nucleic acid comprises at least 50 nucleotides.
[0041] In another embodiment, the guide nucleic acid comprises 10 to 50 nucleotides. In another embodiment, the guide nucleic acid comprises 10 to 40 nucleotides. In another embodiment, the guide nucleic acid comprises 10 to 30 nucleotides. In another embodiment, the guide nucleic acid comprises 10 to 20 nucleotides. In another embodiment, the guide nucleic acid comprises 16 to 28 nucleotides. In another embodiment, the guide nucleic acid comprises 16 to 25 nucleotides. In another embodiment, the guide nucleic acid comprises 16 to 20 nucleotides.
[0042] In one embodiment, the guide nucleic acid comprises at least 70% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 75% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 80% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 85% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 90% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 91% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 92% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 93% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 94% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 95% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 96% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 97% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 98% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises at least 99% sequence complementarity to the target site. In one embodiment, the guide nucleic acid comprises 100% sequence complementarity to the target site. In another embodiment, the guide nucleic acid comprises 70%-100% sequence complementarity to the target site. In another embodiment, the guide nucleic acid comprises 80%-100% sequence complementarity to the target site. In another embodiment, the guide nucleic acid comprises 90%-100% sequence complementarity to the target site.
[0043] In one aspect, the guide nucleic acid is capable of hybridizing to the target site.
[0044] As mentioned above, some RNA-guided CRISPR nucleases, such as CasX and Cas9, require a separate non-coding RNA component called a trans-activating crRNA (tracrRNA) for functional activity. The guide nucleic acid molecules provided herein can combine the crRNA and tracrRNA into a single nucleic acid molecule, referred to herein as a "single guide RNA" (sgRNA). The gRNA guides the active CasX complex to the target site, and CasX can cleave the target site. In other embodiments, the crRNA and tracrRNA are provided as separate nucleic acid molecules.
[0045] In one embodiment, the guide nucleic acid comprises a crRNA. In another embodiment, the guide nucleic acid comprises a tracrRNA. In a further embodiment, the guide nucleic acid comprises an sgRNA.
[0046] In one aspect, the guide nucleic acids provided herein can be expressed from a recombinant vector in vivo. In one aspect, the guide nucleic acids provided herein can be expressed from a recombinant vector in vitro. In one aspect, the guide nucleic acids provided herein can be expressed from a recombinant vector ex vivo. In one aspect, the guide nucleic acids provided herein can be expressed from a nucleic acid molecule in vivo. In one aspect, the guide nucleic acids provided herein can be expressed from a nucleic acid molecule in vitro. In one aspect, the guide nucleic acids provided herein can be expressed from a nucleic acid molecule ex vivo. In another aspect, the guide nucleic acids provided herein can be synthetically synthesized.
[0047] Linkers are short amino acid sequences used to connect two or more proteins or protein domains into a larger protein complex. Linkers do not interfere with the native function of the proteins or protein domains they connect.
[0048] In one aspect, a linker is disposed between the amino acid sequence encoding the first nuclease and the amino acid sequence encoding the second nuclease. In one aspect, the first nuclease is selected from the group consisting of Cas12a nuclease, CasX nuclease, Cas9 nuclease, meganuclease, zinc finger nuclease, transcription activator-like nuclease, and HUH endonuclease. In one aspect, the second nuclease is selected from the group consisting of Cas12a nuclease, CasX nuclease, Cas9 nuclease, meganuclease, zinc finger nuclease, transcription activator-like nuclease, and HUH endonuclease.
[0049] In one embodiment, a linker is positioned between the amino acid sequence encoding the nuclease and the amino acid sequence encoding the functional domain. In one aspect, the nuclease is a meganuclease, a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), an Argonaute (non-limiting examples of Argonaute proteins include Thermus thermophilus Argonaute (TtAgo), Pyrococcus furiosus Argonaute (PfAgo), and Natronobacterium gregoryi Argonaute (NgAgo)), an RNA-guided nuclease, e.g., a CRISPR-associated nuclease (non-limiting examples of CRISPR-associated nucleases include Casl, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, and Cas9). (also known as Csn1 and Csx12), Cas10, Cas12a (also known as Cpf1), Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csl3, Csf4, CasX, CasY, homologs thereof, or modified versions thereof). In one aspect, the functional domain is selected from the group consisting of a deaminase, a uracil-DNA glycosylase (UGI), a transcriptional activator, a recombinase, a transposase, a helicase, and a methylase. In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, the deaminase is an adenine deaminase. In some embodiments, the deaminase is an APOPEC deaminase. In some embodiments, the deaminase is an activation-induced cytidine deaminase (AID).Non-limiting examples of recombinases include tyrosine recombinases linked to linkers provided herein and are selected from the group consisting of Cre recombinase, Gin recombinase, Flp recombinase, and Tnpl recombinase. In another aspect, serine recombinases linked to linkers provided herein are selected from the group consisting of PhiC31 integrase, R4 integrase, and TP-901 integrase. In another aspect, DNA transposases linked to linkers provided herein are selected from the group consisting of TALE-piggyBac and TALE-Mutator.
[0050] In one embodiment, a linker is disposed between a first amino acid sequence encoding a Cas12a nuclease and a second amino acid sequence encoding a HUH endonuclease. In another embodiment, a linker is disposed between a first amino acid sequence encoding a CasX nuclease and a second amino acid sequence encoding a HUH endonuclease.
[0051] In one embodiment, the linker is positioned at the 5' end of the Cas12a nuclease. In another embodiment, the linker is positioned at the 3' end of the Cas12a nuclease. In another embodiment, the linker is positioned at the 5' end of the CasX nuclease. In another embodiment, the linker is positioned at the 3' end of the CasX nuclease. In another embodiment, the linker is positioned at the 5' end of the HUH endonuclease. In another embodiment, the linker is positioned at the 3' end of the HUH endonuclease.
[0052] In one embodiment, the linker comprises at least 5 amino acids. In another embodiment, the linker comprises at least 10 amino acids. In another embodiment, the linker comprises at least 15 amino acids. In another embodiment, the linker comprises at least 20 amino acids. In another embodiment, the linker comprises at least 25 amino acids. In another embodiment, the linker comprises at least 30 amino acids. In another embodiment, the linker comprises at least 40 amino acids. In another embodiment, the linker comprises at least 50 amino acids.
[0053] In one embodiment, the linker comprises between 5 and 50 amino acids. In another embodiment, the linker comprises between 5 and 40 amino acids. In another embodiment, the linker comprises between 5 and 30 amino acids. In another embodiment, the linker comprises between 5 and 20 amino acids. In another embodiment, the linker comprises between 10 and 50 amino acids. In another embodiment, the linker comprises between 10 and 40 amino acids. In another embodiment, the linker comprises between 10 and 30 amino acids. In another embodiment, the linker comprises between 10 and 20 amino acids.
[0054] In one embodiment, the linker comprises 1 amino acid. In another embodiment, the linker comprises 2 amino acids. In another embodiment, the linker comprises 3 amino acids. In another embodiment, the linker comprises 4 amino acids. In one embodiment, the linker comprises 5 amino acids. In another embodiment, the linker comprises 6 amino acids. In another embodiment, the linker comprises 7 amino acids. In another embodiment, the linker comprises 8 amino acids. In another embodiment, the linker comprises 9 amino acids. In another embodiment, the linker comprises 10 amino acids. In another embodiment, the linker comprises 11 amino acids. In another embodiment, the linker comprises 12 amino acids. In another embodiment, the linker comprises 13 amino acids. In another embodiment, the linker comprises 14 amino acids. In another embodiment, the linker comprises 15 amino acids. In another embodiment, the linker comprises 16 amino acids. In another embodiment, the linker comprises 17 amino acids. In another embodiment, the linker comprises 18 amino acids. In another embodiment, the linker comprises 19 amino acids. In another embodiment, the linker comprises 20 amino acids. In another embodiment, the linker comprises 21 amino acids. In another embodiment, the linker comprises 22 amino acids. In another embodiment, the linker comprises 23 amino acids. In another embodiment, the linker comprises 24 amino acids. In another embodiment, the linker comprises 25 amino acids. In another embodiment, the linker comprises 26 amino acids. In another embodiment, the linker comprises 27 amino acids. In another embodiment, the linker comprises 28 amino acids. In another embodiment, the linker comprises 29 amino acids. In another embodiment, the linker comprises 30 amino acids.
[0055] In one aspect, the present disclosure provides an isolated polypeptide comprising an amino acid sequence at least 70% identical or similar to SEQ ID NO:3. In one aspect, the present disclosure provides an isolated polypeptide comprising an amino acid sequence at least 75% identical or similar to SEQ ID NO:3. In one aspect, the present disclosure provides an isolated polypeptide comprising an amino acid sequence at least 80% identical or similar to SEQ ID NO:3. In one aspect, the present disclosure provides an isolated polypeptide comprising an amino acid sequence at least 85% identical or similar to SEQ ID NO:3. In one aspect, the present disclosure provides an isolated polypeptide comprising an amino acid sequence at least 90% identical or similar to SEQ ID NO:3. In one aspect, the present disclosure provides an isolated polypeptide comprising an amino acid sequence at least 91% identical or similar to SEQ ID NO:3. In one aspect, the present disclosure provides an isolated polypeptide comprising an amino acid sequence at least 92% identical or similar to SEQ ID NO:3. In one aspect, the present disclosure provides an isolated polypeptide comprising an amino acid sequence at least 93% identical or similar to SEQ ID NO:3. In one aspect, the present disclosure provides an isolated polypeptide comprising an amino acid sequence at least 94% identical or similar to SEQ ID NO:3. In one aspect, the disclosure provides an isolated polypeptide comprising an amino acid sequence at least 95% identical or similar to SEQ ID NO:3. In one aspect, the disclosure provides an isolated polypeptide comprising an amino acid sequence at least 96% identical or similar to SEQ ID NO:3. In one aspect, the disclosure provides an isolated polypeptide comprising an amino acid sequence at least 97% identical or similar to SEQ ID NO:3. In one aspect, the disclosure provides an isolated polypeptide comprising an amino acid sequence at least 98% identical or similar to SEQ ID NO:3. In one aspect, the disclosure provides an isolated polypeptide comprising an amino acid sequence at least 99% identical or similar to SEQ ID NO:3. In one aspect, the disclosure provides an isolated polypeptide comprising an amino acid sequence 100% identical or similar to SEQ ID NO:3.
[0056] In one embodiment, the linker comprises an amino acid sequence at least 70% identical or similar to SEQ ID NO:3. In another embodiment, the linker comprises an amino acid sequence at least 75% identical or similar to SEQ ID NO:3. In one embodiment, the linker comprises an amino acid sequence at least 80% identical or similar to SEQ ID NO:3. In another embodiment, the linker comprises an amino acid sequence at least 85% identical or similar to SEQ ID NO:3. In another embodiment, the linker comprises an amino acid sequence at least 90% identical or similar to SEQ ID NO:3. In another embodiment, the linker comprises an amino acid sequence at least 91% identical or similar to SEQ ID NO:3. In another embodiment, the linker comprises an amino acid sequence at least 92% identical or similar to SEQ ID NO:3. In another embodiment, the linker comprises an amino acid sequence at least 93% identical or similar to SEQ ID NO:3. In another embodiment, the linker comprises an amino acid sequence at least 94% identical or similar to SEQ ID NO:3. In another embodiment, the linker comprises an amino acid sequence at least 95% identical or similar to SEQ ID NO:3. In another embodiment, the linker comprises an amino acid sequence at least 96% identical or similar to SEQ ID NO: 3. In another embodiment, the linker comprises an amino acid sequence at least 97% identical or similar to SEQ ID NO: 3. In another embodiment, the linker comprises an amino acid sequence at least 98% identical or similar to SEQ ID NO: 3. In another embodiment, the linker comprises an amino acid sequence at least 99% identical or similar to SEQ ID NO: 3. In another embodiment, the linker comprises an amino acid sequence 100% identical or similar to SEQ ID NO: 3.
[0057] In one aspect, a linker provided herein comprises SEQ ID NO:3.
[0058] In one embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 70% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 75% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 80% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 85% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 90% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 91% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 92% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 93% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 94% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 95% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 96% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 97% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 98% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence at least 99% identical to SEQ ID NO:6. In another embodiment, the polynucleotide encoding the linker comprises a polynucleotide sequence 100% identical to SEQ ID NO:6.
[0059] Amino acids contain a carboxylic acid (COOH) group and an amino group (NH) attached to a carbon atom, and each amino acid further contains a variable R group. Based on the properties of the R group, amino acids can be characterized into groups of hydrophobic (non-polar) and hydrophilic (polar) amino acids.
[0060] Hydrophobic amino acids include glycine, alanine, valine, leucine, isoleucine, methionine, proline, phenylalanine, and tryptophan. Hydrophilic amino acids include tyrosine, serine, threonine, cysteine, glutamine, asparagine, glutamic acid, aspartic acid, lysine, histidine, and arginine.
[0061] In one embodiment, the amino acid sequence of the linker comprises at least 5% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 10% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 15% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 20% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 25% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 30% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 35% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 40% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 45% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 50% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 55% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 60% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 65% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 70% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 75% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 80% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 85% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 90% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 95% hydrophobic amino acid residues.
[0062] In one embodiment, the amino acid sequence of the linker comprises 5% to 90% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 5% to 75% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 5% to 50% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 5% to 25% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 25% to 75% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 50% to 75% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 10% to 75% hydrophobic amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 10% to 90% hydrophobic amino acid residues.
[0063] In one embodiment, at least 10% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 15% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 20% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 25% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 30% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 35% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 40% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 45% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 50% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 55% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 60% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 65% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 70% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 75% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 80% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 85% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine. In another embodiment, at least 90% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine in any combination. In another embodiment, at least 95% of the amino acid sequence of the linker is made up of any combination of glycine, threonine, and serine.
[0064] In one embodiment, 10% to 95% of the amino acid sequence of the linker is composed of any combination of glycine, threonine, and serine. In another embodiment, 10% to 75% of the amino acid sequence of the linker is composed of any combination of glycine, threonine, and serine. In another embodiment, 10% to 50% of the amino acid sequence of the linker is composed of any combination of glycine, threonine, and serine. In another embodiment, 10% to 25% of the amino acid sequence of the linker is composed of any combination of glycine, threonine, and serine. In another embodiment, 25% to 75% of the amino acid sequence of the linker is composed of any combination of glycine, threonine, and serine. In another embodiment, 50% to 75% of the amino acid sequence of the linker is composed of any combination of glycine, threonine, and serine.
[0065] At physiological pH (e.g., 7.4), some amino acids have electrically charged R groups: for example, arginine, histidine, and lysine are positively charged amino acids, serine, threonine, asparagine, and glutamine are uncharged amino acids, and aspartic acid and glutamic acid are negatively charged amino acids.
[0066] In one embodiment, the amino acid sequence of the linker comprises at least 5% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 10% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 15% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 20% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 25% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 30% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 35% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 40% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 45% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 50% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 55% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 60% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 65% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 70% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 75% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 80% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 85% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 90% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises at least 95% negatively charged amino acid residues.
[0067] In one embodiment, the amino acid sequence of the linker comprises 5% to 90% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 5% to 75% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 5% to 50% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 5% to 25% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 25% to 75% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 50% to 75% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 10% to 75% negatively charged amino acid residues. In another embodiment, the amino acid sequence of the linker comprises 10% to 90% negatively charged amino acid residues.
[0068] In one aspect, the linkers provided herein can be expressed from a recombinant vector in vivo. In one aspect, the linkers provided herein can be expressed from a recombinant vector in vitro. In one aspect, the linkers provided herein can be expressed from a recombinant vector ex vivo. In one aspect, the linkers provided herein can be expressed from a nucleic acid molecule in vivo. In one aspect, the linkers provided herein can be expressed from a nucleic acid molecule in vitro. In one aspect, the linkers provided herein can be expressed from a nucleic acid molecule ex vivo.
[0069] HUH endonucleases contain a characteristic motif consisting of a first histidine residue (H), a hydrophobic amino acid residue (U), and a second histidine residue (H). HUH endonucleases are known from archaea, bacteria, and eukaryotes. Endogenous HUH endonucleases are involved in cellular processes involving the transition from double-stranded DNA to single-stranded DNA, such as rolling circle replication in viruses and bacterial plasmid conjugation. HUH endonuclease first nicks single-stranded DNA at a specific sequence in the origin of replication (ori), followed by the formation of a covalent phosphotyrosine intermediate, which ligates the 5' end of the DNA strand to a specific tyrosine on the HUH protein. While phosphotyrosine ligation is an intermediate in vivo, purified HUH protein can form stable covalent bonds in vitro with synthetic oligonucleotides containing the ori sequence.
[0070] In one embodiment, the HUH endonuclease hybridizes to an origin of replication (ori) sequence.
[0071] In one embodiment, the HUH endonuclease is a porcine circovirus 2 (PCV) HUH endonuclease. In another embodiment, the HUH endonuclease is a duck circovirus (DCV) HUH endonuclease. In another embodiment, the HUH endonuclease is a broad bean necrotic yellows virus (FBNYV) HUH endonuclease. In another embodiment, the HUH endonuclease is a Streptococcus agalactiae replication protein RepB (RepB). In another embodiment, the HUH endonuclease is a Fructobacillus tropaeoli RepB (RepBm). In another embodiment, the HUH endonuclease is an Escherichia coli conjugation protein TraI (TraI). In another embodiment, the HUH endonuclease is E. coli mobilization protein A (mMobA). In another embodiment, the HUH endonuclease is Staphylococcus aureus nicking enzyme (NES).
[0072] In one aspect, the HUH endonuclease is selected from the group consisting of FBNYV HUH endonuclease, PCV HUH endonuclease, DCV HUH endonuclease, RepB, RepBm, Tral, mMobA, and NES.
[0073] In one aspect, the FBNYV HUH endonuclease comprises an amino acid sequence at least 80% identical or similar to the amino acid sequence of SEQ ID NO:2. In another aspect, the FBNYV HUH endonuclease comprises an amino acid sequence at least 85% identical or similar to the amino acid sequence of SEQ ID NO:2. In another aspect, the FBNYV HUH endonuclease comprises an amino acid sequence at least 90% identical or similar to the amino acid sequence of SEQ ID NO:2. In another aspect, the FBNYV HUH endonuclease comprises an amino acid sequence at least 95% identical or similar to the amino acid sequence of SEQ ID NO:2. In another aspect, the FBNYV HUH endonuclease comprises an amino acid sequence at least 96% identical or similar to the amino acid sequence of SEQ ID NO:2. In another aspect, the FBNYV HUH endonuclease comprises an amino acid sequence at least 97% identical or similar to the amino acid sequence of SEQ ID NO:2. In another aspect, the FBNYV HUH endonuclease comprises an amino acid sequence at least 98% identical or similar to the amino acid sequence of SEQ ID NO:2. In another embodiment, the FBNYV HUH endonuclease comprises an amino acid sequence that is at least 99% identical or similar to the amino acid sequence of SEQ ID NO: 2. In another embodiment, the FBNYV HUH endonuclease comprises an amino acid sequence that is 100% identical or similar to the amino acid sequence of SEQ ID NO: 2.
[0074] In one embodiment, the PCV HUH endonuclease comprises an amino acid sequence at least 80% identical or similar to the amino acid sequence of SEQ ID NO: 14. In another embodiment, the PCV HUH endonuclease comprises an amino acid sequence at least 85% identical or similar to the amino acid sequence of SEQ ID NO: 14. In another embodiment, the PCV HUH endonuclease comprises an amino acid sequence at least 90% identical or similar to the amino acid sequence of SEQ ID NO: 14. In another embodiment, the PCV HUH endonuclease comprises an amino acid sequence at least 95% identical or similar to the amino acid sequence of SEQ ID NO: 14. In another embodiment, the PCV HUH endonuclease comprises an amino acid sequence at least 96% identical or similar to the amino acid sequence of SEQ ID NO: 14. In another embodiment, the PCV HUH endonuclease comprises an amino acid sequence at least 97% identical or similar to the amino acid sequence of SEQ ID NO: 14. In another embodiment, the PCV HUH endonuclease comprises an amino acid sequence at least 98% identical or similar to the amino acid sequence of SEQ ID NO: 14. In another embodiment, the PCV HUH endonuclease comprises an amino acid sequence at least 99% identical or similar to the amino acid sequence of SEQ ID NO: 14. In another embodiment, the PCV HUH endonuclease comprises an amino acid sequence 100% identical or similar to the amino acid sequence of SEQ ID NO: 14.
[0075] In one aspect, the HUH endonuclease comprises an amino acid sequence at least 80% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. In one aspect, the HUH endonuclease comprises an amino acid sequence at least 85% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. In one aspect, the HUH endonuclease comprises an amino acid sequence at least 90% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. In one aspect, the HUH endonuclease comprises an amino acid sequence at least 95% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. In one aspect, the HUH endonuclease comprises an amino acid sequence at least 96% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. In one aspect, the HUH endonuclease comprises an amino acid sequence at least 97% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. In one aspect, the HUH endonuclease comprises an amino acid sequence at least 98% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. In one aspect, the HUH endonuclease comprises an amino acid sequence at least 99% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. In one aspect, the HUH endonuclease comprises an amino acid sequence 100% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14.
[0076] In one embodiment, the FBNYV HUH endonuclease is encoded by a nucleic acid sequence at least 80% identical to the nucleic acid sequence of SEQ ID NO:5. In another embodiment, the FBNYV HUH endonuclease is encoded by a nucleic acid sequence at least 85% identical to the nucleic acid sequence of SEQ ID NO:5. In another embodiment, the FBNYV HUH endonuclease is encoded by a nucleic acid sequence at least 90% identical to the nucleic acid sequence of SEQ ID NO:5. In another embodiment, the FBNYV HUH endonuclease is encoded by a nucleic acid sequence at least 95% identical to the nucleic acid sequence of SEQ ID NO:5. In another embodiment, the FBNYV HUH endonuclease is encoded by a nucleic acid sequence at least 96% identical to the nucleic acid sequence of SEQ ID NO:5. In another embodiment, the FBNYV HUH endonuclease is encoded by a nucleic acid sequence at least 97% identical to the nucleic acid sequence of SEQ ID NO:5. In another embodiment, the FBNYV HUH endonuclease is encoded by a nucleic acid sequence at least 98% identical to the nucleic acid sequence of SEQ ID NO:5. In another embodiment, the FBNYV HUH endonuclease is encoded by a nucleic acid sequence that is at least 99% identical to the nucleic acid sequence of SEQ ID NO: 5. In another embodiment, the FBNYV HUH endonuclease is encoded by a nucleic acid sequence that is 100% identical to the nucleic acid sequence of SEQ ID NO: 5.
[0077] In one embodiment, the PCV HUH endonuclease is encoded by a nucleic acid sequence at least 80% identical to the nucleic acid sequence of SEQ ID NO: 13. In another embodiment, the PCV HUH endonuclease is encoded by a nucleic acid sequence at least 85% identical to the nucleic acid sequence of SEQ ID NO: 13. In another embodiment, the PCV HUH endonuclease is encoded by a nucleic acid sequence at least 90% identical to the nucleic acid sequence of SEQ ID NO: 13. In another embodiment, the PCV HUH endonuclease is encoded by a nucleic acid sequence at least 95% identical to the nucleic acid sequence of SEQ ID NO: 13. In another embodiment, the PCV HUH endonuclease is encoded by a nucleic acid sequence at least 96% identical to the nucleic acid sequence of SEQ ID NO: 13. In another embodiment, the PCV HUH endonuclease is encoded by a nucleic acid sequence at least 97% identical to the nucleic acid sequence of SEQ ID NO: 13. In another embodiment, the PCV HUH endonuclease is encoded by a nucleic acid sequence at least 98% identical to the nucleic acid sequence of SEQ ID NO: 13. In another embodiment, the PCV HUH endonuclease is encoded by a nucleic acid sequence that is at least 99% identical to the nucleic acid sequence of SEQ ID NO: 13. In another embodiment, the PCV HUH endonuclease is encoded by a nucleic acid sequence that is 100% identical to the nucleic acid sequence of SEQ ID NO: 13.
[0078] In one aspect, the HUH endonuclease is encoded by a nucleic acid sequence at least 80% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 5 and 13. In one aspect, the HUH endonuclease is encoded by a nucleic acid sequence at least 85% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 5 and 13. In one aspect, the HUH endonuclease is encoded by a nucleic acid sequence at least 90% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 5 and 13. In one aspect, the HUH endonuclease is encoded by a nucleic acid sequence at least 95% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 5 and 13. In one aspect, the HUH endonuclease is encoded by a nucleic acid sequence at least 96% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 5 and 13. In one aspect, the HUH endonuclease is encoded by a nucleic acid sequence at least 97% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 5 and 13. In one aspect, the HUH endonuclease is encoded by a nucleic acid sequence that is at least 98% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 5 and 13. In one aspect, the HUH endonuclease is encoded by a nucleic acid sequence that is at least 99% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 5 and 13. In one aspect, the HUH endonuclease is encoded by a nucleic acid sequence that is 100% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 5 and 13.
[0079] In one aspect, the HUH endonuclease provided herein can be expressed from a recombinant vector in vivo. In one aspect, the HUH endonuclease provided herein can be expressed from a recombinant vector in vitro. In one aspect, the HUH endonuclease provided herein can be expressed from a recombinant vector ex vivo. In one aspect, the HUH endonuclease provided herein can be expressed from a nucleic acid molecule in vivo. In one aspect, the HUH endonuclease provided herein can be expressed from a nucleic acid molecule in vitro. In one aspect, the HUH endonuclease provided herein can be expressed from a nucleic acid molecule ex vivo.
[0080] In one aspect, the ribonucleoprotein comprises an amino acid sequence at least 80% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23 and 25. In one aspect, the ribonucleoprotein comprises an amino acid sequence at least 85% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23 and 25. In one aspect, the ribonucleoprotein comprises an amino acid sequence at least 90% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23 and 25. In one aspect, the ribonucleoprotein comprises an amino acid sequence at least 95% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23 and 25. In one aspect, the ribonucleoprotein comprises an amino acid sequence at least 96% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23 and 25. In one aspect, the ribonucleoprotein comprises an amino acid sequence at least 97% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23 and 25. In one aspect, the ribonucleoprotein comprises an amino acid sequence at least 98% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23 and 25. In one aspect, the ribonucleoprotein comprises an amino acid sequence at least 99% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23 and 25. In one aspect, the ribonucleoprotein comprises an amino acid sequence 100% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23 and 25.
[0081] In one aspect, the recombinant nucleic acid encodes an amino acid sequence at least 80% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20 to 23 and 25. In one aspect, the recombinant nucleic acid encodes an amino acid sequence at least 85% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20 to 23 and 25. In one aspect, the recombinant nucleic acid encodes an amino acid sequence at least 90% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20 to 23 and 25. In one aspect, the recombinant nucleic acid encodes an amino acid sequence at least 95% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20 to 23 and 25. In one aspect, the recombinant nucleic acid encodes an amino acid sequence at least 96% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20 to 23 and 25. In one aspect, the recombinant nucleic acid encodes an amino acid sequence at least 97% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20 to 23 and 25. In one aspect, the recombinant nucleic acid encodes an amino acid sequence at least 98% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20 to 23 and 25. In one aspect, the recombinant nucleic acid encodes an amino acid sequence that is at least 99% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23 and 25. In one aspect, the recombinant nucleic acid encodes an amino acid sequence that is 100% identical or similar to an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23 and 25.
[0082] In embodiments, the recombinant nucleic acid comprises a polynucleotide sequence at least 80% identical to a sequence selected from the group consisting of SEQ ID NOs: 7, 17-19, and 24. In embodiments, the recombinant nucleic acid comprises a polynucleotide sequence at least 85% identical to a sequence selected from the group consisting of SEQ ID NOs: 7, 17-19, and 24. In embodiments, the recombinant nucleic acid comprises a polynucleotide sequence at least 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 7, 17-19, and 24. In embodiments, the recombinant nucleic acid comprises a polynucleotide sequence at least 95% identical to a sequence selected from the group consisting of SEQ ID NOs: 7, 17-19, and 24. In embodiments, the recombinant nucleic acid comprises a polynucleotide sequence at least 96% identical to a sequence selected from the group consisting of SEQ ID NOs: 7, 17-19, and 24. In embodiments, the recombinant nucleic acid comprises a polynucleotide sequence at least 97% identical to a sequence selected from the group consisting of SEQ ID NOs: 7, 17-19, and 24. In an embodiment, the recombinant nucleic acid comprises a polynucleotide sequence at least 98% identical to a sequence selected from the group consisting of SEQ ID NOs: 7, 17-19, and 24. In an embodiment, the recombinant nucleic acid comprises a polynucleotide sequence at least 99% identical to a sequence selected from the group consisting of SEQ ID NOs: 7, 17-19, and 24. In an embodiment, the recombinant nucleic acid comprises a polynucleotide sequence at least 100% identical to a sequence selected from the group consisting of SEQ ID NOs: 7, 17-19, and 24.
[0083] As used herein, "template nucleic acid molecule" refers to a nucleic acid molecule comprising a nucleic acid sequence to be inserted into a target DNA molecule. In one embodiment, the template nucleic acid molecule comprises single-stranded DNA. In another embodiment, the template nucleic acid molecule comprises double-stranded DNA. In a further embodiment, the template nucleic acid molecule comprises single-stranded RNA. In yet another embodiment, the template nucleic acid molecule comprises double-stranded RNA. In another embodiment, the template nucleic acid molecule comprises DNA and RNA.
[0084] In one embodiment, the ribonucleoprotein comprises at least one template nucleic acid molecule. In another embodiment, the ribonucleoprotein comprises at least two template nucleic acid molecules.
[0085] In one embodiment, the template nucleic acid molecule contains at least 10 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 25 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 50 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 75 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 100 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 250 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 500 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 750 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 1000 nucleotides. In another embodiment, the template nucleic acid molecule contains at least 2500 nucleotides.
[0086] In one embodiment, the template nucleic acid molecule comprises between 10 and 2500 nucleotides. In another embodiment, the template nucleic acid molecule comprises between 25 and 2500 nucleotides. In another embodiment, the template nucleic acid molecule comprises between 50 and 2500 nucleotides. In another embodiment, the template nucleic acid molecule comprises between 75 and 2500 nucleotides. In another embodiment, the template nucleic acid molecule comprises between 100 and 2500 nucleotides. In another embodiment, the template nucleic acid molecule comprises between 250 and 2500 nucleotides. In another embodiment, the template nucleic acid molecule comprises between 500 and 2500 nucleotides. In another embodiment, the template nucleic acid molecule comprises between 25 and 1000 nucleotides. In another embodiment, the template nucleic acid molecule comprises between 25 and 500 nucleotides. In another embodiment, the template nucleic acid molecule comprises between 25 and 250 nucleotides.
[0087] As described above, HUH endonuclease first nicks single-stranded DNA at a specific sequence in the origin of replication (ori), followed by the formation of a covalent phosphotyrosine intermediate, which ligates the 5' end of the DNA strand to a specific tyrosine on the HUH protein. Although the phosphotyrosine ligation is an intermediate in vivo, purified HUH protein can form stable covalent bonds in vitro with synthetic oligonucleotides bearing the ori sequence.
[0088] In embodiments, the template nucleic acid molecule comprises a nucleic acid sequence encoding an origin of replication (ori). In one embodiment, the ori provided herein is capable of hybridizing to HUH endonuclease. In one embodiment, the ori comprises a nucleic acid sequence at least 80% identical to SEQ ID NO: 12. In another embodiment, the ori comprises a nucleic acid sequence at least 85% identical to SEQ ID NO: 12. In another embodiment, the ori comprises a nucleic acid sequence at least 90% identical to SEQ ID NO: 12. In another embodiment, the ori comprises a nucleic acid sequence at least 95% identical to SEQ ID NO: 12. In another embodiment, the ori comprises a nucleic acid sequence at least 96% identical to SEQ ID NO: 12. In another embodiment, the ori comprises a nucleic acid sequence at least 97% identical to SEQ ID NO: 12. In another embodiment, the ori comprises a nucleic acid sequence at least 98% identical to SEQ ID NO: 12. In another embodiment, the ori comprises a nucleic acid sequence at least 99% identical to SEQ ID NO: 12. In another embodiment, the ori comprises a nucleic acid sequence 100% identical to SEQ ID NO: 12. In one embodiment, the ori comprises a nucleic acid sequence at least 80% identical to SEQ ID NO: 27. In another embodiment, the ori comprises a nucleic acid sequence at least 85% identical to SEQ ID NO:27. In another embodiment, the ori comprises a nucleic acid sequence at least 90% identical to SEQ ID NO:27. In another embodiment, the ori comprises a nucleic acid sequence at least 95% identical to SEQ ID NO:27. In another embodiment, the ori comprises a nucleic acid sequence at least 96% identical to SEQ ID NO:27. In another embodiment, the ori comprises a nucleic acid sequence at least 97% identical to SEQ ID NO:27. In another embodiment, the ori comprises a nucleic acid sequence at least 98% identical to SEQ ID NO:27. In another embodiment, the ori comprises a nucleic acid sequence at least 99% identical to SEQ ID NO:27. In another embodiment, the ori comprises a nucleic acid sequence 100% identical to SEQ ID NO:27. In one embodiment, the ori comprises a nucleic acid sequence at least 80% identical to SEQ ID NO:28. In another embodiment, the ori comprises a nucleic acid sequence at least 85% identical to SEQ ID NO:28. In another embodiment, the ori comprises a nucleic acid sequence at least 90% identical to SEQ ID NO:28. In another embodiment, the ori comprises a nucleic acid sequence at least 95% identical to SEQ ID NO:28. In another embodiment, the ori comprises a nucleic acid sequence at least 96% identical to SEQ ID NO: 28. In another embodiment, the ori comprises a nucleic acid sequence at least 97% identical to SEQ ID NO:28.In another embodiment, the ori comprises a nucleic acid sequence at least 98% identical to SEQ ID NO: 28. In another embodiment, the ori comprises a nucleic acid sequence at least 99% identical to SEQ ID NO: 28. In another embodiment, the ori comprises a nucleic acid sequence 100% identical to SEQ ID NO: 28.
[0089] In one embodiment, an ori comprises at least 10 nucleotides. In one embodiment, an ori comprises at least 15 nucleotides. In one embodiment, an ori comprises at least 20 nucleotides. In one embodiment, an ori comprises at least 25 nucleotides. In one embodiment, an ori comprises at least 30 nucleotides. In one embodiment, an ori comprises at least 40 nucleotides.
[0090] In one embodiment, the on is positioned at the 5' end of the template nucleic acid molecule. In another embodiment, the on is positioned at the 3' end of the template nucleic acid molecule.
[0091] In one embodiment, the template nucleic acid molecule comprises a nucleic acid sequence encoding a gene of interest. As used herein, "gene of interest" refers to a polynucleotide sequence encoding a protein or a non-protein-coding RNA molecule to be inserted into a target DNA molecule. In one embodiment, the gene of interest encodes a protein. In another embodiment, the gene of interest encodes a non-protein-coding RNA molecule.
[0092] Non-limiting examples of non-protein-coding RNA molecules include microRNAs (miRNAs), miRNA precursors (pre-miRNAs), small interfering RNAs (siRNAs), small RNAs (18-26 nucleotides in length) and their encoding precursors, heterochromatin siRNAs (hc-siRNAs), Piwi-interacting RNAs (piRNAs), hairpin double-stranded RNAs (hairpin dsRNAs), trans-acting siRNAs (ta-siRNAs), naturally occurring antisense siRNAs (nat-siRNAs), CRISPR RNAs (crRNAs), tracer RNAs (tracrRNAs), guide RNAs (gRNAs), and single guide RNAs (sgRNAs). In one embodiment, the non-protein-coding RNA molecule comprises a miRNA. In one embodiment, the non-protein-coding RNA molecule comprises a siRNA. In one embodiment, the non-protein-coding RNA molecule comprises a ta-siRNA. In one embodiment, the non-protein-coding RNA molecule is selected from the group consisting of a miRNA, an siRNA, and a ta-siRNA.
[0093] In one embodiment, the gene of interest is exogenous to the target DNA molecule. In one embodiment, the gene of interest replaces an endogenous gene of the target DNA molecule.
[0094] As used herein, "target DNA molecule" refers to a selected DNA molecule, or a selected sequence or region of a DNA molecule, where modification (e.g., cleavage, site-specific integration) is desired.
[0095] As used herein, "target region" refers to a portion of a target DNA molecule that is cleaved by a CRISPR nuclease. In contrast to non-target nucleic acids (e.g., non-target ssDNA) or non-target regions, the target site contains significant complementarity to the guide nucleic acid or guide RNA.
[0096] In one embodiment, the target site is 100% complementary to the guide nucleic acid. In another embodiment, the target site is 99% complementary to the guide nucleic acid. In another embodiment, the target site is 98% complementary to the guide nucleic acid. In another embodiment, the target site is 97% complementary to the guide nucleic acid. In another embodiment, the target site is 96% complementary to the guide nucleic acid. In another embodiment, the target site is 95% complementary to the guide nucleic acid. In another embodiment, the target site is 94% complementary to the guide nucleic acid. In another embodiment, the target site is 93% complementary to the guide nucleic acid. In another embodiment, the target site is 92% complementary to the guide nucleic acid. In another embodiment, the target site is 91% complementary to the guide nucleic acid. In another embodiment, the target site is 90% complementary to the guide nucleic acid. In another embodiment, the target site is 85% complementary to the guide nucleic acid. In another embodiment, the target site is 80% complementary to the guide nucleic acid.
[0097] In one embodiment, the target site comprises at least one PAM site. In one embodiment, the target site is adjacent to a nucleic acid sequence comprising at least one PAM site. In another embodiment, the target site is within 5 nucleotides of at least one PAM site. In a further embodiment, the target site is within 10 nucleotides of at least one PAM site. In another embodiment, the target site is within 15 nucleotides of at least one PAM site. In another embodiment, the target site is within 20 nucleotides of at least one PAM site. In another embodiment, the target site is within 25 nucleotides of at least one PAM site. In another embodiment, the target site is within 30 nucleotides of at least one PAM site.
[0098] In one embodiment, the target DNA molecule is single-stranded. In another embodiment, the target DNA molecule is double-stranded.
[0099] In one embodiment, the target DNA molecule comprises genomic DNA. In one embodiment, the target DNA molecule is located within a nuclear genome. In one embodiment, the target DNA molecule comprises chromosomal DNA. In one embodiment, the target DNA molecule comprises plasmid DNA. In one embodiment, the target DNA molecule is located within a plasmid. In one embodiment, the target DNA molecule comprises mitochondrial DNA. In one embodiment, the target DNA molecule is located within a mitochondrial genome. In one embodiment, the target DNA molecule comprises plastid DNA. In one embodiment, the target DNA molecule is located within a plastid genome. In one embodiment, the target DNA molecule comprises chloroplast DNA. In one embodiment, the target DNA molecule is located within a chloroplast genome. In one embodiment, the target DNA molecule is located within a genome selected from the group consisting of a nuclear genome, a mitochondrial genome, and a plastid genome.
[0100] In one embodiment, the target DNA molecule comprises genomic DNA. As used herein, "genomic DNA" refers to DNA that encodes one or more genes. In another embodiment, the target DNA molecule comprises intergenic DNA. In contrast to genomic DNA, "intergenic DNA" comprises non-coding DNA and lacks DNA that encodes genes. In one embodiment, intergenic DNA is located between two genes.
[0101] In one embodiment, the target nucleic acid encodes a gene. As used herein, "gene" refers to a polynucleotide capable of producing a functional unit (e.g., but not limited to, a protein or a non-coding RNA molecule). A gene can include a promoter, an enhancer sequence, a leader sequence, a transcription start site, a transcription stop site, a polyadenylation site, one or more exons, one or more introns, a 5'-UTR, a 3'-UTR, or any combination thereof. A "gene sequence" can include a polynucleotide sequence encoding a promoter, an enhancer sequence, a leader sequence, a transcription start site, a transcription stop site, a polyadenylation site, one or more exons, one or more introns, a 5'-UTR, a 3'-UTR, or any combination thereof. In one embodiment, a gene encodes a non-protein-encoding RNA molecule or a precursor thereof. In another embodiment, a gene encodes a protein. In some embodiments, the target DNA molecule is selected from the group consisting of a promoter, an enhancer sequence, a leader sequence, a transcription start site, a transcription stop site, a polyadenylation site, an exon, an intron, a splice site, a 5'-UTR, a 3'-UTR, a protein-coding sequence, a non-protein-coding sequence, an miRNA, a pre-miRNA, and an miRNA binding site.
[0102] In one aspect, the disclosure provides a method of generating an edit in a target DNA molecule, comprising contacting the target DNA molecule with a ribonucleoprotein, wherein the ribonucleoprotein comprises: (a) a recombinant polypeptide comprising: (i) an amino acid sequence encoding a Cas12a nuclease; (ii) an amino acid sequence encoding a linker; and (iii) an amino acid sequence encoding a HUH endonuclease; (b) at least one guide nucleic acid; and (c) at least one template nucleic acid molecule, wherein the ribonucleoprotein edits the target DNA molecule. In one aspect, the disclosure provides a method of generating an edit in a target DNA molecule, comprising contacting the target DNA molecule with a ribonucleoprotein, wherein the ribonucleoprotein comprises: (a) a recombinant polypeptide comprising: (i) an amino acid sequence encoding a CasX nuclease; (ii) an amino acid sequence encoding a linker; and (iii) an amino acid sequence encoding a HUH endonuclease; (b) at least one guide nucleic acid; and (c) at least one template nucleic acid molecule, wherein the ribonucleoprotein edits the target DNA molecule. In one aspect, the methods provided herein further comprise the step of (d) detecting at least one edit in the target DNA molecule.
[0103] In one aspect, the disclosure provides a method of generating an edit in a target DNA molecule, the method comprising providing to a cell: (a) a recombinant polypeptide, or one or more nucleic acid molecules encoding the recombinant polypeptide, comprising: (i) an amino acid sequence encoding a Cas12a nuclease; (ii) an amino acid sequence encoding a linker; and (iii) an amino acid sequence encoding a HUH endonuclease; (b) at least one guide nucleic acid, or at least one nucleic acid molecule encoding the at least one guide nucleic acid; and (c) at least one template nucleic acid molecule, or at least one nucleic acid molecule encoding the at least one template nucleic acid molecule; wherein the recombinant polypeptide, the at least one guide nucleic acid, and the at least one template nucleic acid form a ribonucleoprotein, and the ribonucleoprotein generates at least one edit in the target DNA molecule in the cell. In one aspect, the present disclosure provides a method for generating an edit in a target DNA molecule, comprising providing to a cell: (a) a recombinant polypeptide or one or more nucleic acid molecules encoding the recombinant polypeptide, the recombinant polypeptide comprising: (i) an amino acid sequence encoding a CasX nuclease; (ii) an amino acid sequence encoding a linker; and (iii) an amino acid sequence encoding a HUH endonuclease; (b) at least one guide nucleic acid or at least one nucleic acid molecule encoding at least one guide nucleic acid; and (c) at least one template nucleic acid molecule or at least one nucleic acid molecule encoding at least one template nucleic acid molecule, wherein the recombinant polypeptide, the at least one guide nucleic acid, and the at least one template nucleic acid form a ribonucleoprotein, and the ribonucleoprotein generates at least one edit in the target DNA molecule in the cell. In one aspect, the method provided herein further comprises (d) detecting at least one edit in the target DNA molecule. In one aspect, the ribonucleoprotein is formed intracellularly. In one aspect, the ribonucleoprotein is formed extracellularly. In one aspect, the ribonucleoprotein is formed in vivo. In one embodiment, the ribonucleoprotein is formed in vitro.
[0104] In one aspect, the editing provided herein includes mutations. As used herein, "mutation" refers to a non-naturally occurring change to a nucleic acid or amino acid sequence when compared with a naturally occurring reference nucleic acid or amino acid sequence from the same organism. When identifying a mutation, it is understood that the reference sequence must be from the same nucleic acid (e.g., gene, non-coding RNA) or amino acid (e.g., protein). When determining whether a difference between two sequences includes a mutation, it is understood in the art that comparison should not be made between homologous sequences from two different species, or between homologous sequences from two different species within a single species. Rather, a comparison should be made between the edited sequence and the endogenous, unedited (e.g., "wild-type") sequence of the same organism.
[0105] In one embodiment, the mutation comprises an insertion of at least one nucleotide or amino acid. In another embodiment, the mutation comprises a deletion of at least one nucleotide or amino acid. In a further embodiment, the mutation comprises a substitution of at least one nucleotide or amino acid. In a still further embodiment, the mutation comprises an inversion of at least two nucleotides or amino acids. In another embodiment, the mutation is selected from the group consisting of an insertion, a deletion, a substitution, and an inversion.
[0106] In one embodiment, the mutation comprises site-specific integration. In one embodiment, the site-specific integration comprises inserting all or part of a template nucleic acid molecule into a target DNA molecule.
[0107] As used herein, "site-specific integration" refers to all or a portion of a desired sequence (e.g., a template nucleic acid molecule) being inserted or integrated into a desired site or locus within a plant genome (e.g., a target DNA molecule). The desired sequence can comprise a transgene or construct. In one aspect, the template nucleic acid molecule comprises one or two homologous arms flanking the desired sequence to facilitate a targeted insertion event via homologous recombination and / or homology-directed repair.
[0108] Any site or locus within the genome of a plant, animal, fungus, or bacteria can potentially be selected for site-specific integration of a transgene or construct of the present disclosure.
[0109] For site-specific integration, a double-strand break (DSB) or nick can first be created in the target DNA molecule via the RNA-guided CRISPR nuclease or ribonucleoprotein provided herein. In the presence of a template nucleic acid molecule, the DSB or nick can then be repaired by homologous recombination (HR) or non-homologous end joining (NHEJ) between the homologous arm of the template nucleic acid molecule and the target DNA molecule, resulting in the site-specific integration of all or part of the template nucleic acid molecule into the target DNA molecule, creating a targeted insertion event at the site of the DSB or nick.
[0110] In one embodiment, the site-specific integration involves the use of an NHEJ repair mechanism endogenous to the cell. In another embodiment, the site-specific integration involves the use of an HR repair mechanism endogenous to the cell.
[0111] In one embodiment, the mutation comprises the incorporation of at least 5 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of at least 10 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of at least 15 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of at least 20 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of at least 25 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of at least 50 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of at least 100 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of at least 250 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of at least 500 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of at least 1000 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of at least 2000 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule.
[0112] In one embodiment, the mutation comprises the incorporation of between 5 and 3500 consecutive nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 5 and 2500 consecutive nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 5 and 1500 consecutive nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 5 and 750 consecutive nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 5 and 500 consecutive nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 5 and 250 consecutive nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 5 and 150 consecutive nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 25 and 2500 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 25 and 1500 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 25 and 750 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 50 and 2500 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 50 and 1500 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 50 and 750 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation of between 100 and 2500 contiguous nucleotides of the template nucleic acid molecule into the target DNA molecule. In one embodiment, the mutation comprises the incorporation into the target DNA molecule of between 100 and 1500 consecutive nucleotides of the template nucleic acid molecule. It involves the incorporation of between 00 and 750 consecutive nucleotides into the target DNA molecule.
[0113] In one aspect, the method provided herein includes detecting an edit or mutation in a target DNA molecule. Any method available in the art that can detect an edit or mutation in a target DNA molecule can be used. Suitable methods for detecting edits or mutations include, but are not limited to, Southern blot, polymerase chain reaction (PCR), and nucleic acid sequencing.
[0114] Any of the methods provided herein can involve transient transfection or stable transformation of a cell of interest (e.g., a eukaryotic cell, a prokaryotic cell). In one aspect, a nucleic acid molecule provided herein is stably transformed into a cell. In one aspect, a nucleic acid molecule provided herein is transiently transfected into a cell.
[0115] In one embodiment, the nucleic acid molecule encoding the Cas12a nuclease is stably transformed into the cell. In another embodiment, the nucleic acid molecule encoding the Cas12a nuclease is transiently transfected into the cell. In another embodiment, the Cas12a nuclease is transfected into the cell.
[0116] In one embodiment, the nucleic acid molecule encoding the CasX nuclease is stably transformed into the cell. In another embodiment, the nucleic acid molecule encoding the CasX nuclease is transiently transfected into the cell. In another embodiment, the CasX nuclease is transfected into the cell.
[0117] In one embodiment, the nucleic acid molecule encoding the linker is stably transformed into the cell. In another embodiment, the nucleic acid molecule encoding the linker is transiently transfected into the cell. In another embodiment, the linker is transfected into the cell.
[0118] In one embodiment, the nucleic acid molecule encoding the HUH endonuclease is stably transformed into the cell. In another embodiment, the nucleic acid molecule encoding the HUH endonuclease is transiently transfected into the cell. In another embodiment, the HUH endonuclease is transfected into the cell.
[0119] In one embodiment, the nucleic acid molecule encoding the FBNYV HUH endonuclease is stably transformed into the cell. In another embodiment, the nucleic acid molecule encoding the FBNYV HUH endonuclease is transiently transfected into the cell. In another embodiment, the FBNYV HUH endonuclease is transfected into the cell.
[0120] In one embodiment, the nucleic acid molecule encoding the PCV HUH endonuclease is stably transformed into the cell. In another embodiment, the nucleic acid molecule encoding the PCV HUH endonuclease is transiently transfected into the cell. In another embodiment, the PCV HUH endonuclease is transfected into the cell.
[0121] In one embodiment, the nucleic acid molecule encoding the guide nucleic acid is stably transformed into the cell. In another embodiment, the nucleic acid molecule encoding the guide nucleic acid is transiently transfected into the cell. In another embodiment, the guide nucleic acid is transfected into the cell.
[0122] In one embodiment, the nucleic acid molecule encoding the template nucleic acid molecule is stably transformed into the cell. In another embodiment, the nucleic acid molecule encoding the template nucleic acid molecule is transiently transfected into the cell. In another embodiment, the template nucleic acid molecule is transfected into the cell.
[0123] In one embodiment, a nucleic acid molecule encoding one or more components of a ribonucleoprotein is stably transformed into a cell. In another embodiment, a nucleic acid molecule encoding one or more components of a ribonucleoprotein is transiently transfected into a cell. In another embodiment, a ribonucleoprotein is transfected into a cell.
[0124] Numerous methods for transforming cells with recombinant nucleic acid molecules or constructs are known in the art and can be used in accordance with the methods of the present application. Any suitable method or technique for transforming cells known in the art can be used in accordance with the present methods. Effective methods for plant transformation include bacterial-mediated transformation, such as Agrobacterium-mediated transformation or Rhizobium-mediated transformation, and microprojectile bombardment-mediated transformation. Various methods are known in the art for regenerating or generating transgenic plants, such as by transforming explants with transformation vectors via bacterial-mediated transformation or microprojectile bombardment and then culturing the explants.
[0125] In one aspect, the method comprises providing a nucleic acid molecule, protein, or ribonucleoprotein to a cell via Agrobacterium-mediated transformation. In one aspect, the method comprises providing a nucleic acid molecule, protein, or ribonucleoprotein to a cell via polyethylene glycol-mediated transformation. In one aspect, the method comprises providing a nucleic acid molecule, protein, or ribonucleoprotein to a cell via biolistic transformation. In one aspect, the method comprises providing a nucleic acid molecule, protein, or ribonucleoprotein to a cell via liposome-mediated transfection. In one aspect, the method comprises providing a nucleic acid molecule, protein, or ribonucleoprotein to a cell via viral transduction. In one aspect, the method comprises providing a nucleic acid molecule, protein, or ribonucleoprotein to a cell via the use of one or more delivery particles. In one aspect, the method comprises providing a nucleic acid molecule, protein, or ribonucleoprotein to a cell via microinjection. In one aspect, the method comprises providing a nucleic acid molecule, protein, or ribonucleoprotein to a cell via electroporation.
[0126] In one aspect, the nucleic acid molecule is provided to the cell via a method selected from the group consisting of Agrobacterium-mediated transformation, polyethylene glycol-mediated transformation, biolistic transformation, liposome-mediated transfection, viral transduction, the use of one or more delivery particles, microinjection, and electroporation.
[0127] In one aspect, the protein is provided to the cell via a method selected from the group consisting of Agrobacterium-mediated transformation, polyethylene glycol-mediated transformation, biolistic transformation, liposome-mediated transfection, viral transduction, the use of one or more delivery particles, microinjection, and electroporation.
[0128] In one aspect, the ribonucleoprotein is provided to the cell via a method selected from the group consisting of Agrobacterium-mediated transformation, polyethylene glycol-mediated transformation, biolistic transformation, liposome-mediated transfection, viral transduction, the use of one or more delivery particles, microinjection, and electroporation.
[0129] Other methods for transformation, such as vacuum infiltration, pressure, sonication, and silicon carbide fiber agitation, are also known in the art and are contemplated for use with any of the methods provided herein.
[0130] Methods for transforming cells are well known to those skilled in the art. For example, specific instructions for transforming plant cells by microprojectile bombardment with recombinant DNA-coated particles (e.g., biolistic transformation) can be found in U.S. Patent Nos. 5,550,318, 5,538,880, 6,160,208, 6,399,861, and 6,153,812, and Agrobacterium-mediated transformation can be found in U.S. Patent Nos. 5,159,135, 5,824,877, 5,591,616, 6,384,301, 5,750,871, 5,463,174, and 5,188,958, all of which are incorporated herein by reference. Further methods for transforming plants can be found, for example, in Compendium of Transgenic Crop Plants (2009) Blackwell Publishing. Any suitable method known to those of skill in the art can be used to transform plant cells with any of the nucleic acid molecules provided herein.
[0131] Lipofection is described, for example, in U.S. Patent Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include those described in Felgner, WO 91 / 17424, WO 91 / 16024. Delivery can be to cells (e.g., in vitro or ex vivo administration) or to target tissues (e.g., in vivo administration).
[0132] Delivery vehicles, vectors, particles, nanoparticles, formulations, and components thereof for expressing one or more elements of a nucleic acid molecule or protein are used in WO 2014 / 093622 (PCT / US2013 / 074667). In one aspect, the method for providing a nucleic acid molecule or protein to a cell comprises delivery via a delivery particle. In one aspect, the method for providing a nucleic acid molecule or protein to a cell comprises delivery via a delivery vesicle. In one aspect, the delivery vesicle is selected from the group consisting of an exosome and a liposome. In one aspect, the method for providing a nucleic acid molecule or protein to a cell comprises delivery via a viral vector. In one aspect, the viral vector is selected from the group consisting of an adenoviral vector, a lentiviral vector, and an adeno-associated viral vector. In another aspect, the method for providing a nucleic acid molecule or protein to a cell comprises delivery via a nanoparticle. In one aspect, the method for providing a nucleic acid molecule or protein to a cell comprises microinjection. In one aspect, the method for providing a nucleic acid molecule or protein to a cell comprises a polycation. In one aspect, the method for providing a nucleic acid molecule or protein to a cell includes a cationic oligopeptide.
[0133] In one aspect, the delivery particle is selected from the group consisting of an exosome, an adenoviral vector, a lentiviral vector, an adeno-associated viral vector, a nanoparticle, a polycation, and a cationic oligopeptide. In one aspect, the method provided herein comprises the use of one or more delivery particles. In another aspect, the method provided herein comprises the use of two or more delivery particles. In another aspect, the method provided herein comprises the use of three or more delivery particles.
[0134] Agents suitable for facilitating the transfer of proteins, nucleic acids, mutagens, and ribonucleoproteins into plant cells include agents that increase the permeability of the outside of the plant cell or agents that increase the permeability of the plant cell to oligonucleotides, polynucleotides, proteins, or ribonucleoproteins. Such agents that facilitate the transfer of compositions into plant cells include chemical agents, physical agents, or combinations thereof. Chemical agents for conditioning include (a) detergents, (b) organic solvents or aqueous solutions or aqueous mixtures of organic solvents, (c) oxidizing agents, (e) acids, (f) bases, (g) oils, (h) enzymes, or combinations thereof.
[0135] Organic solvents useful for conditioning plants to be permeabilized with polynucleotides include DMSO, DMF, pyridine, N-pyrrolidine, hexamethylphosphoramide, acetonitrile, dioxane, polypropylene glycol, and other solvents that are miscible with water or that dissolve phosphonucleotides in non-aqueous systems (such as those used in synthetic reactions). Naturally occurring or synthetic oils, with or without surfactants or emulsifiers, can be used, such as oils from plant sources, crop oils (e.g., those listed in the 9th Compendium of Herbicide Adjuvants, publicly available online at www.herbicide.adjuvants.com), paraffin oil, polyol fatty acid esters, or oils with short-chain molecules modified with amides or polyamines, such as polyethyleneimine or N-pyrrolidine.
[0136] Examples of useful surfactants include sodium or lithium salts of fatty acids (such as tallow or tallowamine or phospholipids) and organosilicone surfactants. Other useful surfactants include organosilicone surfactants, such as nonionic organosilicone surfactants, such as trisiloxane ethoxylate surfactants or silicone polyether copolymers, such as a copolymer of polyalkylene oxide-modified heptamethyltrisiloxane and allyloxypolypropylene glycol methyl ether (commercially available as Silwet® L-77).
[0137] Useful physical agents can include (a) abrasives, such as carborundum, corundum, sand, calcite, pumice, garnet, etc.; (b) nanoparticles, such as carbon nanotubes; or (c) physical forces. Carbon nanotubes are disclosed by Kam et al. (2004) / . Am. Chem. Soc., 126 (22):6850-6851; Liu et al. (2009) Nano Lett., 9(3):1007-1010; and Khodakovskaya et al. (2009) ACS Nano, 3(10):3221-3227. Physical forces can include heating, cooling, application of positive pressure, or ultrasonic treatment. Embodiments of the present methods can optionally include an incubation step, a neutralization step (e.g., to neutralize acids, bases, or oxidizing agents, or to inactivate enzymes), a rinsing step, or a combination thereof. The methods of the present invention can further include the application of other agents that have enhanced effects due to the silencing of specific genes. For example, if a polynucleotide is designed to regulate a gene that confers herbicide resistance, subsequent application of the herbicide can dramatically affect the efficacy of the herbicide.
[0138] Agents for laboratory acclimation of plant cells to permeabilization with polynucleotides include, for example, application of chemical agents, enzyme treatment, heat or cold, positive or negative pressure treatment, or sonication. Agents for acclimating plants in the field include chemical agents such as surfactants and salts.
[0139] In one aspect, the transformed or transfected cell is a prokaryotic cell. In another aspect, the transformed or transfected cell is a eukaryotic cell. In another aspect, the transformed or transfected cell is a plant cell. In another aspect, the transformed or transfected cell is an animal cell. In another aspect, the transformed or transfected cell is a fungal cell.
[0140] Recipient plant cells or explant targets for transformation include, but are not limited to, seed cells, fruit cells, leaf cells, cotyledon cells, hypocotyl cells, meristem cells, embryo cells, endosperm cells, root cells, shoot cells, stem cells, sheath cells, flower cells, inflorescence cells, stem cells, pedicel cells, style cells, stigma cells, receptacle cells, petal cells, sepal cells, pollen cells, anther cells, filament cells, ovary cells, ovule cells, pericarp cells, phloem cells, bud cells, or vascular tissue cells. In another aspect, the present disclosure provides plant chloroplasts. In a further aspect, the present disclosure provides epidermal cells, stomatal cells, trichome cells, root hair cells, storage root cells, or tuber cells. In another aspect, the present disclosure provides protoplasts. In another aspect, the present disclosure provides plant callus cells. Any cell from which a fertilized plant can be regenerated is contemplated as a useful recipient cell for the practice of the present disclosure. Callus can be initiated from a variety of tissue sources, including, but not limited to, immature embryos or parts of embryos, seedling apical meristems, microspores, etc. Those cells capable of growing as callus can serve as recipient cells for transformation. Actual transformation methods and materials for producing the transgenic plants of the present disclosure (e.g., various media and recipient target cells, transformation of immature embryos, and subsequent regeneration of fertilized transgenic plants) are disclosed, for example, in U.S. Pat. Nos. 6,194,636 and 6,232,526 and U.S. Patent Application Publication No. 2004 / 0216189, all of which are incorporated herein by reference. Transformed explants, cells, or tissues can be subjected to additional culture steps, such as callus induction, selection, and regeneration, as known in the art. Transformed cells, tissues, or explants containing recombinant DNA inserts can be propagated, developed, or regenerated into transgenic plants in culture, plugs, or soil according to methods known in the art. In one aspect, the present disclosure provides plant cells that are not propagation material and do not mediate natural propagation of plants. In another aspect, the present disclosure also provides plant cells that are propagation material and mediate natural propagation of plants. In another aspect, the present disclosure provides plant cells that cannot sustain themselves through photosynthesis. In another aspect, the present disclosure provides somatic plant cells.Somatic cells, in contrast to germline cells, do not mediate plant reproduction. In one aspect, the present disclosure provides non-reproducing plant cells.
[0141] In one aspect, any nucleic acid molecule, polypeptide, or ribonucleoprotein provided herein is in a cell. In another aspect, any nucleic acid molecule, polypeptide, or ribonucleoprotein provided herein is in a prokaryotic cell. In one aspect, any nucleic acid molecule, polypeptide, or ribonucleoprotein provided herein is in a eukaryotic cell.
[0142] In one aspect, any cell provided herein is a host cell. In one aspect, the host cell comprises any ribonucleoprotein provided herein. In one aspect, the host cell comprises any polypeptide provided herein. In one aspect, the host cell comprises any nucleic acid molecule provided herein.
[0143] In one aspect, the host cell is selected from the group consisting of a plant cell, a bacterial cell, a mammalian cell, a fungal cell, an insect cell, an arachnid cell, an avian cell, a fish cell, a reptilian cell, and an amphibian cell. In another aspect, the host plant cell is selected from the group consisting of a corn cell, a soybean cell, a cotton cell, a canola cell, a rice cell, a wheat cell, a sorghum cell, an alfalfa cell, a sugarcane cell, a millet cell, a tomato cell, a potato cell, and an algae cell. In another aspect, the host cell is an E. coli cell.
[0144] In one embodiment, the prokaryotic cell is selected from the group consisting of Acidobacteria, Actinobacteria, Aquificae, Armatimonadetes, Bacteroidetes, Caldiserica, Chlamydie, Chlorobi, Chloroflexi, Chrysiogenetes, Coprothermobacterota, Cyanobacteria, Deferribacteres, Deinococcus-Thermus, Dictyoglomi, Yersinia, and the like. In another embodiment, the prokaryotic cell is a cell derived from a phylum selected from the group consisting of Elusimicrobia, Fibrobacteres, Firmicutes, Fusobacteria, Gemmatimonadetes, Lentisphaerae, Nitrospirae, Planctomycetes, Proteobacteria, Spirochaetes, Synergistetes, Tenericutes, Thermodesulfobacteria, Thermotogae, and Verrucomicrobia. In another embodiment, the prokaryotic cell is an Escherichia coli cell. In another aspect, the prokaryotic cell is selected from a genus selected from the group consisting of Escherichia, Agrobacterium, Rhizobium, Sinorhizobium, and Staphylococcus.
[0145] In one embodiment, the eukaryotic cell is an ex vivo cell. In another embodiment, the eukaryotic cell is a plant cell. In another embodiment, the eukaryotic cell is a plant cell in culture. In another embodiment, the eukaryotic cell is an angiosperm cell. In another embodiment, the eukaryotic cell is a gymnosperm cell. In another embodiment, the eukaryotic cell is a monocotyledonous plant cell. In another embodiment, the eukaryotic cell is a dicotyledonous plant cell. In another embodiment, the eukaryotic cell is a corn cell. In another embodiment, the eukaryotic cell is a rice cell. In another embodiment, the eukaryotic cell is a sorghum cell. In another embodiment, the eukaryotic cell is a wheat cell. In another embodiment, the eukaryotic cell is a canola cell. In another embodiment, the eukaryotic cell is an alfalfa cell. In another embodiment, the eukaryotic cell is a soybean cell. In another embodiment, the eukaryotic cell is a cotton cell. In another embodiment, the eukaryotic cell is a tomato cell. In another embodiment, the eukaryotic cell is a potato cell. In a further embodiment, the eukaryotic cell is a cucumber cell. In another embodiment, the eukaryotic cell is a millet cell. In another embodiment, the eukaryotic cell is a barley cell. In another embodiment, the eukaryotic cell is a rapeseed cell. In another embodiment, the eukaryotic cell is a grass cell. In another embodiment, the eukaryotic cell is a Setaria cell. In another embodiment, the eukaryotic cell is an Arabidopsis cell. In a further embodiment, the eukaryotic cell is an algae cell.
[0146] In one embodiment, the plant cell is an epidermal cell. In another embodiment, the plant cell is a stomatal cell. In another embodiment, the plant cell is a trichome cell. In another embodiment, the plant cell is a root cell. In another embodiment, the plant cell is a leaf cell. In another embodiment, the plant cell is a callus cell. In another embodiment, the plant cell is a protoplast cell. In another embodiment, the plant cell is a pollen cell. In another embodiment, the plant cell is an ovary cell. In another embodiment, the plant cell is a flower cell. In another embodiment, the plant cell is a meristem cell. In another embodiment, the plant cell is an endosperm cell. In another embodiment, the plant cell does not contain propagation material and does not mediate natural reproduction of the plant. In another embodiment, the plant cell is a somatic plant cell.
[0147] Additional provided plant cells, tissues and organs can be derived from seeds, fruits, leaves, cotyledons, hypocotyls, meristems, embryos, endosperms, roots, shoots, stems, pods, flowers, inflorescences, stems, pedicels, styles, stigmas, receptacles, petals, sepals, pollen, anthers, filaments, ovaries, ovules, pericarp, phloem, and vascular tissue.
[0148] In a further embodiment, the eukaryotic cell is an animal cell. In another embodiment, the eukaryotic cell is an animal cell in culture. In a further embodiment, the eukaryotic cell is a human cell. In another embodiment, the eukaryotic cell is not a human stem cell. In a further embodiment, the eukaryotic cell is a human cell in culture. In a further embodiment, the eukaryotic cell is a somatic human cell. In a further embodiment, the eukaryotic cell is a cancer cell. In a further embodiment, the eukaryotic cell is a mammalian cell. In a further embodiment, the eukaryotic cell is a mouse cell. In a further embodiment, the eukaryotic cell is a porcine cell. In a further embodiment, the eukaryotic cell is a bovine cell. In a further embodiment, the eukaryotic cell is an avian cell. In a further embodiment, the eukaryotic cell is a reptilian cell. In a further embodiment, the eukaryotic cell is an amphibian cell. In a further embodiment, the eukaryotic cell is an insect cell. In a further embodiment, the eukaryotic cell is an arthropod cell. In a further embodiment, the eukaryotic cell is a cephalopod cell. In a further aspect, the eukaryotic cell is an arachnid cell. In a further aspect, the eukaryotic cell is a mollusk cell. In a further aspect, the eukaryotic cell is a nematode cell. In a further aspect, the eukaryotic cell is a fish cell.
[0149] In another embodiment, the eukaryotic cell is a protozoan cell. In another embodiment, the eukaryotic cell is a fungal cell. In one embodiment, the fungal cell is a yeast cell. In one embodiment, the yeast cell is a Schizosaccharomyces pombe cell. In another embodiment, the yeast cell is a Saccharomyces cerevisiae cell.
[0150] The use of the terms "polynucleotide" or "nucleic acid molecule" is not intended to limit the present disclosure to polynucleotides comprising deoxyribonucleic acid (DNA). For example, ribonucleic acid (RNA) molecules are also contemplated. Those skilled in the art will recognize that polynucleotides and nucleic acid molecules can comprise deoxyribonucleotides, ribonucleotides, or a combination of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogues. Polynucleotides of the present disclosure also encompass all forms of sequences, including, but not limited to, single-stranded forms, double-stranded forms, hairpins, stem-and-loop structures, and the like. In one aspect, the nucleic acid molecules provided herein are DNA molecules. In another aspect, the nucleic acid molecules provided herein are RNA molecules. In one aspect, the nucleic acid molecules provided herein are single-stranded. In another aspect, the nucleic acid molecules provided herein are double-stranded.
[0151] As used herein, the term "recombinant," with respect to nucleic acid (DNA or RNA) molecules, proteins, constructs, vectors, etc., refers to a nucleic acid or amino acid molecule or amino acid sequence that is artificial and not normally found in nature and / or is present in a context not normally found in nature, including nucleic acid (DNA or RNA) molecules, proteins, constructs, etc. that contain combinations of polynucleotide or protein sequences that are not found contiguous or close together in nature without human intervention, and / or polynucleotide molecules, proteins, constructs, etc. that contain at least two polynucleotide or protein sequences that are heterologous to each other.
[0152] In one aspect, the methods and compositions provided herein include vectors. As used herein, the terms "vector" and "plasmid" are used interchangeably and refer to circular double-stranded DNA molecules that are physically separated from chromosomal DNA. In one aspect, the plasmid or vector used herein can replicate in vivo.
[0153] In one aspect, the present disclosure provides a plasmid comprising any of the nucleic acid molecules provided herein. In another aspect, the present disclosure provides a plasmid encoding any of the amino acid sequences provided herein. In another aspect, the present disclosure provides a plasmid encoding any of the ribonucleoproteins provided herein.
[0154] In another embodiment, a nucleic acid encoding a Cas12a nuclease or a CasX nuclease is provided in the vector. In a further embodiment, a nucleic acid encoding a linker is provided in the vector. In a further embodiment, a nucleic acid encoding a HUH endonuclease is provided in the vector. In a further embodiment, a nucleic acid encoding a guide nucleic acid is provided in the vector. In a further embodiment, a nucleic acid encoding a template nucleic acid molecule is provided in the vector.
[0155] As used herein, the term "polypeptide" refers to a chain of at least two covalently linked amino acids. Polypeptides can be encoded by the polynucleotides provided herein. An example of a polypeptide is a protein. Proteins provided herein can be encoded by the nucleic acid molecules provided herein.
[0156] Nucleic acids can be isolated using routine techniques in the art. For example, nucleic acids can be isolated using any method, including, but not limited to, recombinant nucleic acid technology and / or polymerase chain reaction (PCR). General PCR techniques are described, for example, in "PCR Primer: A Laboratory Manual," edited by Dieffenbach and Dveksler, Cold Spring Harbor Laboratory Press, 1995. Recombinant nucleic acid technology includes, for example, restriction enzyme digestion and ligation, which can be used to isolate nucleic acids. Isolated nucleic acids can also be chemically synthesized as single nucleic acid molecules or as a series of oligonucleotides. Polypeptides can be purified from natural sources (e.g., biological samples) by known methods, such as DEAE ion exchange, gel filtration, and hydroxyapatite chromatography. Polypeptides can also be purified, for example, by expressing nucleic acids in expression vectors. Additionally, purified polypeptides can be obtained by chemical synthesis. The degree of purity of a polypeptide can be measured using any appropriate method, such as column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.
[0157] Without limitation, nucleic acids can be detected using hybridization. Hybridization between nucleic acids is discussed in detail in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).
[0158] Polypeptides can be detected using antibodies. Techniques for detecting polypeptides using antibodies include enzyme-linked immunosorbent assay (ELISA), Western blot, immunoprecipitation, and immunofluorescence. The antibodies provided herein can be polyclonal or monoclonal. Antibodies having specific binding affinity for the polypeptides provided herein can be generated using methods well known in the art. The antibodies provided herein can be bound to a solid support, such as a microtiter plate, using methods well known in the art.
[0159] The term "percent identity" or "percent identical," as used herein with respect to two or more nucleotide or protein sequences, is calculated by (i) comparing two optimally aligned sequences (nucleotide or protein) over a comparison window, (ii) determining the number of positions where the same nucleic acid base (for nucleotide sequences) or amino acid residue (for proteins) occurs in both sequences to yield the number of matched positions, (iii) dividing the number of matched positions by the total number of positions within the comparison window, and then (iv) multiplying this quotient by 100% to yield the percent identity. When calculating "percent identity" with respect to a reference sequence where a specific comparison window is not designated, the percent identity is determined by dividing the number of matched positions over the region of alignment by the total length of the reference sequence. Thus, for purposes of this application, when two sequences (query and subject) are optimally aligned (allowing for gaps in their alignment), the "percent identity" of a query sequence is equal to the number of identical positions between the two sequences divided by the total number of positions of the query sequence over its length (or comparison window), then multiplied by 100%. When percentages of sequence identity are used in the context of proteins, it is recognized that non-identical residue positions often differ by conservative amino acid substitutions, in which amino acid residues are substituted with other amino acid residues having similar chemical properties (e.g., charge or hydrophobicity), thus not altering the functional properties of the molecule. Where conservative substitutions vary between sequences, the percent sequence identity can be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity."
[0160] The terms "percent sequence complementarity" or "percent complementarity," as used herein in reference to two nucleotide sequences, are similar to the concept of percent identity, but refer to the percentage of nucleotides in a query sequence that optimally base-pair or hybridize with nucleotides in a subject sequence when the query and subject sequences are linearly arranged and optimally base-paired, without secondary folds such as loops, stems, or hairpins. Such percent complementarity can be between two DNA strands, two RNA strands, or one DNA and one RNA strand. "Percent complementarity" can be calculated by (i) optimally base-pairing or hybridizing two nucleotide sequences in a linear, fully extended configuration (without folding or secondary structure) over the comparison window, (ii) determining the number of base-pairing positions between the two sequences over the comparison window to obtain the number of complementary positions, (iii) dividing the number of complementary positions by the total number of positions in the comparison window, and (iv) multiplying this quotient by 100% to obtain the percent complementarity of the two sequences. Optimal base pairing of two sequences can be determined based on known pairing of nucleotide bases such as GC, AT, and AU through hydrogen bonding. When calculating "percent complementarity" with respect to a reference sequence without specifying a specific comparison window, the percent identity is determined by dividing the number of complementary positions between the two linear sequences by the total length of the reference sequence. Thus, for the purposes of this application, when two sequences (query and subject) are optimally base-paired (allowing for mismatched or non-base-paired nucleotides), the "percent complementarity" of a query sequence is equal to the number of base-paired positions between the two sequences divided by the total number of positions of the query sequence over its length, then multiplied by 100%.
[0161] Various pairwise or multiple sequence alignment algorithms and programs are known in the art for optimal alignment of sequences to calculate their percent identity, such as ClustalW or basic local alignment search tool (BLAST®), which can be used to compare sequence identity or similarity between two or more nucleotide or protein sequences. Although other alignment and comparison methods are known in the art, the alignment and percent identity between two sequences (including the percent identity ranges noted above) can be determined by the ClustalW algorithm. See, for example, Chenna R. et al., "Multiple sequence alignment with the Clustal series of programs," Nucleic Acids Research 31: 3497-3500 (2003); Thompson JD et al., "Clustal W: Improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice," Nucleic Acids Research 22: 4673-4680 (1994); Larkin MA et al., "Clustal W and Clustal X version 2.0," Bioinformatics 23: 2947-48 (2007); and Altschul, S. F., Gish, W., Miller, W., Myers, E. W., and Lipman, D. J. (1990) "Basic local alignment search tool," J. Mol. Biol. 215: 403-410, the entire contents and disclosures of which are incorporated herein by reference. (1990).
[0162] As used herein, a first nucleic acid molecule can "hybridize" to a second nucleic acid molecule through non-covalent interactions (e.g., Watson-Crick base pairing) in a sequence-specific antiparallel manner (i.e., the nucleic acid specifically binds to a complementary nucleic acid) under appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength. As is known in the art, standard Watson-Crick base pairing includes adenine pairing with thymine, adenine pairing with uracil, and guanine (G) pairing with cytosine (C) [DNA, RNA]. Furthermore, it is also known in the art that guanine bases pair with uracil for hybridization between two RNA molecules (e.g., dsRNA). For example, G / U base pairing is partially responsible for the degeneracy (redundancy) of the genetic code in the context of base pairing between the anticodon of a tRNA and the codon of an mRNA. In the context of this disclosure, a guanine in the protein-binding segment (dsRNA duplex) of a subject DNA-targeting RNA molecule is considered complementary to a uracil, and vice versa. Thus, if a G / U base pair can be made in the protein-binding segment (dsRNA duplex) of a subject DNA-targeting RNA molecule at a given nucleotide position, that position is not considered non-complementary, but instead is considered complementary.
[0163] Hybridization and washing conditions are well known and are exemplified in Sambrook, J., Fritsch, E. F. and Maniatis, T. Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), particularly Chapter 11 and Table 11.1 therein; and Sambrook, J. and Russell, W., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor (2001).
[0164] Hybridization requires that the two nucleic acids contain complementary sequences, although mismatches between bases are possible. Suitable conditions for hybridization between two nucleic acids depend on the length of the nucleic acids and the degree of complementarity, variables well known in the art. The greater the degree of complementarity between two nucleotide sequences, the greater the melting temperature (Tm) of hybrids of nucleic acids having those sequences. For hybridization between nucleic acids with short stretches of complementarity (e.g., complementarity of 35 nucleotides or less), the position of mismatches becomes important (see Sambrook et al.). Typically, the length of a hybridizable nucleic acid is at least 10 nucleotides. Exemplary minimum lengths for hybridizable nucleic acids are at least 15 nucleotides; at least 18 nucleotides; at least 20 nucleotides; at least 22 nucleotides; at least 25 nucleotides; and at least 30 nucleotides. Furthermore, those skilled in the art will recognize that temperature and wash solution salt concentration can be adjusted as necessary depending on factors such as the length of the complementary region and the degree of complementarity.
[0165] It is understood in the art that the sequence of a polynucleotide does not need to be 100% complementary to the sequence of its target nucleic acid to be specifically hybridizable or hybridizable.Furthermore, a polynucleotide can hybridize in one or more segments so that intervening or adjacent segments are not involved in the hybridization event (for example, a loop structure or a hairpin structure).For example, an antisense nucleic acid in which 18 nucleotides out of 20 nucleotides of an antisense compound are complementary to a target region and therefore specifically hybridizes represents 90% complementarity.In this example, the remaining non-complementary nucleotides can be clustered or interspersed with complementary nucleotides, and do not need to be adjacent to each other or to complementary nucleotides. The percent complementarity between specific stretches of nucleic acid sequences within a nucleic acid can be routinely determined using the BLAST® program (Basic Local Alignment Search Tool) and PowerBLAST programs (see Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656), which are known in the art, by using the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), using the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489), using default settings.
[0166] As used herein, "in vivo" refers to within a living cell, tissue, or organism. As used herein, "in vitro" refers to within laboratory equipment. Non-limiting examples of laboratory equipment include test tubes, flasks, beakers, graduated cylinders, pipettes, Petri dishes, and microtiter plates. As used herein, "ex vivo" refers to within cells or tissues from an organism in an external environment. "Ex planta" refers to plant cells or tissues in an external environment, while "in planta" refers to cells within a living plant. As a non-limiting example, plant protoplasts in a Petri dish or test tube are considered both ex vivo and ex planta.
[0167] In one aspect, any nucleic acid molecule or polypeptide provided herein can be used in vivo. In one aspect, any nucleic acid molecule or polypeptide provided herein can be used in vitro. In one aspect, any nucleic acid molecule or polypeptide provided herein can be used ex vivo. In one aspect, any nucleic acid molecule or polypeptide provided herein can be used in a plant body. In one aspect, any nucleic acid molecule or polypeptide provided herein can be used outside a plant body.
[0168] In one aspect, any cell provided herein is present in vivo. In one aspect, any cell provided herein is present in vitro. In one aspect, any cell provided herein is present ex vivo. In one aspect, any plant cell provided herein is present within a plant. In one aspect, any plant cell provided herein is present outside a plant.
[0169] As commonly understood in the art, the term "promoter" refers to a DNA sequence that contains an RNA polymerase binding site, a transcription initiation site, and / or a TATA box and that supports or facilitates the transcription and expression of an associated transcribable polynucleotide sequence and / or gene (or transgene). Promoters can be synthetically produced, altered, or derived from known or naturally occurring promoter sequences or other promoter sequences. Promoters can also include chimeric promoters, which contain a combination of two or more heterologous sequences. Thus, promoters of the present application can include variants of promoter sequences that are similar in composition but not identical to other promoter sequences known or provided herein. Promoters can be classified according to various criteria related to the expression pattern of the associated coding or transcribable sequence or gene (including transgene) operably linked to the promoter, such as constitutive, developmental, tissue-specific, inducible, etc.
[0170] In one embodiment, the recombinant nucleic acid provided herein comprises at least one promoter. In another embodiment, the polynucleotide encoding the Cas12a nuclease is operably linked to at least one promoter. In another embodiment, the polynucleotide encoding the CasX nuclease is operably linked to at least one promoter. In another embodiment, the polynucleotide encoding the HUH endonuclease is operably linked to at least one promoter. In another embodiment, the polynucleotide encoding the FBNYV HUH endonuclease is operably linked to at least one promoter. In another embodiment, the polynucleotide encoding the PCV HUH endonuclease is operably linked to at least one promoter. In another embodiment, the polynucleotide encoding the guide nucleic acid is operably linked to at least one promoter. In another embodiment, the polynucleotide encoding the template nucleic acid molecule is operably linked to at least one promoter.
[0171] A promoter that drives expression in all or most tissues of a plant is called a "constitutive" promoter. A promoter that drives expression during a certain period or developmental stage is called a "developmental" promoter. A promoter that drives enhanced expression in certain tissues of an organism compared to other tissues of the organism is called a "tissue-preferred" promoter. Thus, a "tissue-preferred" promoter causes relatively high or preferential expression in certain tissues of a plant, but low expression levels in other tissues of the plant. A promoter that expresses in certain tissues of an organism, with little or no expression in other tissues, is called a "tissue-specific" promoter. An "inducible" promoter is a promoter that initiates transcription in response to environmental stimuli such as heat, cold, drought, light, or other stimuli such as wounding or chemical application. Promoters can also be classified in terms of their origin, for example, heterologous, homologous, chimeric, synthetic, etc.
[0172] As used herein, the term "heterologous" refers to a promoter sequence that has a different origin to its associated transcribable DNA sequence, coding sequence, or gene (or transgene) and / or does not naturally occur in the plant species being transformed. The term "heterologous" can refer more broadly to the combination of two or more DNA molecules or sequences, such as a promoter and an associated transcribable DNA sequence, coding sequence, or gene, when such a combination is created artificially and is not normally found in nature.
[0173] In one aspect, the promoters provided herein are constitutive promoters. In another aspect, the promoters provided herein are tissue-specific promoters. In a further aspect, the promoters provided herein are tissue-preferred promoters. In yet another aspect, the promoters provided herein are inducible promoters. In one aspect, the promoters provided herein are selected from the group consisting of constitutive promoters, tissue-specific promoters, tissue-preferred promoters, and inducible promoters.
[0174] An RNA polymerase III (Pol III) promoter can be used to drive the expression of non-protein-encoding RNA molecules. In one embodiment, the promoter provided herein is a Pol III promoter. In another embodiment, the Pol III promoter provided herein is operably linked to a nucleic acid molecule encoding a non-protein-encoding RNA. In yet another embodiment, the Pol III promoter provided herein is operably linked to a nucleic acid molecule encoding a guide nucleic acid. In yet another embodiment, the Pol III promoter provided herein is operably linked to a nucleic acid molecule encoding a single guide RNA. In a further embodiment, the Pol III promoter provided herein is operably linked to a nucleic acid molecule encoding a CRISPR RNA (crRNA). In another embodiment, the Pol III promoter provided herein is operably linked to a nucleic acid molecule encoding a tracer RNA (tracrRNA).
[0175] Non-limiting examples of Pol III promoters include the U6 promoter, the H1 promoter, the 5S promoter, the adenovirus 2 (Ad2) VAI promoter, the tRNA promoter, and the 7SK promoter. See, e.g., Schramm and Hernandez, 2002, Genes & Development, 16:2593-2620, which is incorporated herein by reference in its entirety. In one aspect, the Pol III promoter provided herein is selected from the group consisting of the U6 promoter, the H1 promoter, the 5S promoter, the adenovirus 2 (Ad2) VAI promoter, the tRNA promoter, and the 7SK promoter. In another aspect, the guide RNA provided herein is operably linked to a promoter selected from the group consisting of the U6 promoter, the H1 promoter, the 5S promoter, the adenovirus 2 (Ad2) VAI promoter, the tRNA promoter, and the 7SK promoter. In another aspect, the single guide RNA provided herein is operably linked to a promoter selected from the group consisting of a U6 promoter, an H1 promoter, a 5S promoter, an adenovirus 2 (Ad2) VAI promoter, a tRNA promoter, and a 7SK promoter. In another aspect, the CRISPR RNA provided herein is operably linked to a promoter selected from the group consisting of a U6 promoter, an H1 promoter, a 5S promoter, an adenovirus 2 (Ad2) VAI promoter, a tRNA promoter, and a 7SK promoter. In another aspect, the tracer RNA provided herein is operably linked to a promoter selected from the group consisting of a U6 promoter, an H1 promoter, a 5S promoter, an adenovirus 2 (Ad2) VAI promoter, a tRNA promoter, and a 7SK promoter.
[0176] In one embodiment, the promoter provided herein is a dahlia mosaic virus (DaMV) promoter. In another embodiment, the promoter provided herein is a U6 promoter. In another embodiment, the promoter provided herein is an actin promoter.
[0177] Examples describing promoters that can be used herein include, but are not limited to, U.S. Pat. No. 6,437,217 (maize RS81 promoter), U.S. Pat. No. 5,641,876 (rice actin promoter), U.S. Pat. No. 6,426,446 (maize RS324 promoter), U.S. Pat. No. 6,429,362 (maize PR-1 promoter), U.S. Pat. No. 6,232,526 (maize A3 promoter), U.S. Pat. No. 6,177,611 (constitutive maize promoters), U.S. Pat. Nos. 5,322,938, 5,352,605, 5,359,142, and 5,530,196 (35S promoter), U.S. Pat. No. 6,433,252 (maize L3 oleosin promoter), U.S. Pat. No. 6,429,357 (rice actin 2 promoter and rice actin 2 intron), U.S. Pat. No. 5,837,848 (root-specific promoter), U.S. Pat. No. 6,294,714 (light-inducible promoter), U.S. Pat. No. 6,140,078 (salt-inducible promoter), U.S. Pat. No. 6,252,138 (pathogen-inducible promoter), U.S. Pat. No. 6,175,060 (phosphorus deficiency-inducible promoter), U.S. Pat. No. 6,635,806 (gamma coixin promoter), and U.S. Patent Application No. 09 / 757,089 (maize chloroplast aldolase promoter).Additional promoters that may find use include the nopaline synthase (NOS) promoter (Ebert et al., 1987), the octopine synthase (OCS) promoter (carried on the tumor-inducing plasmid of Agrobacterium tumefaciens), caulimovirus promoters such as the cauliflower mosaic virus (CaMV) 19S promoter (Lawton et al., Plant Molecular Biology (1987) 9: 315-324), the CaMV 35S promoter (Odell et al., Nature (1985) 313: 810-812), the figwort mosaic virus 35S promoter (U.S. Patent Nos. 6,051,753; 5,378,619), the sucrose synthase promoter (Yang and Russell, Proceedings of the National Academy of Sciences, USA (1990) 87: 4144-4148), the R gene complex promoter (Chandler et al., Plant Cell (1989) 1: 1175-1183), and the chlorophyll a / b binding protein gene promoters, PC1SV (U.S. Patent No. 5,850,019), and AGRtu.nos (GenBank accession number V00087; Depicker et al., Journal of Molecular and Applied Genetics (1982) 1: 561-573; Bevan et al., 1983) promoters.
[0178] Promoter hybrids can also be used and constructed to enhance transcriptional activity (see U.S. Pat. No. 5,106,739) or combine desired transcriptional activity, inducibility, and tissue or developmental specificity. Promoters that function in plants include, but are not limited to, inducible, viral, synthetic, constitutive, temporally regulated, spatially regulated, and spatiotemporally regulated promoters. Other tissue-enhanced, tissue-specific, or developmentally regulated promoters are also known in the art and are contemplated as useful in practicing the present disclosure.
[0179] It is recognized in the art that a fragment of a promoter sequence can function to drive the transcription of an operably linked nucleic acid molecule.For example, but not limited to, if a 1000bp promoter is cut into 500bp, and the 500bp fragment can drive transcription, the 500bp fragment is called a "functional fragment".
[0180] As used herein, " nuclear localization signal " (NLS) refers to the amino acid sequence that " tags " protein for import into the nucleus of cell.In one aspect, the nucleic acid molecule provided herein encodes a nuclear localization signal.In another aspect, the nucleic acid molecule provided herein encodes two or more nuclear localization signals.
[0181] In one embodiment, the Cas12a nuclease provided herein comprises a nuclear localization signal. In one embodiment, the nuclear localization signal is located at the N-terminus of the Cas12a nuclease. In a further embodiment, the nuclear localization signal is located at the C-terminus of the Cas12a nuclease. In yet another embodiment, the nuclear localization signal is located at both the N-terminus and the C-terminus of the Cas12a nuclease.
[0182] In one embodiment, the CasX nucleases provided herein comprise a nuclear localization signal. In one embodiment, the nuclear localization signal is located at the N-terminus of the CasX nuclease. In a further embodiment, the nuclear localization signal is located at the C-terminus of the CasX nuclease. In yet another embodiment, the nuclear localization signal is located at both the N-terminus and the C-terminus of the CasX nuclease.
[0183] In one embodiment, the HUH endonuclease provided herein comprises a nuclear localization signal. In one embodiment, the nuclear localization signal is located at the N-terminus of the HUH endonuclease. In a further embodiment, the nuclear localization signal is located at the C-terminus of the HUH endonuclease. In yet another embodiment, the nuclear localization signal is located at both the N-terminus and the C-terminus of the HUH endonuclease.
[0184] In one embodiment, the FBNYV HUH endonuclease provided herein comprises a nuclear localization signal. In one embodiment, the nuclear localization signal is located at the N-terminus of the FBNYV HUH endonuclease. In a further embodiment, the nuclear localization signal is located at the C-terminus of the FBNYV HUH endonuclease. In yet another embodiment, the nuclear localization signal is located at both the N-terminus and the C-terminus of the FBNYV HUH endonuclease.
[0185] In one embodiment, the PCV HUH endonuclease provided herein comprises a nuclear localization signal. In one embodiment, the nuclear localization signal is located at the N-terminus of the PCV HUH endonuclease. In a further embodiment, the nuclear localization signal is located at the C-terminus of the PCV HUH endonuclease. In yet another embodiment, the nuclear localization signal is located at both the N-terminus and C-terminus of the PCV HUH endonuclease.
[0186] In one embodiment, the ribonucleoprotein comprises at least one nuclear localization signal, hi another embodiment, the ribonucleoprotein comprises at least two nuclear localization signals.
[0187] In one aspect, the nuclear localization signal provided herein is encoded by SEQ ID NO: 8. In one aspect, the nuclear localization signal provided herein is encoded by SEQ ID NO: 9. In another aspect, the nuclear localization signal is selected from the group consisting of SEQ ID NOs: 8 and 9.
[0188] Various species exhibit specific biases toward specific codons for specific amino acids. Codon bias (differences in codon usage among organisms) is often correlated with the efficiency of messenger RNA (mRNA) translation and is thought to depend, inter alia, on the characteristics of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell generally reflects the codons most frequently used in peptide synthesis. Therefore, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the "Codon Usage Database," available at www.kazusa.or.jp / codon, and these tables can be adapted in many ways. See Nakamura et al., 2000, Nucl. Acids Res. 28:292. Computer algorithms for optimizing codons for a particular sequence for expression in a particular host cell are also available, for example, Gene Forge (Aptagen; Jacobus, PA).
[0189] As used herein, "codon optimization" refers to the process of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon of the sequence (e.g., at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) with a codon that is more frequently or most frequently used in the genes of the host cell, while maintaining the original amino acid sequence (e.g., introducing silent mutations).
[0190] In one aspect, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons, or all codons) in a sequence encoding a Cas12a nuclease, a CasX nuclease, or a HUH endonuclease correspond to the most frequently used codon for a particular amino acid. Regarding codon usage in plants, including algae, see Campbell and Gowri, 1990, Plant Physiol., 92:1-11; and Murray et al., 1989. Nucleic Acids Res. 17:477-98, each of which is incorporated by reference in its entirety.
[0191] In one embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for prokaryotic cells. In one embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for E. coli cells. In one embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for eukaryotic cells. In one embodiment, the nucleic acid molecule provided herein is codon-optimized for animal cells. In one embodiment, the nucleic acid molecule provided herein is codon-optimized for fungal cells. In one embodiment, the nucleic acid molecule provided herein is codon-optimized for yeast cells. In another embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for plant cells. In another embodiment, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for monocotyledonous plant species. In another embodiment, the protein-encoding nucleic acid molecule has been codon-optimized for dicotyledonous plant species. In further aspects, nucleic acid molecules provided herein encode a Cas12a nuclease that has been codon-optimized for gymnosperm species. In further aspects, nucleic acid molecules provided herein encode a Cas12a nuclease that has been codon-optimized for angiosperm species. In further aspects, nucleic acid molecules provided herein encode a Cas12a nuclease that has been codon-optimized for corn cells. In further aspects, nucleic acid molecules provided herein encode a Cas12a nuclease that has been codon-optimized for soybean cells. In further aspects, nucleic acid molecules provided herein encode a Cas12a nuclease that has been codon-optimized for rice cells. In further aspects, nucleic acid molecules provided herein encode a Cas12a nuclease that has been codon-optimized for wheat cells. In further aspects, nucleic acid molecules provided herein encode a Cas12a nuclease that has been codon-optimized for cotton cells. In further aspects, nucleic acid molecules provided herein encode a Cas12a nuclease that has been codon-optimized for sorghum cells. In a further aspect, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for alfalfa cells. In a further aspect, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for sugarcane cells. In a further aspect, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for Arabidopsis cells. In a further aspect, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for tomato cells. In a further aspect, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for cucumber cells. In a further aspect, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for potato cells. In a further aspect, the nucleic acid molecule provided herein encodes a Cas12a nuclease that has been codon-optimized for algal cells.
[0192] In one embodiment, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for a prokaryotic cell. In one embodiment, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for an E. coli cell. In one embodiment, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for a eukaryotic cell. In one embodiment, a nucleic acid molecule provided herein is codon-optimized for an animal cell. In one embodiment, a nucleic acid molecule provided herein is codon-optimized for a fungal cell. In one embodiment, a nucleic acid molecule provided herein is codon-optimized for a yeast cell. In another embodiment, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for a plant cell. In another embodiment, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for a monocotyledonous plant species. In another embodiment, a nucleic acid molecule encoding a protein has been codon-optimized for a dicotyledonous plant species. In a further aspect, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for a gymnosperm species. In a further aspect, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for an angiosperm species. In a further aspect, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for a corn cell. In a further aspect, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for a soybean cell. In a further aspect, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for a rice cell. In a further aspect, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for a wheat cell. In a further aspect, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for a cotton cell. In a further aspect, a nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for a sorghum cell. In a further aspect, the nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for alfalfa cells. In a further aspect, the nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for sugarcane cells. In a further aspect, the nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for Arabidopsis cells. In a further aspect, the nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for tomato cells. In a further aspect, the nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for cucumber cells. In a further aspect, the nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for potato cells. In a further aspect, the nucleic acid molecule provided herein encodes a CasX nuclease that has been codon-optimized for algal cells.
[0193] In one aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for prokaryotic cells. In one aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for E. coli cells. In one aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for eukaryotic cells. In one aspect, the nucleic acid molecule provided herein is codon-optimized for animal cells. In one aspect, the nucleic acid molecule provided herein is codon-optimized for fungal cells. In one aspect, the nucleic acid molecule provided herein is codon-optimized for yeast cells. In another aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for plant cells. In another aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for monocotyledonous plant species. In another aspect, the nucleic acid molecule encoding the protein has been codon-optimized for dicotyledonous plant species. In a further aspect, the nucleic acid molecules provided herein encode a HUH endonuclease that has been codon-optimized for gymnosperm species. In a further aspect, the nucleic acid molecules provided herein encode a HUH endonuclease that has been codon-optimized for angiosperm species. In a further aspect, the nucleic acid molecules provided herein encode a HUH endonuclease that has been codon-optimized for corn cells. In a further aspect, the nucleic acid molecules provided herein encode a HUH endonuclease that has been codon-optimized for soybean cells. In a further aspect, the nucleic acid molecules provided herein encode a HUH endonuclease that has been codon-optimized for rice cells. In a further aspect, the nucleic acid molecules provided herein encode a HUH endonuclease that has been codon-optimized for wheat cells. In a further aspect, the nucleic acid molecules provided herein encode a HUH endonuclease that has been codon-optimized for cotton cells.In a further aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for sorghum cells. In a further aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for alfalfa cells. In a further aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for sugarcane cells. In a further aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for Arabidopsis cells. In a further aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for tomato cells. In a further aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for cucumber cells. In a further aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for potato cells. In a further aspect, the nucleic acid molecule provided herein encodes a HUH endonuclease that has been codon-optimized for algal cells.
[0194] In one aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for prokaryotic cells. In one aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for E. coli cells. In one aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for eukaryotic cells. In one aspect, a nucleic acid molecule provided herein is codon-optimized for animal cells. In one aspect, a nucleic acid molecule provided herein is codon-optimized for fungal cells. In one aspect, a nucleic acid molecule provided herein is codon-optimized for yeast cells. In another aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for plant cells. In another aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for monocotyledonous plant species. In another aspect, a nucleic acid molecule encoding a protein has been codon-optimized for dicotyledonous plant species. In a further aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for gymnosperm species. In a further aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for angiosperm species. In a further aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for corn cells. In a further aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for soybean cells. In a further aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for rice cells. In a further aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for wheat cells. In a further aspect, a nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for cotton cells. In a further aspect, the nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for sorghum cells. In a further aspect, the nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for alfalfa cells. In a further aspect, the nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for sugarcane cells. In a further aspect, the nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for Arabidopsis cells. In a further aspect, the nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for tomato cells. In a further aspect, the nucleic acid molecule provided herein encodes an FBNYV HUH endonuclease that has been codon-optimized for cucumber cells. In a further aspect, the nucleic acid molecules provided herein encode an FBNYV HUH endonuclease that has been codon-optimized for potato cells. In a further aspect, the nucleic acid molecules provided herein encode an FBNYV HUH endonuclease that has been codon-optimized for algal cells.
[0195] In one embodiment, a nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for prokaryotic cells. In one embodiment, a nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for E. coli cells. In one embodiment, a nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for eukaryotic cells. In one embodiment, a nucleic acid molecule provided herein is codon-optimized for animal cells. In one embodiment, a nucleic acid molecule provided herein is codon-optimized for fungal cells. In one embodiment, a nucleic acid molecule provided herein is codon-optimized for yeast cells. In another embodiment, a nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for plant cells. In another embodiment, a nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for monocotyledonous plant species. In another embodiment, a nucleic acid molecule encoding a protein has been codon-optimized for dicotyledonous plant species. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for gymnosperm species. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for angiosperm species. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for corn cells. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for soybean cells. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for rice cells. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for wheat cells. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for cotton cells. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for sorghum cells. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for alfalfa cells. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for sugarcane cells. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for Arabidopsis cells. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for tomato cells. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for cucumber cells. In a further aspect, the nucleic acid molecule provided herein encodes a PCV HUH endonuclease that has been codon-optimized for potato cells. In a further aspect, the nucleic acid molecules provided herein encode a PCV HUH endonuclease that is codon-optimized for an algal cell.
[0196] In one aspect, a nucleic acid molecule provided herein encodes a linker that is codon-optimized for a prokaryotic cell. In one aspect, a nucleic acid molecule provided herein encodes a linker that is codon-optimized for an E. coli cell. In one aspect, a nucleic acid molecule provided herein encodes a linker that is codon-optimized for a eukaryotic cell. In one aspect, a nucleic acid molecule provided herein is codon-optimized for an animal cell. In one aspect, a nucleic acid molecule provided herein is codon-optimized for a fungal cell. In one aspect, a nucleic acid molecule provided herein is codon-optimized for a yeast cell. In another aspect, a nucleic acid molecule provided herein encodes a linker that is codon-optimized for a plant cell. In another aspect, a nucleic acid molecule provided herein encodes a linker that is codon-optimized for a monocotyledonous plant species. In another aspect, a nucleic acid molecule encoding a protein is codon-optimized for a dicotyledonous plant species. In a further aspect, a nucleic acid molecule provided herein encodes a linker that is codon-optimized for a gymnosperm species. In a further aspect, a nucleic acid molecule provided herein encodes a linker that is codon-optimized for an angiosperm species. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for corn cells. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for soybean cells. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for rice cells. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for wheat cells. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for cotton cells. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for sorghum cells. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for alfalfa cells.In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for sugarcane cells. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for Arabidopsis cells. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for tomato cells. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for cucumber cells. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for potato cells. In a further aspect, the nucleic acid molecules provided herein encode a linker that is codon-optimized for algal cells. [Example]
[0197] Example 1 Expression vector design The Cys-free LbCas12a protein (SEQ ID NO: 1) was fused to the HUH endonuclease (SEQ ID NO: 2) from Broad Bean Necrotic Yellows Virus (FBNYV) via a novel 16-amino acid flexible linker (SEQ ID NO: 3). One of the linker's optimal features is the presence of numerous glycine, serine, and threonine residues. Without being limited to any scientific theory, these residues are not hydrophobic and can easily associate with the solvent between the two protein domains. Without being limited to any scientific theory, these residues are also among the most flexible amino acids, thereby giving the linker the ability to adopt many different conformations and thus allowing the two protein domains freedom to move relative to each other. Also, without being limited to any scientific theory, aspartic acid and glutamic acid residues are highly charged, a feature that promotes interactions with the solvent and reduces the likelihood of protein aggregation.
[0198] The sequences encoding the three components of the fusion protein (cys-free LbCas12a, linker, and HUH) were codon-optimized for optimal expression in E. coli. The nucleotide sequence encoding cys-free LbCas12a is shown as SEQ ID NO:4. The nucleotide sequence encoding the HUH endonuclease is shown as SEQ ID NO:5, and the nucleotide sequence encoding the linker is shown as SEQ ID NO:6. Both N- and C-terminal HUH fusion proteins were designed. The nucleotide sequence encoding cys-free LbCas12a:linker:HUH is shown as SEQ ID NO:7, and the nucleotide sequence encoding HUH:linker:cys-free LbCas12a is shown as SEQ ID NO:17. Nuclear localization signals (NLS) (SEQ ID NOs:8 and 9) were introduced at the 5' and 3' ends of the LbCas12a-linker:HUH open reading frame, respectively. Additionally, a nucleotide sequence encoding a histidine (HIS) tag (SEQ ID NO:10) was introduced 5' to the 5' NLS sequence. A nucleotide sequence encoding the HIS-tag:NLS1:cys-free LbCas12a:linker:HUH:NLS2 (hereinafter referred to as "cys-free LbCas12a:HUH") sequence was placed under the control of the P-Ec.tac promoter (SEQ ID NO: 11) and inserted into a bacterial expression vector. A schematic representation of the N- and C-terminal HUH fusion protein expression cassettes is shown in Figure 1B. A control expression vector containing the HIS-tag:NLS1:cys-free LbCas12a:NLS2 (hereinafter referred to as "cys-free LbCas12a control") sequence under the control of the P-Ec.tac promoter was also designed. See Figure 1B. Cys-free LbCas12a:HUH and cys-free LbCas12a control proteins were expressed and purified from E. coli cells transformed with the above expression vectors.
[0199] Example 2 Fusion protein testing To test whether cys-free LbCas12a:HUH can recognize and cleave chromosomal DNA in the presence of its cognate guide RNA, three unique target sites were selected within the soybean (Glycine max) genome. crRNAs were designed to guide cys-free LbCas12a:HUH and cys-free LbCas12a control proteins to each target site. Ribonucleoprotein (RNP) complexes containing purified cys-free LbCas12a:HUH fusion protein or cys-free LbCas12a and the cognate crRNA were assembled.
[0200] A 70-nucleotide long ssDNA template was designed. The template contains a 10-nucleotide signature motif flanked by homology arms containing homology to sequences adjacent to the GmTS1 site. The ssDNA template was fused to the 15-nucleotide origin (ori) HUH recognition sequence of PCV (SEQ ID NO: 12). As shown in Example 6, the ori sequence from PCV is compatible with the FBNYV HUH endonuclease.
[0201] This single-stranded DNA template (ssDNA template) was added to the RNP complex. This tripartite RNP complex (see Figure 1A) was then tested in vitro for the predicted functionality of the two fusion domains: (a) targeted DNA cleavage and (b) covalent tethering of the ssDNA template. To examine the covalent tethering of the ssDNA template to cys-free LbCas12a:HUH, a gel shift assay was performed (see Table 1). The gel shift assay confirmed that in the presence of cys-free LbCas12a:HUH, the ssDNA template migrated significantly slower than in the presence of either the cys-free LbCas12a control or its unbound ssDNA template (see Table 1). The slower migration indicates the formation of a complex tethered with cys-free LbCas12a:HUH::ssDNA.
[0202] [Table 1]
[0203] The ability of cys-free LbCas12a:HUH to cleave target DNA was assessed by in vitro assay. A 1020-nucleotide PCR amplicon spanning two target sites (GmTS1 and GmTS2) was generated from wild-type soybean germplasm A3555. The ability of cys-free LbCas12a:HUH to cleave the PCR amplicon at the expected site in the presence of the cognate crRNA was tested. Cys-free LbCas12a was used as a positive control, and an assay lacking nuclease served as a negative control.
[0204] Targeted cleavage of the PCR amplicon by cys-free LbCas12a in the presence of crRNA-TS1 is predicted to result in 892-nucleotide and 128-nucleotide digestion products. Targeted cleavage of the PCR amplicon by cys-free LbCas12a in the presence of crRNA-TS2 is predicted to result in 734-nucleotide and 286-nucleotide digestion products. As shown in Table 2, the PCR amplicon was digested in the expected pattern by both the cys-free LbCas12a:HUH complex and the cys-free LbCas12a control complex.
[0205] [Table 2]
[0206] Example 3 Testing of fusion proteins in protoplasts Trimeric cys-free LbCas12a:HUH / crRNA / ssDNA template RNP complexes were tested for chromosome breakage in protoplast cells. Protoplasts were prepared from soybean embryos and then transformed with RNP complexes containing the cys-free LbCas12a:HUH / crRNA / ssDNA template and their positive control counterparts carrying cys-free LbCas12a by standard methods known in the art. Negative controls included matched crRNA and ssDNA templates but no fusion protein (see Table 3).
[0207] A qualitative assay (Table 3) was developed to examine the activity of cys-free LbCas12a:HUH. After 2 days of incubation at room temperature, total genomic DNA was isolated from all samples, and 1020-nucleotide PCR amplicons spanning the target sites GmTS1 and GmTS2 were generated. The amplicons contained a heterogeneous population of DNA sequences encompassing wild-type and edited target sites, the latter of which were expected to contain indels. The PCR-amplified amplicons were then re-digested in vitro with cys-free LbCas12a. Undigested amplicons, indicative of edited target sites, were observed in both the test and positive controls. Undigested amplicons were not detectable in the negative control. Sequencing of the undigested fractions confirmed the presence of targeted indels in both the test and positive controls.
[0208] [Table 3]
[0209] For quantitative comparisons between treatments, amplicons spanning the target sites were generated and sequenced by next-generation sequencing (NGS) using standard methods known in the art. Sequencing reads were considered mutant if they contained an indel within the 7-nucleotide LbCas12a cleavage site, located 18–24 nucleotides downstream of the LbCas12a PAM site. In all three target sites tested, indel rates of 7–19% were detected in both the test and positive controls. Negative controls had indel rates of 0.3% or less in the same regions, which was significantly lower than the test / positive control cleavage rates for all sites (Figure 2).
[0210] Example 4 In vivo testing of fusion proteins The cys-free LbCas12a:HUH / crRNA / template complex was tested for chromosome breakage and targeted integration in plants. Test and control samples contained both the cys-free LbCas12a:HUH fusion protein and the crRNA for the GmTS1 target site, as described in Example 1. The template for the test assay contained an ssDNA template containing the 15-nucleotide ori HUH recognition sequence from PCV, as described in Example 2. The control assay contained an ssDNA template lacking the ori sequence. See Figure 3A.
[0211] The RNP complex was assembled and used for plant transformation. Two independent transformations were performed using particle bombardment. The RNP (cys-free LbCas12a:HUH / crRNA / template) complex, along with a linear dsDNA fragment encoding the spectinomycin resistance gene aadA, was coated onto 0.6 μm gold particles using a mixture of CaCl2 and spermidine as a coating agent. The coated particles were delivered into desiccated excised soybean embryos using the Biolistic PDS-1000 / HE particle delivery system (Biorad).
[0212] Transformed seedlings were grown in tissue culture using spectinomycin for selection. Given near-stoichiometric delivery of the RNP and spectinomycin selectable marker genes, surviving plants were expected to be enriched for plants carrying the RNP complex. Seedlings with at least one trifoliate leaf were sampled for DNA analysis.
[0213] Amplicons spanning the GmTS1 target site were generated and subjected to next-generation sequencing (NGS) using standard methods known in the art. The chromosome cleavage activity of cys-free LbCas12a:HUH was determined by quantifying the presence of targeted indels within the 7-nucleotide GmTS1 cleavage site. Plants that generated at least 20% targeted indel reads were designated mutants (see Table 4 and Figure 3B). The mutation rate exceeded 30% in all treatments, suggesting robust cleavage activity by the cys-free LbCas12a:HUH fusion protein (see Table 4 and Figure 3B).
[0214] [Table 4]
[0215] Targeted integration was defined by the presence of a unique signature motif of the template in the NGS amplicon spanning the GmTS1 target site (see Figure 4). Integration can occur via two fundamental DNA repair processes: non-homologous end joining (NHEJ) and homologous recombination (HR), or a combination of NHEJ and HR. Different template-chromosome junctions created by NHEJ and HR can be used to determine the mechanism of integration (see Figure 4).
[0216] Overall, HUH-mediated tethering of exogenous ssDNA to cys-free LbCas12a promoted targeted integration (see Table 5; Table 6; and Figure 5). For example, integration of single-copy templates by NHEJ was significantly higher in the test compared to the control in all batches. Homologous recombination detected up to 3.8% integration in the test, compared to 1.3% or less in the control. Covalent tethering of ssDNA templates to LbCas12a:HUH enhanced integration of targeted templates in plants.
[0217] [Table 5]
[0218] [Table 6]
[0219] Example 5 Testing of N- and C-terminal HUH fusion proteins in protoplasts The N- and C-terminal configurations of LbCas12a using the FBNYV HUH fusion protein described in Example 1 and Figure 1B were tested in vitro for targeted chromosome breakage and subsequent DNA repair at the GmTS1 target site in soybean protoplasts (see Figure 6). Protoplasts were transformed using standard methods known in the art with various combinations of reagents listed in Table 7. Reagents included "N-terminal" (HUH:cys-free LbCas12a) or "C-terminal" (cys-free LbCas12a:HUH) protein, a crRNA for the GmTS1 target site, and a 70-nt ssDNA template (described in Example 4) with or without a 5' PCVori extension to facilitate ligation to the HUH domain and a 90-bp dsDNA oligonucleotide (dsDNA oligo).
[0220] [Table 7]
[0221] After incubation at room temperature for 2 days, total genomic DNA was isolated. PCR was performed using primers flanking the GmTS1 target site. Amplicons were sequenced using MiSeq technology (www.illmina.com). Sequences were analyzed for (1) targeted insertions and deletions (indels) and (2) targeted integration of ssDNA templates by either HR or alternative mechanisms, collectively labeled as NHEJ-mediated integration. In the latter case, single- and multiple-copy integrations were also distinguished. HR, by definition, can integrate only one copy of the template. (3) Multiple- and single-copy integration of dsDNA oligonucleotides by NHEJ was also quantified.
[0222] Chromosome cleavage was confirmed for both N- and C-terminal enzyme configurations (HUH:cys-free LbCas12a and cys-free LbCas12a:HUH), as evidenced by the presence of indels at the target site in the presence of either the fusion protein or the cognate gRNA. See, for example, treatments 1-8, 11, and 12 compared to control treatments 13-16 in Figure 7. Cys-free LbCas12a:HUH ("C-terminal" fusion) exhibited a higher indel rate and therefore better cleavage rate than HUH:cys-free LbCas12a ("N-terminal fusion"). Tethering the ssDNA template to the HUH domain negatively affected chromosome cleavage rates. Without being bound to any particular theory, spatial interference between the DNA template and the Cas12a-chromosome interface may interfere with nuclease activity. However, NHEJ-mediated template integration was significantly higher when the template was tethered to the HUH domain for both protein configurations (see Figure 8). Similarly, HR-mediated integration of ssDNA templates was significantly higher when the template was tethered to the HUH domain for both protein configurations (see Figure 8 and Table 8). The integration rate of dsDNA oligos was higher than that of ssDNA templates, excluding the effect of HUH tethering (see Figure 9). Consistent with the indel rate, C-terminal fusion proteins showed improved activity for incorporating ds-oligos compared to N-terminal fusion proteins, even though the difference was not statistically significant. Also consistent with the indel rate findings, tethering ssDNA templates to fusion proteins adversely affected targeted integration of ds-oligos. In general, NHEJ-mediated integration of either ssDNA or dsDNA oligonucleotides showed a close correlation with chromosome breakage rates.
[0223] [Table 8]
[0224] Example 6 Compatibility of viral replication origins from various species with FBNYV-HUH endonuclease fused to Cys-free LbCas12a Both Broad Bean Necrotic Yellows Virus (FBNYV) and Porcine Circovirus 2 (PCV) belong to the Arfiviricetes class and possess HUH proteins with conserved structures. Similarly, their origins of replication (ori) share the same minimal nonamer core sequence (agtattacc) required for recognition and cleavage. See Vega-Rocha et al., 2007, Biochemistry 46:6201 and Timchenko et al., 1999, J. of Virol. 73:10173. Recognition of the FBNYV ori (SEQ ID NO:27) and PCV ori (SEQ ID NO:12) sequences by Cas12a fusion proteins carrying the FBNYV HUH endonuclease was examined. The Tral ori sequence (SEQ ID NO: 28) from the unrelated HUH relaxase Tral derived from the bacterial conjugation F plasmid (see Dostal et al., 2011, Nucleic Acids Res 39: 2658) and a "mock" ori sequence containing a 15-bp soybean fragment similar in size to the PCV2 ori were incorporated as negative controls. The three ori sequences and the mock ori sequence were fused to a 70-nucleotide-long ssDNA template, and gel shift assays similar to those described in Example 2 were performed. Gel shift assays (see Table 6) confirmed that in the presence of cys-free LbCas12a:HUH, ssDNA templates containing the FBNYV ori and PCV ori migrated significantly slower than those containing the Tral ori or mock ori. The slower migration indicates the formation of a complex tethered to the cys-free LbCas12a:HUH::ssDNA. Gel shifts of tethered oligos show that the ori sequences from PCV and FBNYV match the HUH protein of FBNYV, while the ori from the distant homolog Tral and a mimic ori of similar size do not.
[0225] [Table 9]
Claims
1. (a)(i) an amino acid sequence of a Cas12a nuclease comprising an amino acid sequence having at least 90% sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26; (ii) an amino acid sequence of a linker comprising an amino acid sequence that is at least 93% identical to the amino acid sequence of SEQ ID NO: 3; and (iii) a HUH nuclease selected from the group consisting of Broad bean necrotic yellows virus (FBNYV) HUH endonuclease, Porcine circovirus 2 (PCV) HUH endonuclease, Duck circovirus (DCV) HUH endonuclease, Streptococcus agalactiae replication protein RepB (RepB), Fructobacillus tropaeoli RepB (RepBm), Escherichia coli conjugation protein TraI (TraI), Escherichia coli mobilization protein A (mMobA), and Staphylococcus aureus nicking enzyme (NES), or an amino acid sequence of a HUH nuclease comprising an amino acid sequence having at least 90% sequence identity to the amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. a recombinant polypeptide comprising: (b) at least one guide nucleic acid Ribonucleoproteins containing
2. (a) a first nucleic acid sequence encoding a Cas12a nuclease comprising an amino acid sequence having at least 90% sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26; (b) a second nucleic acid sequence encoding a linker comprising an amino acid sequence that is at least 93% identical to the amino acid sequence of SEQ ID NO: 3; and (c) a HUH nuclease selected from the group consisting of Broad Bean Necrotic Yellows Virus (FBNYV) HUH endonuclease, Porcine Circovirus 2 (PCV) HUH endonuclease, Duck Circovirus (DCV) HUH endonuclease, Streptococcus agalactiae replication protein RepB (RepB), Fructobacillus tropaeoli RepB (RepBm), Escherichia coli conjugation protein TraI (TraI), E. coli mobilization protein A (mMobA), and Staphylococcus aureus nicking enzyme (NES), or a third nucleic acid sequence encoding a HUH nuclease comprising an amino acid sequence having at least 90% sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. A recombinant nucleic acid comprising:
3. The ribonucleoprotein of claim 1, further comprising at least one single-stranded DNA template nucleic acid molecule.
4. 3. The recombinant nucleic acid of claim 2, further comprising: (d) a fourth nucleic acid sequence encoding at least one guide nucleic acid, or a fourth nucleic acid sequence encoding at least one template nucleic acid molecule, or both.
5. 5. The ribonucleoprotein of claim 3, or the recombinant nucleic acid of claim 4, wherein the template nucleic acid molecule comprises a nucleic acid sequence of an origin of replication (ori), the ori being capable of being nicked by the HUH nuclease followed by the formation of a covalent phosphotyrosine intermediate, whereby the 5' end of a DNA strand is linked to a tyrosine in the HUH protein, and the origin of replication is located at the 5' or 3' end of the template nucleic acid molecule.
6. 3. The recombinant nucleic acid of claim 2, wherein the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, or any combination thereof is codon-optimized for a prokaryotic cell or for a eukaryotic cell selected from the group consisting of a plant cell, a mammalian cell, a fungal cell, an insect cell, an arachnid cell, an avian cell, a fish cell, a reptile cell, and an amphibian cell.
7. 10. A host cell comprising the ribonucleoprotein of claim 1 or the recombinant nucleic acid of claim 2, wherein the host cell is selected from the group consisting of a plant cell, a bacterial cell, a mammalian cell, a fungal cell, an insect cell, an arachnid cell, an avian cell, a fish cell, a reptilian cell, and an amphibian cell.
8. A plasmid comprising the recombinant nucleic acid of claim 2 or 4.
9. 1. A method for generating an edit in a target DNA molecule, comprising contacting the target DNA molecule with a ribonucleoprotein, wherein the ribonucleoprotein (a)(i) an amino acid sequence of a Cas12a nuclease comprising an amino acid sequence having at least 90% sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26; (ii) an amino acid sequence of a linker comprising an amino acid sequence that is at least 93% identical to the amino acid sequence of SEQ ID NO: 3; and (iii) a HUH nuclease selected from the group consisting of Broad bean necrotic yellows virus (FBNYV) HUH endonuclease, Porcine circovirus 2 (PCV) HUH endonuclease, Duck circovirus (DCV) HUH endonuclease, Streptococcus agalactiae replication protein RepB (RepB), Fructobacillus tropaeoli RepB (RepBm), Escherichia coli conjugation protein TraI (TraI), Escherichia coli mobilization protein A (mMobA), and Staphylococcus aureus nicking enzyme (NES), or an amino acid sequence of a HUH nuclease comprising an amino acid sequence having at least 90% sequence identity to the amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. a recombinant polypeptide comprising: (b) at least one guide nucleic acid; and (c) at least one template nucleic acid molecule Including, wherein the ribonucleoprotein generates at least one edit in the target DNA molecule.
10. 1. A method for generating an edit in a target DNA molecule, comprising administering to a cell: (a)(i) an amino acid sequence of a Cas12a nuclease comprising an amino acid sequence having at least 90% sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1, 15, and 26; (ii) an amino acid sequence of a linker comprising an amino acid sequence that is at least 93% identical to the amino acid sequence of SEQ ID NO: 3; and (iii) a HUH nuclease selected from the group consisting of Broad bean necrotic yellows virus (FBNYV) HUH endonuclease, Porcine circovirus 2 (PCV) HUH endonuclease, Duck circovirus (DCV) HUH endonuclease, Streptococcus agalactiae replication protein RepB (RepB), Fructobacillus tropaeoli RepB (RepBm), Escherichia coli conjugation protein TraI (TraI), Escherichia coli mobilization protein A (mMobA), and Staphylococcus aureus nicking enzyme (NES), or an amino acid sequence of a HUH nuclease comprising an amino acid sequence having at least 90% sequence identity to the amino acid sequence selected from the group consisting of SEQ ID NOs: 2 and 14. a recombinant polypeptide comprising or one or more nucleic acid molecules encoding said recombinant polypeptide; (b) at least one guide nucleic acid, or at least one nucleic acid molecule encoding said at least one guide nucleic acid; and (c) at least one template nucleic acid molecule, or at least one nucleic acid molecule encoding said at least one template nucleic acid molecule; providing wherein the recombinant polypeptide, at least one guide nucleic acid, and at least one template nucleic acid molecule form a ribonucleoprotein, and the ribonucleoprotein generates at least one edit in the target DNA molecule in the cell.
11. 10. The method of claim 9, wherein the ribonucleoprotein is formed intracellularly or extracellularly.
12. 11. The method of Claim 9 or 10, wherein the at least one edit comprises a deletion, substitution, or inversion, or the incorporation of at least 5 consecutive nucleotides from the at least one template nucleic acid molecule into the target DNA molecule.
13. 1. An isolated polypeptide comprising the amino acid sequence of SEQ ID NO:3, wherein the amino acid sequence is disposed between: a) the amino acid sequence of a first nuclease selected from the group consisting of Cas12a nuclease, CasX nuclease, Cas9 nuclease, meganuclease, zinc finger nuclease, transcription activator-like nuclease, and HUH endonuclease; and b) the amino acid sequence of a second nuclease selected from the group consisting of Cas12a nuclease, CasX nuclease, Cas9 nuclease, meganuclease, zinc finger nuclease, transcription activator-like nuclease, and HUH endonuclease.
14. A polypeptide comprising an amino acid sequence encoded by a nucleic acid sequence selected from SEQ ID NOs: 7, 17-19, and 24, or an amino acid sequence selected from SEQ ID NOs: 20-23 and 25.
15. An expression cassette encoding an amino acid sequence encoded by a nucleic acid sequence selected from SEQ ID NOs: 7, 17-19, and 24, or an amino acid sequence selected from SEQ ID NOs: 20-23 and 25.
16. 16. A cell comprising a polypeptide according to claim 14 and / or an expression cassette according to claim 15.
17. 8. The host cell of claim 7, wherein the host cell is a plant cell selected from the group consisting of a corn cell, a soybean cell, a cotton cell, a canola cell, a rice cell, a wheat cell, a sorghum cell, an alfalfa cell, a sugarcane cell, a millet cell, a tomato cell, a potato cell, an oilseed rape cell, and an algae cell, or the bacterial cell is an Escherichia coli cell.
Citation Information
Patent Citations
Engineering and optimization of improved systems, methods and enzyme compositions for sequence engineering
JP2016501532A
CRISPR-CPF1-Related Methods, Compositions, and Components for Cancer Immunotherapy
JP2019507599A
US1990874144-4148
Maize chloroplast aldolase promoter compositions and methods for use thereof
US20040216189A1
Novel crispr enzymes and systems
US20190233814A1