Compositions, constructs, plant viral vectors, and methods for plant gene editing
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- RGT UNIV OF CALIFORNIA
- Filing Date
- 2024-06-28
- Publication Date
- 2026-05-06
AI Technical Summary
Current methods for delivering gene editing reagents into plants are inefficient, requiring time-consuming and resource-intensive tissue culture techniques, and are limited by the need for pre-established transgenic lines and species-specific applications, with plant viruses like TRV facing cargo capacity constraints that prevent the delivery of larger CRISPR-Cas enzymes.
Employing small TnpB-type bacterial enzymes, less than 500 amino acids in length, which can be cloned into plant viral vectors like TRV, enabling efficient gene editing without tissue culture and allowing for the delivery of both the enzyme and guide RNA, overcoming cargo size limitations.
Facilitates efficient gene editing in plants by bypassing traditional tissue culture methods and accommodating larger gene editing components, enhancing editing efficiency and broadening applicability across different plant species.
Smart Images

Figure US2024036209_02012025_PF_FP_ABST
Abstract
Description
COMPOSITIONS, CONSTRUCTS, PLANT VIRAL VECTORS, AND METHODS FOR PLANT GENE EDITINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional App. No. 63 / 511,070, filed on June 29, 2023. U.S. Provisional App. No. 63 / 520.258, filed on August 17, 2023, and U.S. Provisional App. No. 63 / 566,666, filed on March 18, 2024. The entire contents of each of the aforementioned applications are incorporated by reference herein.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with government support under award number 2334027 awarded by the National Science Foundation. The government has certain rights in the invention.REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0003] The contents of the electronic sequence listing (790482.00489.xml; Size: 785,615 bytes; and Date of Creation: June 26, 2024) is herein incorporated by reference in its entirety'.BACKGROUND
[0004] Given current crop productivity projections, agricultural practices will be inadequate to meet the global food demands by the year 2050 (Ray et al., 2013). Our ability to address this problem will largely depend on the efficiency that novel genetic diversity can be created and introduced into current plant breeding pipelines. Historically, novel phenotypes were generated through random mutation or transgenic breeding, followed by successive backcrossing to an elite crop variety, taking upwards of 10-12 years to complete (Chen et al., 2019). With the advent of genome editing, we now have the ability to more efficiently genetically modify' crop genomes, resulting in beneficial phenotypes(Haun et al., 2014; Lemmon et al., 2018; Liu et al., 2021; Rodriguez-Leal et al., 2017; Zsogon et al., 2018). Despite this advance, a primary bottleneck remains: fast and efficient delivery of the gene editing reagents into crop plants (Altpeter et al.. 2016).SUMMARY
[0005] In one aspect, a composition is provided. The composition can include a TnpB bacterial enzy me. The composition can also include a polynucleotide including a first portion adapted to bind to at least a portion of the TnpB bacterial enzyme and a second portion for binding to a plant genomic target.
[0006] In another aspect, a construct is provided. The construct can include a plant promotor operably connected to a polynucleotide encoding for a TnpB bacterial enzyme.
[0007] In yet another aspect, a plant viral vector is provided. The plant viral vector can include a construct that can include a plant promotor operably connected to a polynucleotide encoding for a TnpB bacterial enzyme.
[0008] In another aspect, a plant cell is provided. The plant cell can include a construct that can include a plant promotor operably connected to a polynucleotide encoding for a TnpB bacterial enzyme, or a plant viral vector that includes the construct.
[0009] In another aspect, a method for gene editing a plant is provided. The method can include introducing a construct or a plant viral vector into one or more plant cells. The construct can include a plant promotor operably connected to a polynucleotide encoding for a TnpB bacterial enzyme. The plan viral vector can include the construct.BRIEF DESCRIPTION OF THE FIGURES
[0010] FIGS. 1 A and IB. (FIG. lA)The '‘DOTS’’ pipeline for detecting bona-fide RNA- guided TnpB systems in metagenomes. (FIG. IB) A phylogenetic tree of TnpBs from metagenomes (red) as well as a public IS database, ISfinder (black), with the corresponding amino acid length shown in the bar graph below. Distinct TAM sequences and small-sized TnpB enzymes from metagenomes are also highlighted in red circles and black arrows, respectively.
[0011] FIGS. 2A-2C. (FIG. 2A) Representative loci architectures for TnpB-associated systems characterized in this study. (FIG. 2B) The Web logo of the TAMs for ISBagOl and ISXfaOl TnpB systems. (FIG. 2C) Comparable E. coli interference activities between Casl2a, ISBagOl and ISXfaOl, when the target site is flanked by a PAM / TAM at 26°C.
[0012] FIGS. 3A-3C. (FIG. 3A) Schematics of the constructs used to test gene editing activity by ISBagOl. Top panel, ISBagOl coRNA with a 20bp spacer targeting the AtPDS3 gene were driven by the AtU6-26 promoter and followed by the HDV ribozyme. 2X indicates that the construct also contains a second similar AtU6-26 driven coRNA cassette with a second gRNA sequence that is not shown. Zea mays codon optimized ISBagOl coding sequence was expressed in a separate transcription cassette driven by the UBQ10 gene promoter and an rbcS- E9 terminator. Bottom panel, native ISBagOl sequence encoding the ISBagOl protein and the overlapping coRNA were driven by the UBQ10 gene promoter, followed by a 20bp spacer targeting the AtPDS3 gene the HDV ribozyme, and the rbcS-E9 terminator. NLS, nuclear localization signal. (FIG. 3B) Left panel, editing efficiency of ISBagOl at the AtPDS3 gene at regions targeted by four different guide RNAs in Arabidopsis protoplasts, with ISBagOl protein coding sequence and coRNA coding sequence in split cassettes. Right panel, editing efficiency by ISBagOl and .4 / / 7A3 gRNAl, with native ISBagOl coding sequence and coRNA sequencein the same cassette. (FIG. 3C) Examples of detailed editing profiles by ISBagOl in Arabidopsis protoplasts detected by amplicon sequencing. Left panel, AtPDSS gRNAl. Right panel, AtPDS3 gRNAl 8. Editing events are shown on the left as: (position where the editing starts): (number of nucleotides of)D(deletion). position 0 is between the 18th and 19th nucleotides of the guide RNA sequence, such that the 18th nucleotide is position -1, the 19th nucleotide is position +1. Read number, number of reads showing a particular edit. AtPDS3 gRNAl reference sequence is SEQ ID NO: 610, amplicons in descending order are SEQ ID NOs: 61 1 -620. AtPDS3 gRNA18 reference sequence is SEQ ID NO: 621 , amplicons in descending order are SEQ ID NOs: 622-625.
[0013] FIG. 4A and 4B. Schematics of a bacteria selection assay. (FIG. 4A) The E.coli cell harbors the selection plasmid with the inducible toxic gene induced by Arabinose. The cell can survive in the absence of Arabinose but not in the presence of Arabinose. (FIG. 4B) Transforming another plasmid encoding TnpB and coRNA targeting the toxic gene will result in survival of the cell.
[0014] FIGS. 5A-5C. (FIG. 5A) The distribution of amino acid lengths of Cas9, Casl2, IscB and TnpB endonucleases. (FIG. 5B) Representative loci architecture of CRISPR-Cas systems versus TnpB-associated transposon systems (FIG. 5C) Molecular mechanism of DNA targeting by CRISPR-Cas enzymes and TnpB-coRNA.
[0015] FIG. 6. An example of the split expression cassette to express the coRNA and the TnpB as described in Example 8. In this example, the UBQ10 promoter is used to drive expression of the TnpB, followed by the rbcS-E9 terminator. Two SV40 NLS signals and two FLAG tags are located at the 3’ end, however these sequences may also be tested at the 5’ end of the TnpB. In this example the coRNA is being driven by the U6 promoter, with an HDV ribozyme sequence downstream of the spacer; however, a variety of promoters and RNA processing sequences may be tested.
[0016] FIG. 7A and 7B. (FIG. 7A) An example schematic of the TnpB and coRNA architecture commonly found in nature, where the coRNA overlaps with the TnpB at the 3’ end of the TnpB sequence. (FIG. 7B) An example of a single expression cassette to express the TnpB and coRNA. In this example the UBQ10 promoter is used to drive expression of the TnpB-coRNA sequence, and terminated by the rbcS-E9 terminator. In this example the HDV ribozyme sequence downstream of the spacer is depicted, however a variety of RNA processing systems such as tRNA-Gly may be used.
[0017] FIG. 8. An example of the single expression cassette to express the TnpB and coRNA, followed by a tandem repeat of the coRNA and spacer sequences. In this example theUBQ10 promoter is used to drive expression of the TnpB-coRNA sequence, and terminated by the rbcS-E9 terminator. In this example the HDV ribozyme sequence upstream of the terminator is depicted, however a variety of RNA processing systems such as tRNA-Gly may be used.
[0018] FIGS. 9A-9D. TRV2 schematic depicting the RNA virus sequence. The red box indicates the part of the virus where TnpB and coRNA will most likely be cloned. The light gray boxes indicate the promoter and terminator used to drive initial expression of the RNA virus. The light blue boxes indicate the RNA virus. (FIG. 9A) TnpB-ojRNA single expression cassette cloned into the “cargo’’ site. (FIG. 9B) TnpB and coRNA split expression design. (FIG. 9C) TRV2 schematic showing an example of a RNA processing sequence, HDV ribozyme. (FIG. 9D) TnpB and tandem coRNA cloned into the “cargo” design.
[0019] FIG. 10A-10C. ISYmul, ISDra2 and ISAaml editing frequency. The target sites are plotted along the X-axis and the editing efficiency (percent indel reads (%)) are plotted on the Y-axis. For ISYmul and ISDra2, additional transfections w ere performed to replicate initial results, as indicated by “experiment_2” and “experiment_3” along the X-axis. Each dot indicates a single transfection. The standard error of the mean (SEM) was calculated for each target site.
[0020] FIGS. 11. Editing frequency for TnpBs capable of editing an average greater than 0.1% at least one target site. The target sites are plotted along the X-axis and the editing efficiency (percent indel reads (%)) are plotted on the Y-axis. Each dot indicates a single transfection. The standard error of the mean (SEM) was calculated for each target site.
[0021] FIG. 12A-12D. Small RNA-Seq and structure analysis of ISBagOl and ISYmul ®RNA sequences. (FIG. 12A and FIG. 12B) Small RNA-Seq results for ISYmul and ISBagOl with the read count plotted along the Y-axis, and the TnpB and coRNA positions in the bacterial expression vectors plotted along the X-axis. (FIG. 12C and FIG. 12D). The secondary structures of coRNA for ISYmul and ISBagOl predicted by Vienna RNA Fold (Gruber et al. 2008), consisting of coRNA scaffold (ISYmul=127-nt; ISBag01=186-nt) and guide regions. The stem-loops and pseudoknot (PK) interactions are also indicated on the structures. ISYmul (SEQ ID NO: 626) and ISBagOl (SEQ ID NO: 627).
[0022] FIG. 13A and 13B. ISBagOl ®RNA variants for improved editing efficiency. (FIG. 13A) coRNA (WT, SEQ ID NO: 628) variant structure for 1.1 (SEQ ID NO: 629) and 1.2 (SEQ ID NO: 630). (FIG. 13B) ISBagOl gl was used in this experiment. The target sites are plotted along the X-axis and the editing efficiency (percent indel reads (%)) are plotted on the Y-axis. The X-axis label gl_native indicates the TnpB with coRNA overlap Native DNA sequence;gl_variant_l. l and gl_variant_1.2 are novel roRNA variants. Each dot indicates a single transfection. The standard error of the mean (SEM) was calculated for each target site. P value was calculated using a two-tailed t-Test assuming equal variances.
[0023] FIG. 14. Editing frequency for ISYmul using TRV vectors for genome editing in Arabidopsis protoplast cells. The ISYmul TRV2 cargo configurations are plotted along the X- axis and the editing efficiency (percent indel reads (%)) are plotted on the Y-axis. The vertical dashed line separates experiment one and two. Each dot indicates a single transfection. The standard error mean (SEM) was calculated for each target site.
[0024] FIG. 15A-15G. Gene editing outcomes in transgenic T1 plants expressing ISYmul or ISDra2 TnpB. (FIG. 15A-15D) The genotypes are plotted along the X-axis and the editing efficiency (percent indel reads (%)) are plotted on the Y-axis. Each dot indicates an independent T1 plant. The standard error of the mean (SEM) was calculated for each target site. The number above each bar indicates the average editing frequency. WT is an abbreviation for wild type, and rdr6 is an abbreviation for the mutant in the RNA-DEP ENDENT RNA POLYMERASE 6 gene. (FIG. 15E-15G) DNA repair indel profile for ISDra2 g9 and ISYmul g!2 wildtype samples. The six. two, and twenty most common indels are shown for ISDra2 g9, ISDra2 gl2, and ISYmul gl2, respectively. The indel type is listed on the left. For example - 3 means a deletion of three base pairs. The read counts for each indel are listed on the right. The TAM is identified by the red box, and the target site is located by the black box in the Reference sequence. ISDra2 g9 WT reference sequence is SEQ ID NO: 631, amplicons in descending order are SEQ ID NOs: 632-637. ISDra2 g!2 WT reference sequence is SEQ ID NO: 638, amplicons in descending order are SEQ ID NO 639-640. ISYmul g!2 WT reference sequence is SEQ ID NO: 641, amplicons in descending order are SEQ ID NOs: 642-661.
[0025] FIG. 16A-16E. TRV delivery of ISYmul for editing in whole plants. (FIG. 16A and 16B) Violin plots depicting the genotypes along the X-axis and the editing efficiency (percent indel reads (%)) on the Y-axis. Each dot indicates an independent plant. The standard error of the mean (SEM) was calculated for each target site as indicated by the blue dot and vertical line. WT is an abbreviation for wild type, and rdr6 is an abbreviation for mutant in RNA-DEPENDENT RNA POLYMERASE 6 gene. (FIG. 16C and 16D) DNA repair indel profile for ISYmul g2. The indel type is listed on the left. The read counts for each indel are listed on the right. The TAM is identified by the red box, and the target site is located by the black box in the Reference sequence. ISYmul -g2-tRNA reference sequence is SEQ ID NO: 662, amplicons in descending order are SEQ ID NOs: 663-672. ISYmul-g2-HDV-tRNA reference sequence is SEQ ID NO: 673, amplicons in descending order are SEQ ID NOs: 674-686. (FIG. 16E) Examples of plants infected with TRV vectors showing white sectors, an ISYmul-g2-tRNA-Isoleucine TRV infected plant (top) and an ISYmul -g2-HDV-tRNA- Isoleucine TRV infected plant (bottom). Yellow arrow points to a leaf with clear white sectors.
[0026] FIGS. 17A-17E. Testing activities of novel TnpBs in bacteria. (FIG. 17A) Cassette of dual plasmids for plasmid interference assay. (FIG. 17B) Normalized CFU (Norm. CFU) of five identified novel TnpBs. All the interference assays were performed at 26°C. Error bars represent the standard error of the mean (SEM) of three replicas. (FIG. 17C) TAM sequences of TnpB_30 from the TAM depletion assay. (FIG. 17D) Small RNA-seq confirmed expression of the coRNA of TnpB_30. (FIG. 17E) Secondary structures of TnpB_30 coRNA (SEQ ID NO 687).
[0027] FIG. 18. TnpB_30 editing efficiency m Arabidopsis protoplast cells. The target sites are plotted along the X-axis and the editing efficiency (percent indel reads (%)) is plotted on the Y-axis. Each dot indicates a single transfection. The standard error of the mean (SEM) was calculated for each target site.
[0028] FIGS. 19A-19D. Spacer length optimization for ISDra2, ISYmul. ISBagOl, and TnpB_30. The spacer lengths (nucleotides) are plotted along the X-axis and the editing efficiency (percent indel reads (%)) is plotted on the Y-axis. Each dot indicates a single transfection. The standard error of the mean (SEM) was calculated for each target site. (FIG. 19A) ISDra2. (FIG. 19B) ISYmul. (FIG. 19C) ISBagOl. (FIG. 9D) TnpB_30.
[0029] FIG. 20A-20C. Comparison of ISYmul expression using the single or split expression cassette designs. (FIG. 20A) Plasmid design for the single and split plasmids. Green arrow boxes symbolize promoters, red boxes symbolize terminators, and the black arrows indicate the orientation of the expression cassette. (FIG. 20B) Small RNA-seq results for ISYmul with the read count plotted along the Y-axis, and the TnpB and wRNA positions in the bactenal expression vectors plotted along the X-axis. (FIG. 20C) TnpB and wRNA single and split expression configurations are plotted along the X-axis and the editing efficiency (percent indel reads (%)) is plotted on the Y-axis. Each dot indicates a single transfection. The standard error of the mean (SEM) w as calculated for each target site.
[0030] FIG. 21. ISYmul wRNA variants for improved editing efficiency. (A) ISYmul g2 was used in this experiment. The coRNA variants are plotted along the X-axis and the editing efficiencies (percent indel reads (%)) are plotted on the Y-axis. The X-axis label WT vO.O indicates the TnpB with coRNA overlap native DNA sequence; all other variants along the X- axis are novel ISYmul wRN A variants. Each dot indicates a single transfection. The standard error of the mean (SEM) w as calculated for each coRNA design tested.
[0031] FIG. 22. Representative ISYmul wRNA variants at stem-loop 2. The stem-loop 2 are highlighted in the secondary structures of ISYmul wRNA variants. Representative in RNA variants of stem-loop 2 are compared to the wild-type sequences. The topologies of the RNA two-way junctions are labeled within the region of interest. Stem-loop 2 (top) is SEQ ID NO: 688, WT is SEQ ID NO: 689, v3.2 is SEQ ID NO: 690, v3. 16 is SEQ ID NO: 691, v2. 1 is SEQ ID NO: 692.
[0032] FIG. 23A and 23B. ISYmul somatic editing in Tl transgenic plants. ISYmul g2 and ISYmul gl2 were used in this experiment. Panels A and B display box and whisker plots. Each dot indicates a single T1 transgenic plant. The room and HS treatments stand for room temperature and heat shock plant growth conditions, respectively. The genotypes are plotted along the X-axis and the editing efficiencies (percent indel reads (%)) are plotted on the Y- axis. WT is an abbreviation for wild type, and rdr6 is an abbreviation for mutant in RNA- DEP ENDENT RNA POLYMERASE 6 gene. (FIG. 23 A) ISYmul g2 Tl. (FIG. 23B) ISYmul g!2 Tl.
[0033] FIG. 24A and 24B. ISDra2 somatic editing in T1 transgenic plants. ISDra2 g9 and ISDra2 g!2 were used in this expenment. Panels A and B display box and whisker plots. Each dot indicates a single Tl transgenic plant. The room and HS treatments stand for room temperature and heat shock plant growth conditions, respectively. The genotypes are plotted along the X-axis and the editing efficiencies (percent indel reads (%)) are plotted on the Y- axis. WT is an abbreviation for wild type, and rdr6 is an abbreviation for mutant in RNA- DEPENDENT RNA POLYMERASE 6 gene. (FIG. 24A) ISDra2 g9 Tl. (FIG. 24B) ISDra2 g!2 Tl.
[0034] FIG. 25A-25D. TRV delivery of ISYmul for editing in whole plants. (FIG. 25A) Schematic of the TRV1 and TRV2 plasmids. (FIG. 25B and 25C) The TRV cargo configurations are plotted along the X-axis and the editing efficiency (percent indel reads (%)) is plotted on the Y -axis. Each dot indicates a single plant. The standard error of the mean (SEM) was calculated for each target site. WT samples are black, rdr6 samples are blue, and ku70 samples are red. WT is an abbreviation for wild type, and rdr6 is an abbreviation for mutant in RNA-DEPENDENT RNA POLYMERASE 6 gene. (FIG. 25B) ISYmul g2. (FIG. 25C) ISYmul gl2. (FIG. 25D) DNA repair indel profile for ISYmul g!2 using TRV to deliver ISYmul -gl2-HDV-tRNA-Isoleucine. The top six most common indel types are listed on the left. The read counts for each indel are listed on the right. The TAM is identified by the red box, and the target site is located by the black box in the reference sequence. The reference sequence is SEQ ID NO: 693, amplicons in descending order are SEQ ID NOs: 694-699.
[0035] FIG. 26A-26C. Heritability of edits generated with ISYmul g2 encoded on the TRV vector. (FIG. 26 A) Picture of a plant that underwent TRV delivery’ using ISYmul -g2-HDV- tRNA-Isoleucine. The white sectors, indicated by the yellow arrows, indicate biallelic mutations in the PDS3 gene. The 54.54% in the upper left comer is the editing frequency determined by random tissue sampling of three leaves distal from the TRV delivery’ location. (FIG. 26B) Image of progeny seedlings from the plant in Figure 26A, depicting two albino seedlings. (FIG. 26C) Sanger sequencing trace file screenshot for one of the albino plants in FIG. 26B (top sequence, SEQ ID NO: 700, bottom sequence SEQ ID NO: 701).
[0036] FIG. 27A-27F. Testing the impact of wRNA processing on the activities of TnpB in bacteria. (FIG. 27A) Different cassettes of TnpB and wRNA expression in bacteria. (FIG. 27B) Normalized CFU of ISDra2, ISBagOl. ISYmul and ISAaml TnpB expressed in cassette (a). (FIG. 27C) Normalized CFU of ISDra2, ISBagOl, ISYmul and ISAaml TnpB expressed in cassette (b). (FIG. 27D) Normalized CFU of ISBagOl expressed in cassette (c) and (d). (FIG. 27E) Normalized CFU of ISYmul expressed in cassette (c) and (e). (FIG. 27F) Normalized CFU of ISYmul expressed in cassette (f) and (g). All the interference assays were performed in 26°C. Error bars represent the standard error of the mean (SEM) of three replicas.DETAILED DESCRIPTION
[0037] Overview
[0038] The most common method of delivery of gene editing reagents into plants is to encode RNA-guided genome editors (e.g. CRISPR-Cas enzymes) within transgenes and use standard tissue culture transformation approaches (Altpeter et al., 2016; Chen et al., 2019; Nash & Voytas, 2021). However, this approach has substantial disadvantages. Tissue culture methods require considerable time, resources, technical expertise, and can cause unintended changes to the genome and epigenome (Altpeter et al., 2016). Further, tissue culture is restricted in its application, as only a limited number species and genotypes are amenable to this technique (Altpeter et al., 2016). As such, approaches have recently been developed to circumvent tissue culture transformation, such as de novo induction of meristems (Maher et al., 2020), the ‘‘graft-mobile'’ gene editing system(Y ang et al., 2023), and viral delivery’ of editing reagents (Ellison et al.. 2020; Ghoshal et al.. 2020; Liu et al., 2022; Nagalakshmi et al.. 2022). These techniques hint at the potential to overcome this bottleneck; however, they are limited by the need for pre-established transgenic lines, they require highly experienced personnel, are likely difficult to scale, and appear to be species and genotype dependent.
[0039] A second, slightly more efficient way that works in some crops, is to introduce CRISPR protein and guide RNAs directly into plant cells during the tissue culture process,followed by plant regeneration and backcrossing to eliminate epigenetic variation induced bytissue culture. However, this method also requires difficult tissue culture procedures that are costly, time consuming, and that only work in some crop species.
[0040] Plant viruses are ideal vectors for delivering editing reagents into plants given their natural ability' to amplify and spread throughout the plant. In addition, many plant viruses, for example the broad host range RNA virus Tobacco rattle virus (TRV). are able to infect almost all parts of the plant, and yet the virus is rarely transmitted into future sexual generations because viral genome does not integrate into the plant genome (Bradamante et al., 2021 ). Recently, it has been shown that delivery- of single guide RNA (sgRNA) sequences fused to mobile tRNA sequences encoded in TRV vectors can induce highly efficient biallelic gene editing or epigenome editing in plants (Ellison et al., 2020; Ghoshal et al., 2020; Nagalakshmi et al., 2022). However, delivery of both the Cas nuclease and the sgRNA has not yet been possible with TRV or other viruses as their low cargo capacity- (-1.5-2 kb RNA) limits the size of the Cas nuclease to be fewer than 400 amino acids (aa). For example, ,S / n'Cas9 (1369 aa) and / ACasl2a (1228 aa). two of the most extensively utilized CRISPR-Cas genome editors in plants, are substantially larger than this cargo limit (FIG. 5A). Even the most compact Cas9d (750 aa) (Aliaga Goltsman et al., 2022) and Casl2 (420 aa) (Bigelyte et al., 2021 ; Wu et al., 2021) genome editors still don’t meet such a stringent limit. Several ancestral Cas9 and Casl2 enzymes (-400 aa or less) have recently been demonstrated to be RNA-guided endonucleases with genome editing activities in mammalian cells (Altae-Tran et al.. 2021; Karvelis et al., 2021). These ancestral enzymes, including IscB (Cas9 ancestors) and TnpB (Casl2 ancestors), are derived from prokaryotic transposons (e.g., IS200 / IS605) rather than CRISPR systems (FIG. 5B). Yet, they share common structural architectures and similar mechanisms of RNA- guided DNA cleavage with CRISPR-Cas9 and Casl2 (FIG. 5C). For example, the target DNA binding and cleavage of both systems involve the recognition of a protospacer adjacent motif (PAM) or transposon adjacent motif (TAM) and the formation of an J?-loop structure betw een the target DNA and guide RNA (FIG. 5C) (Nakagawa et al., 2023; Sasnauskas et al., 2023; Schuler et al., 2022). In contrast to CRISPR-Cas enzymes, these ancestral proteins typically form a complex with a relatively large (-200 nt) guide RNA called OMEGA (for obligate mobile element-guided activity) RNA (coRNA) (Altae-Tran et al., 2021) which is encoded at either the left or right boundary- of the transposons (FIG. 5B). Interestingly, while the PAM sequences of CRISPR-Cas systems are often difficult to predict, due to the natural function of these transposon systems, the TAM sequences of these ancestral RNA-guided systems are directly encoded on one end of the transposon boundary, providing a catalog of ‘ TAM”sequences from nature (FIG. 5B). Compared to IscB, TnpB enzymes bear significantly larger diversities (Altae-Tran et al., 2021) and higher genome editing efficiencies (Altae-Tran et al., 2021; Karvelis et al., 2021) and the more compact size of TnpB offers advantages for the viral delivery underlying our strategy (FIG. 5A).
[0041] The present disclosure herein relates, in part, to one or more small TnpB type bacterial enzymes that are less than about 500 amino acids in length and are capable of editing plant cells. An advantage of this small enzyme is that it is small enough to be cloned in a plant viral vector, such as for example TRV. This allows for an easier method to edit plant genomes that does not require tissue culture or plant transformation.
[0042] Definitions and Terminology
[0043] The disclosed polypeptides, compositions, constructs, viral vectors, and methods for plant gene editing may be further described using definitions and terminology as follows. The definitions and terminology used herein are for the purpose of describing particular aspects only and are not intended to be limiting.
[0044] As used in this specification and the claims, the singular forms “a,” ”an.” and “the” include plural forms unless the context clearly dictates otherwise.
[0045] As used herein, “about”, “approximately,” “substantially,” and “significantly” will be understood by persons of ordinary skill in the art and w ill vary to some extent on the context in which they are used. If there are uses of the term which are not clear to persons of ordinary skill in the art given the context in which it is used, “about” and “approximately” will mean up to plus or minus 10% of the particular term and “substantially” and “significantly” will mean more than plus or minus 10% of the particular term.
[0046] As used herein, the terms “include” and “including” have the same meaning as the terms “comprise” and “comprising.” The terms “comprise” and “comprising” should be interpreted as being “open” transitional terms that permit the inclusion of additional components further to those components recited in the claims. The terms “consist” and “consisting of’ should be interpreted as being “closed” transitional terms that do not permit the inclusion of additional components other than the components recited in the claims. The term “consisting essentially of’ should be interpreted to be partially closed and allowing the inclusion only of additional components that do not fundamentally alter the nature of the claimed subj ect matter.
[0047] The phrase “such as” should be interpreted as “for example, including.” Moreover, the use of any and all exemplary language, including but not limited to “such as”, is intendedmerely to better illuminate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed.
[0048] Furthermore, in those instances where a convention analogous to “at least one of A,B and C, etc.” is used, in general such a construction is intended in the sense of one having ordinary skill in the art would understand the convention (e.g. , “a system having at least one of A, B and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description or figures, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B.”
[0049] All language such as “up to,” “at least,” “greater than,” “less than,” and the like, include the number recited and refer to ranges which can subsequently be broken down into ranges and subranges. A range includes each individual member. Thus, for example, a group having 1-3 members refers to groups having 1, 2. or 3 members. Similarly, a group having 6 members refers to groups having 1, 2, 3, 4, or 6 members, and so forth.
[0050] The modal verb “may” refers to the preferred use or selection of one or more options or choices among the several described embodiments or features contained within the same. Where no options or choices are disclosed regarding a particular embodiment or feature contained in the same, the modal verb “may” refers to an affirmative act regarding howto make or use an aspect of a described embodiment or feature contained in the same, or a definitive decision to use a specific skill regarding a described embodiment or feature contained in the same. In this latter context, the modal verb “may” has the same meaning and connotation as the auxiliary verb “can.”
[0051] The terms “identical” or percent “identity,” in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (i.e., 50%, 55%, 60%, 65%, 70%. 75%. 80%. 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identity over a specified region, e.g., of an entire nucleic acid or polypeptide sequence or individual portions or domains of a nucleic acid or polypeptide), when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection. Such sequences are then said to be “substantially identical.”This definition also refers to the complement of a test sequence, in the context of nucleic acids. By way of example, in embodiments, the identify exists over a region that is about or at least about 5, 10, 15, 20, 50, 100, or 1000, amino acids in length, to about, less than about, or at least about 220, 100 or 1000 amino acids or nucleotides in length. Optionally, the identity exists over a region that is at least about 5, 10, 15, or 16 amino acids in length to about 100, about 20 to about 75, about 30 to about 50 amino acids or nucleotides in length.
[0052] For sequence comparison, typically one sequence acts as a reference sequence, to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary', and sequence algorithm program parameters are designated. Preferably, default program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters.
[0053] An example of algorithms suitable for determining percent sequence identify and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al., Nuc. Acids Res. 25:3389-3402 (1977) and Altschul et al., J. Mol. Biol. 215:403-410 (1990), respectively. As will be appreciated by one of skill in the art, the software for performing BLAST analyses is publicly available through the website of the National Center for Biotechnology Information (NCBI). In embodiments, BLAST and BLAST 2.0 are used, with the parameters described herein, to determine percent sequence identify for the nucleic acids and proteins. In embodiments, a BLAST algorithm involves first identifying high scoring sequence pairs (HSPs) by identify ing short words of length W in the query sequence, which either match or satisfy' some positive-valued threshold score T when aligned wi th a word of the same length in a database sequence. In embodiments, T is referred to as the neighborhood word score threshold (Altschul et al., supra). In embodiments, these initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. In embodiments, the word hits are extended in both directions along each sequence for as far as the cumulative alignment score can be increased. In embodiments, cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty' score for mismatching residues; always <0). In embodiments, for amino acid sequences, a scoring matrix is used to calculate the cumulative score. In embodiments, extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantify X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or theend of either sequence is reached. In embodiments, the BLAST algorithm parameters W. T, and X determine the sensitivity and speed of the alignment. In embodiments, the NCBI BLASTN or BLASTP program is used to align sequences. In embodiments, the BLASTN or BLASTP program uses the defaults used by the NCBI. In embodiments, the BLASTN program (for nucleotide sequences) uses as defaults: a word size (W) of 28; an expectation threshold (E) of 10; max matches in a query range set to 0; match / mismatch scores of 1. -2; linear gap costs; the filter for low complexity regions used; and mask for lookup table only used. In embodiments, the BLASTP program (for amino acid sequences) uses as defaults: a word size (W) of 3; an expectation threshold (E) of 10; max matches in a query range set to 0; the BLOSUM62 matrix (see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89: 10915 (1992)); gap costs of existence: 11 and extension: 1; and conditional compositional score matrix adjustment.
[0054] The terms “polypeptide,” “peptide” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. The terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymer.
[0055] The term “amino acid” refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, y- carboxyglutamate, and O-phosphoserine. Amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Amino acid mimetics refers to chemical compounds that have a structure that is different from the general chemical structure of an amino acid, but that functions in a manner similar to a naturally occurring amino acid.
[0056] Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes.
[0057] “Conservatively modified variants” applies to both amino acid and nucleic acid sequences. With respect to particular nucleic acid sequences, conservatively modified variants refers to those nucleic acids which encode identical or essentially identical amino acid sequences, or where the nucleic acid does not encode an amino acid sequence, to essentially identical sequences. Because of the degeneracy of the genetic code, a large number of functionally identical nucleic acids encode any given protein. For instance, the codons GCA, GCC, GCG and GCU all encode the amino acid alanine. Thus, at every position where an alanine is specified by a codon, the codon can be altered to any of the corresponding codons described without altering the encoded polypeptide. Such nucleic acid variations are “silent variations,” which are one species of conservatively modified variations. Every nucleic acid sequence herein which encodes a polypeptide also describes every possible silent variation of the nucleic acid. One of skill will recognize that each codon in a nucleic acid (except AUG, which is ordinarily the only codon for methionine, and TGG, which is ordinarily the only codon for tryptophan) can be modified to yield a functionally identical molecule. Accordingly, each silent variation of a nucleic acid which encodes a polypeptide is implicit in each described sequence with respect to the expression product, but not with respect to actual probe sequences.
[0058] As to amino acid sequences, one of skill will recognize that individual substitutions to a peptide, polypeptide, or protein sequence which alters a single amino acid is a “conservatively modified variant” where the alteration results in the substitution of an amino acid with a chemically similar amino acid. Conservative substitution tables providing functionally similar amino acids are well known in the art. Such conservatively modified variants are in addition to and do not exclude polymorphic variants, interspecies homologs, and alleles.
[0059] The following eight groups each contain amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G); 2) Aspartic acid (D). Glutamic acid (E); 3) Asparagine (N), Glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); 6) Phenylalanine (F), Tyrosine (Y), Try ptophan (W); 7) Serine (S), Threonine (T); and 8) Cysteine (C), Methionine (M) (see, e.g., Creighton. Proteins (1984)).
[0060] The terms “polynucleotide,” “polynucleotide sequence,” “nucleic acid” and “nucleic acid sequence” refer to a nucleotide, oligonucleotide, polynucleotide (which terms may be used interchangeably), or any fragment thereof. These phrases also refer to DNA or RNA of genomic, natural, or synthetic origin (which may be single-stranded or double-stranded and may represent the sense or the antisense strand).
[0061] The terms “nucleic acid” and “oligonucleotide,” as used herein, may refer to polydeoxyribonucleotides (containing 2-deoxy-D-ribose). polyribonucleotides (containing D- ribose), and to any other type of polynucleotide that is an N glycoside of a purine or pyrimidine base. There is no intended distinction in length between the terms “nucleic acid”, “oligonucleotide” and “polynucleotide”, and these terms will be used interchangeably. These terms refer only to the primary structure of the molecule. Thus, these terms include double- and single-stranded DNA, as well as double- and single-stranded RNA. For use in the present methods, an oligonucleotide also can comprise nucleotide analogs in which the base, sugar, or phosphate backbone is modified as well as non-purine or non-pyrimidine nucleotide analogs.
[0062] Oligonucleotides can be prepared by any suitable method, including direct chemical synthesis by a method such as the phosphotriester method of Narang et al., 1979, Meth. Enzymol. 68:90-99; the phosphodiester method of Brown etal., (919, Meth. Enzymol. 68:109- 151; the diethylphosphoramidite method of Beaucage et al., 1981, Tetrahedron Letters 22: 1859-1862; and the solid support method of U.S. Pat. No. 4,458,066, each incorporated herein by reference. A review of synthesis methods of conjugates of oligonucleotides and modified nucleotides is provided in Goodchild, 1990. Bioconjugate Chemistry 1(3): 165-187, incorporated herein by reference.
[0063] Regarding polynucleotide sequences, the terms “percent identity” and “% identity ” refer to the percentage of residue matches between at least two polynucleotide sequences aligned using a standardized algorithm. Such an algorithm may insert, in a standardized and reproducible way, gaps in the sequences being compared in order to optimize alignment between two sequences, and therefore achieve a more meaningful comparison of the two sequences. Percent identity' for a nucleic acid sequence may be determined as understood in the art. (See, e.g., U.S. Patent No. 7,396,664, which is incorporated herein by reference in its entirety). A suite of commonly used and freely available sequence comparison algorithms is provided by the National Center for Biotechnology Information (NCBI) Basic Local Alignment Search Tool (BLAST), which is available from several sources, including the NCBI, Bethesda, Md.. at its website. The BLAST software suite includes various sequence analysis programs including “blastn,” that is used to align a known polynucleotide sequence with other polynucleotide sequences from a variety of databases. Also available is a tool called “BLAST 2 Sequences” that is used for direct pairwise comparison of two nucleotide sequences. “BLAST 2 Sequences” can be accessed and used interactively at the NCBI website. The “BLAST 2 Sequences” tool can be used for both blastn and blastp (discussed above).
[0064] Regarding polynucleotide sequences, percent identity may be measured over the length of an entire defined polynucleotide sequence, for example, as defined by a particular SEQ ID number, or may be measured over a shorter length, for example, over the length of a fragment taken from a larger, defined sequence, for instance, a fragment of at least 20, at least 30, at least 40, at least 50, at least 70, at least 100, or at least 200 contiguous nucleotides. Such lengths are exemplary only, and it is understood that any fragment length supported by the sequences shown herein, in the tables, figures, or Sequence Listing, may be used to describe a length over which percentage identity may be measured.
[0065] Regarding polynucleotide sequences, “variant,” “mutant,” or “derivative” may be defined as a nucleic acid sequence having at least 50% sequence identity to the particular nucleic acid sequence over a certain length of one of the nucleic acid sequences using blastn with the “BLAST 2 Sequences” tool available at the National Center for Biotechnology Information’s website. (See Tatiana A. Tatusova, Thomas L. Madden (1999), "Blast 2 sequences - a new tool for comparing protein and nucleotide sequences", FEMS Microbiol Lett. 174:247-250). Such a pair of nucleic acids may show, for example, at least 60%, at least 70%. at least 80%. at least 85%. at least 90%. at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or greater sequence identity over a certain defined length.
[0066] Nucleic acid sequences that do not show a high degree of identity' may nevertheless encode similar amino acid sequences due to the degeneracy of the genetic code where multiple codons may encode for a single amino acid. It is understood that changes in a nucleic acid sequence can be made using this degeneracy to produce multiple nucleic acid sequences that all encode substantially the same protein. For example, polynucleotide sequences as contemplated herein may encode a protein and may be codon-optimized for expression in a particular host. In the art, codon usage frequency tables have been prepared for a number of host organisms including humans, mouse, rat, pig, E. coli, plants, and other host cells.
[0067] A “recombinant nucleic acid” is a sequence that is not naturally occurring or has a sequence that is made by an artificial combination of two or more otherwise separated segments of sequence. This artificial combination is often accomplished by chemical synthesis or, more commonly, by the artificial manipulation of isolated segments of nucleic acids, e.g. , by genetic engineering techniques known in the art. The term recombinant includes nucleic acids that have been altered solely by addition, substitution, or deletion of a portion of the nucleic acid. Frequently, a recombinant nucleic acid may include a nucleic acid sequence operably linked toa promoter sequence. Such a recombinant nucleic acid may be part of a vector that is used, for example, to transform a cell.
[0068] The nucleic acids disclosed herein may be '‘substantially isolated or purified.” The term “substantially isolated or purified” refers to a nucleic acid that is removed from its natural environment, and is at least 60% free, preferably at least 75% free, and more preferably at least 90% free, even more preferably at least 95% free from other components with which it is naturally associated.
[0069] The polynucleotide sequences contemplated herein may be present in expression vectors. For example, the vectors may comprise a polynucleotide encoding an ORF of a protein operably linked to a promoter. “Operably linked” refers to the situation in w hich a first nucleic acid sequence is placed in a functional relationship with a second nucleic acid sequence. For instance, a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence. Operably linked DNA sequences may be in close proximity or contiguous and, where necessary to join two protein coding regions, in the same reading frame. Vectors contemplated herein may comprise a heterologous promoter operably linked to a polynucleotide that encodes a protein. A “heterologous promoter” refers to a promoter that is not the native or endogenous promoter for the protein or RNA that is being expressed.
[0070] As used herein, "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (such as into mRNA or another RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides may be collectively referred to as "gene product. "
[0071] The term “vector” refers to some means by which nucleic acid (e.g, DNA) can be introduced into a host organism or host tissue. There are various types of vectors including plasmid vector, bacteriophage vectors, cosmid vectors, bacterial vectors, and viral vectors. As used herein, a “vector” may refer to a recombinant nucleic acid that has been engineered to express a heterologous polypeptide (e.g, the fusion proteins disclosed herein). The recombinant nucleic acid typically includes c / .s- acting elements for expression of the heterologous polypeptide.
[0072] Compositions, Constructs, Viral Vectors, and Plant Cells
[0073] In certain aspects, disclosed herein are compositions for plant gene editing. In various aspects, the composition can include a TnpB bacterial enzyme. In various aspects, the TnpB bacterial enzyme can be from any bacterial TnpB system. In certain aspects, the TnpBbacterial enzy me can have a length of less than about 500 amino acids, less than about 450 amino acids, less than about 420 amino acids, or less than about 400 amino acids.
[0074] In various aspects, the TnpB can be from, or derived from, Brevibacillus agri. A TnpB derived from Brevibacillus agri refers to a variant of the TnpB of Brevibacilhis agri. In the same or alternative aspects, the TnpB can have an amino acid sequence that is 70 % or more, 75 % or more. 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the amino acid sequence of the TnpB of Brevibacillus agri. In various aspects, the TnpB can have an amino acid sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the amino acid sequence of SEQ ID NO: 5. In various aspects, the TnpB can have the amino acid sequence of SEQ ID NO: 5. It should be understood that a variant of the TnpB of Brevibacillus agri can be a synthetically designed variant or can be a TnpB that is of similar sequence but from another species or genus of bacteria.
[0075] In various aspects, the TnpB can be from, or derived from, Xylella fastidiosa. A TnpB derived from Xylella fastidiosa refers to a variant of the TnpB of Xylella fastidiosa. In the same or alternative aspects, the TnpB can have an amino acid sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the amino acid sequence of the TnpB of Xylella fastidiosa. In various aspects, the TnpB can have an amino acid sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more. 90 % or more, 95 % or more, or 99 % or more identical to the amino acid sequence of SEQ ID NO: 10. In various aspects, the TnpB can have the amino acid sequence of SEQ ID NO: 10. It should be understood that a variant of the TnpB of Xylella fastidiosa can be a synthetically designed variant or can be a TnpB that is of similar sequence but from another species or genus of bacteria.
[0076] In various aspects, the TnpB can have an amino acid sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more, or 100 % identical to one or more of the amino acid sequences of SEQ ID NOs: 46, 47, 48, 49, and 702.
[0077] In various aspects, the TnpB can comprise a peptide tag. In such aspects, the peptide tag can include any suitable peptide tag known in the art. As one example aspect, the peptide tag can be a FLAG tag. In some aspects, the peptide tag(s) is located at the N-terminus of the TnpB protein, while in others, the peptide tag(s) is located at the C-terminus of the TnpB protein, while in still further aspects, the peptide tag(s) is located within the interior portion of the TnpB, or any combination thereof.
[0078] In various aspects, the compositions disclosed herein can also include a polynucleotide. In various aspects, the polynucleotide can include a first portion adapted to bind to at least a portion of the TnpB and a second portion for binding to a plant genomic target. In various aspects, the polynucleotide can be RNA. In one or more aspects, the polynucleotide can comprise an coRNA. In one or more aspects, the coRNA is about 100-400 nucleotides in length, or about 200-300 nucleotides in length. In various aspects, a DNA sequence that is 70 % or more. 75 % or more, 80 % or more, 85 % or more. 90 % or more, 95 % or more, or 99 % or more, or 100% identical to the polynucleotide sequence of one or more of SEQ ID NOs: 11 - 13, 50-53, 359-430, and 580-606 can encode for the coRNA. In certain aspects, the polynucleotide can include a spacer element. In various aspects, the space element can be any length suitable for use in plant gene editing using the systems and methods disclosed herein. It should be understood that in alternative aspect, the first portion of the polynucleotide and the second portion of the polynucleotide can also be present as separate polynucleotides. In such an aspect, the separate polynucleotides may include sequences to allow for joining or hybridization in a plant cell.
[0079] In various aspects, constructs are disclosed herein. The constructs can include a plant promoter operably connected to a polynucleotide encoding for a TnpB. The TnpB encoded for in the polynucleotide can exhibit any or all of the properties discussed above with respect to TnpB.
[0080] In various aspects, the polynucleotide encoding for a TnpB can include a polynucleotide sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the polynucleotide sequence of one or more of SEQ ID NOs: 1-4, 6-9, or 34-45. In various aspects, the polynucleotide encoding for a TnpB can have the polynucleotide sequence of one or more of SEQ ID NOs: 1-4, 6-9, or 34- 45.
[0081] In various aspects, the polynucleotide encoding for a TnpB can be codon optimized for expression in plant cells. In one or more aspects, the polynucleotide encoding for a TnpB can be Zea mays codon optimized.
[0082] As discussed above, in various aspects, the construct can include a plant promoter operably connected to the polynucleotide encoding for a TnpB. The plant promoter can be any suitable promoter for expressing the TnpB in the plant cell of interest. In one example aspect, the plant promoter can include a UBQ10 gene promoter.
[0083] In certain aspects, the polynucleotide can further encode for an coRNA having a first portion adapted to bind to at least a portion of the TnpB bacterial enzyme and a second portionfor binding to a plant genomic target. The coRNA can include any or all of the features and parameters discussed above with respect to coRNA. In one or more aspects, the polynucleotide encoding for TnpB and the coRNA can utilize the same promoter and / or be part of the same cassette. In various aspects, a sequence of the coRNA and a sequence of the TnpB bacterial enzyme can at least partly overlap in a cassette or construct. For instance, in certain aspects, a polynucleotide that encodes for an coRNA and a TnpB that at least partly overlap can have a sequence that is 70 % or more, 75 % or more. 80 % or more, 85 % or more, 90 % or more. 95 % or more, or 99 % or more identical to the polynucleotide sequence of one or more of SEQ ID NOs: 140-147, 150-170, 431-433, or 475-481. In alternative aspects, a sequence of the coRNA and a sequence of the TnpB bacterial enzyme can be present as a split cassette and under the control of separate promoters.
[0084] In various aspects, the polynucleotide encoding for an coRNA can be a separate polynucleotide and / or construct than the one encoding for the TnpB.
[0085] In various aspects, plant viral vectors are disclosed. In one or more aspects, the plant viral vector can include one or more of the constructs mentioned above. In various aspects, the plant viral vector can be any suitable vector for use in plant gene editing. In one example aspect, the plant viral vector can be, or can be derived from, the Tobacco rattle virus (TRV).
[0086] In various aspects, plant cells are disclosed. In certain aspects, the plant cells can include one or more of the constructs mentioned above and / or the plant viral vectors mentioned above. In one or more aspects, the constructs and plant viral vectors can include any or all of the respective parameters and properties described above. For example, the cell can be a cell of a major agricultural plant, e.g., Barley, Beans (Dry Edible), Canola, Com, Cotton (Pima), Cotton (Upland). Flaxseed, Hay (Alfalfa). Hay (Non-Alfalfa), Oats, Peanuts, Rice, Sorghum, Soybeans, Sugarbeets. Sugarcane. Sunflowers (Oil), Sunflowers (Non-Oil), Sweet Potatoes , Tobacco (Burley), Tobacco (Flue-cured), Tomatoes, Wheat (Durum), Wheat (Spring), Wheat (Winter), and the like. As another example, the cell is a cell of a vegetable crops which include but are not limited to. e.g., alfalfa sprouts, aloe leaves, arrow root, arrowhead, artichokes, asparagus, bamboo shoots, banana flowers, bean sprouts, beans, beet tops, beets, bittermelon, bok choy, broccoli, broccoli rabe (rappini), brussels sprouts, cabbage, cabbage sprouts, cactus leaf (nopales), calabaza, cardoon, carrots, cauliflower, celery, chayote, Chinese artichoke (crosnes), Chinese cabbage, Chinese celery, Chinese chives, choy sum, chrysanthemum leaves (tung ho), collard greens, com stalks, com-sweet, cucumbers, daikon, dandelion greens, dasheen, dau mue (pea tips), donqua (winter melon), eggplant, endive, escarole, fiddle headfems, field cress, frisee, gai choy (chinese mustard), gallon, galanga (siam, thai ginger), garlic, ginger root, gobo, greens, hanover salad greens, huauzontle. j erusalem artichokes, jicama, kale greens, kohlrabi, lamb's quarters (quilete), lettuce (bibb), lettuce (boston), lettuce (boston red), lettuce (green leaf), lettuce (iceberg), lettuce (lolla rossa), lettuce (oak leaf - green), lettuce (oak leaf - red), lettuce (processed), lettuce (red leaf), lettuce (romaine), lettuce (ruby romaine), lettuce (russian red mustard), linkok, lo bok, long beans, lotus root, mache, maguey (agave) leaves, malanga. mesculin mix. mizuna, moap (smooth luffa). moo, moqua (fuzzy squash), mushrooms, mustard, nagaimo, okra, ong choy, onions green, opo (long squash), ornamental com, ornamental gourds, parsley, parsnips, peas, peppers (bell type), peppers, pumpkins, radicchio, radish sprouts, radishes, rape greens, rape greens, rhubarb, romaine (baby red), rutabagas, salicomia (sea bean), sinqua (angled / ridged luffa), spinach, squash, straw bales, sugarcane, sweet potatoes, swiss chard, tamarindo, taro, taro leaf, taro shoots, tatsoi, tepeguaje (guaje), tindora, tomatillos, tomatoes, tomatoes (cherry), tomatoes (grape type), tomatoes (plum type), tumeric, turnip tops greens, turnips, water chestnuts, yampi, yams, yu choy, yuca (cassava), and the like. In one aspect, in addition to agricultural plants, ornamental plants, conifer plants, forage and turf grass are also contemplated for use with the systems disclosed herein.In various aspects, any plant cell and / or type of plant may be used in the present disclosure so long as it remains viable after being transformed with a sequence of nucleic acids. In certain aspects, the plant cell is not adversely affected by the transduction of the necessary nucleic acid sequences, the subsequent expression of the proteins or the resulting intermediates.
[0087] Methods
[0088] In various aspects, methods for gene editing in a plant are disclosed. In one or more aspects, the methods can include introducing one or more of the constructs mentioned above and / or one or more of the plant viral vectors mentioned above into one or more plant cells. The contracts and / or plant viral vectors can be introduced into the plant cells using any suitable techniques know n by those of skill in the art.
[0089] The methods for gene editing in a plant can be utilized on any desired plant genus or species. In certain aspects, the methods for gene editing can be performed on any agricultural crop or other plant.EXAMPLES
[0090] The following Examples are illustrative and are not intended to limit the scope of the claimed subject matter.
[0091] Example 1 - Detection of bona-fide RNA-guided TnpB systems in metagenomes
[0092] In order to determine the coRNA architecture in functional RNA-guided TnpB systems and the TAM sequences, the boundaries of the TnpB-associated transposons need to be precisely determined. TnpB associated transposons are moving from one location to another and the genomic regions flanking the transposon ends therefore are different. The exact region of the system can be determined by common strategies such as multiple genome alignments (Altae-Tran et al., 2021). However, multiple genome alignments require the genomes to be from the species or strains that are relatively highly abundant in the corresponding samples or that can be cultivated; otherwise, no genome will be obtained from such analyses. Alternatively, in metagenomic analyses based on short paired-end reads, the same region of the transposon systems can be determined by read mapping (called the read mapping approach hereafter), if only a fraction of the corresponding genomes encode the system while the others do not. Based on this knowledge, a high-throughput computational pipeline was developed based on the read mapping approach to identify the transposon boundaries (LE and RE) and the TAM of the TnpB-associated transposon systems from metagenomic datasets. This pipeline was named ‘"DOTS” (for Detectives of Transposon Systems) (FIG. 1A). The read mapping approach has several advantages over multiple genome alignments: (1) it requires only one metagenomic sample; (2) the species or strains representing the genome do not need to be dominant in the sample; and (3) the analysis can essentially be applied to any assembled sequences, including mobile genetic elements which typically require detection in multiple habitats or samples. In a preliminary analysis, we applied the "DOTS ” pipeline to identify a total of 96 putative RNA-guided IS605 TnpB systems from only 500 ggKbase metagenomes (FIG. IB). These RNA-guided IS605 TnpB systems bear diverse TAM sequences, including CTAT, TCAC and ACAG, which are not present in public IS databases such as ISfinder, have a range of molecular sizes (336-478 aa) (FIG. IB), and have not been reported previously.
[0093] Example 2 -TnpB systems with robust RNA-guided activities in E.coli
[0094] To establish a robust pipeline for experimentally onboarding RNA-guided TnpB systems, ten compact TnpB (< 420 aa) systems were selected that are from publicly accessible genomes on the National Center for Biotechnology Information (NCBI) for reconstituting their RNA-guided DNA cleavage activities in E.coli. These TnpB-associated transposons are all from organisms in ambient environments, such as soil and plant-associated microbes, making them ideal candidates for plant editing. Preliminary . results show that two TnpB systems have robust RNA-guided activities in E.coli'. one is a hypercompact IS200 / IS605 TnpB (359 aa) from Brevibacillus agri (ISBagOl). the other one is an IS607 TnpB (398 aa) from Xylella fastidiosa (ISXfaOl) which is a plant pathogen (FIGS. 2A-2C). Interestingly, the RNA-guidedDNA cleavage activities of TnpB from IS607 transposons have yet to be reported, and our preliminary results suggests that beyond IS200 / IS605, it should be fruitful to mine TnpBs from additional IS607s.
[0095] Example 3 -TnpB activity in Arabidopsis protoplasts
[0096] The ISBagOl sequence and 18 different gRNA sequences targeting the AtPDS3 gene were cloned into plasmids and tested for activity in Arabidopsis protoplasts (FIG. 3A). Ten of the gRNA sequences tested have thus far shown a low level of editing activity, with highest editing efficiencies yielded by gRNAl , gRNAlO and gRNA18 (FIG. 3B), showing deletions of between 4 and 17 base pairs (FIG. 3C). Two configurations were tested, one in which the ISBagOl protein and the coRNA were expressed as separate expression cassettes, and one resembling the native bacterial architecture in which the native sequence encoding the ISBagO 1 protein and the oiRN A were partially overlapping in a single expression cassette (FIG. 3A). Editing activity was detected with both configurations (FIG. 3B). Importantly, the single expression cassette architecture should allow the sy stem to be further compacted, decreasing the cargo DNA size that will need to be cloned into plant viral vectors. While the editing efficiency of ISBagOl has thus far been very low (FIG. 3B), we are highly encouraged to see activity of this tiny nuclease in plant cells, as it is the first candidate we have tested. A very similar set of experiments was carried out with ISXfaOl, however we did not observe any editing activity. These initial experiments suggest that a wise course of action will be to test many TnpB candidates to find ones with the highest editing efficiency in plant cells. Table 1 includes the construct sequences for the constructs used in Example 3. Table 2 lists the spacer sequences of ISBagOl that were tested.Table 1: Construct Sequences for Example 3Table 2: spacer sequences of ISBagOl that were tested in Example 3
[0097] Example 4 - Identification and testing of tiny TnpB systems in plant cells
[0098] Utilizing ggKbase, new RNA-guided TnpB systems will be identified from IS605 (with IS605 TnpA), IS607 (with IS607 TnpA), and IS 1341 (without TnpA) transposons using the “DOTS’" approach described above. Thousands of new RNA-guided TnpB systems in the range of 300-400 amino acids in size will be selected, manually curated and subjected to subsequent sanity checks: (1) the catalytic domain is not broken (2) structure prediction algorithms (e.g., Alphafold2) can produce a reasonable 3D folding of the TnpB protein and (3) the RE of the transposons comprise putative coRNA-like secondary structures by RNA folding programs (e.g.. RNAfold). A large number of these TnpB (-100) will be tested, first in bacteria to identify their TAM sequences and confirm that they are active RNA-guided nucleases. Tests in bacteria will be run at 26 °C to screen for nucleases that are active at temperatures that are optimal for most crop plants. All active nucleases will then be tested in Arabidopsis and maize protoplasts to quickly test if these enzymes work in plant cells, and to screen for ones that work in both monocots and dicots. Enzymes whose catalytic activity is optimal around 23-28 °C will be prioritized as these are the temperatures that most crop plants are grown. This will be testedby running parallel protoplast editing assays at different temperatures. Protoplasts will be transfected with the small Cas systems encoded on plasmids, and editing outcomes will be assayed by deep amplicon sequencing to determine cutting efficiency, similar to the methods used above to test ISBagOl. Prioritized Cas systems will initially be tested to determine if (1) they are compatible with a single expression cassette format containing both the TnpB protein and oiRNA (FIG. 3A). (2) an N terminal or C terminal nuclear localization sequence shows activity, (3) if codon optimization is necessary. We will also vary the length of the ©RNA and spacer sequences and we will test a number of different gRNA sequences. It is believed that several TnpB systems with good editing activity will be identified, ideally ones with different TAM sequences which will allow for editing of a larger variety of plant sequences.
[0099] Example 5 - Testing tiny TnpB systems in planta
[0100] Small RNA-guided TnpB systems that show good activity in protoplasts will be tested in whole plant settings. We will test in a Nicotiana benthamiana transient expression system. We will test in transgenic Arabidopsis plants to test for germline editing and heritable transmission. The constructs which are used to test TnpB activities in protoplasts are binary vectors, thus, they can be used directly for Agrobacterium-mediated transformation of the Arabidopsis plants. In addition, the Arabidopsis protoplast assays will target the AtPDS3 gene, which provides a convenient readout for editing in whole plants. Constructs with TnpB systems and the best performing guide RNA sequences will be transformed into Arabidopsis plants and the TO plants used for transformation, as well as the T1 transgenic plants will be cultured at the optimum temperature determined for the specific TnpB system being tested. The somatic editing efficiency of target regions in transgenic T1 plants will be evaluated by amplicon sequencing. In addition, sy stems with high efficiency will be screened for white sectors caused by the loss of function of the AtPDSS gene, as seen previously for CasO (Li et al., 2023). T2 populations of T1 plants with relatively high somatic editing efficiency will be screened for albino seedlings and tested for transgene segregation. Fully albino seedlings in the T2 population in which the transgene has been segregated away (null segregants) indicate that the mutation in the AtPDS3 gene has been meiotically inherited.
[0101] Example 6 - Testing tiny TnpB systems in plant viral systems
[0102] The most successful tiny RNA-guided TnpB systems will be initially encoded into tobacco rattle virus (TRV) derived viral vectors together with the isoleucine tRNA sequence which was shown to be the most effective RNA mobility sequence in our previous study (Nagalakshmi et al., 2022). Arabidopsis and Nicotiana benthamiana plants will be infected in order to test for both somatic and germline editing, using protocols established in our previouswork (Ghoshal et al., 2020; Nagalakshmi et al.. 2022). For Arabidopsis, the best performing AtPDS3 guide RNA sequences will be used so that we can observe white sectors on the initially infected plants, and also screen for albino seedling in the next generation which will be indicative of germline transmission of edits (Li et al., 2023).
[0103] Those systems showing the most promise will also be encoded into viral vectors being developed that are capable of infecting monocots such as maize, and tested for both somatic and germline editing.
[0104] Example 7 - Optimizing the efficiency of tiny TnpB systems
[0105] The efficiency of the best systems will be further optimized by protein and guide RNA engineering using both directed evolution and rational design in E.coli and / or in vitro assays, following retesting in protoplasts and whole plants using the approaches in Examples 4-6. We will first employ a bacterial selection system to perform high-throughput directed evolution on the TnpB proteins of candidate systems for higher efficiency (Ma et al., 2022). In the selection system, TnpB-mediated cleavage of a toxic gene in a selection plasmid enables bacterial survival (FIG. 4A and 4B). Candidate TnpBs will be selectively mutagenized in all regions of the protein using standard protocols (Hand et al.. 2019). Mutants with fast DNA targeting kinetics will be selected under conditions such as shortened bacterial recovery time after plasmid transformation, where faster targeting kinetics can compensate for the shortened time to destroy the toxic gene.
[0106] As a second approach we will determine the 3D structures of the candidate TnpB systems and use structural guided mutagenesis of both the TnpB protein and coRNA sequences, and then test variants for improved activity in in vitro assays. We will first express and reconstitute the candidate TnpB and coRNA complexes in E.coli and conduct in-vitro biochemical analysis on these candidate systems. Using polyacrylamide gel electrophoresis (PAGE), the DNA cleavage activity of these systems on short oligonucleotides as well as plasmid DNA will be assessed. We will then use cryo-electron microscopy to perform structural studies on systems that are biochemically well-behaved with detectable DNA binding and / or cleavage activities. We will identify potential engineering sites based on the 3D structures and prior CRISPR-Cas enzyme engineering efforts (Kim et al.. 2022; Tsuchida et al., 2022; Xu et al., 2021) and design variants that may improve genome editing activities.
[0107] Variants from these approaches will then be tested in protoplast editing assays, after which they will be incorporated into the best performing viral editing systems.
[0108] Example 8 - Testing vecotrs for TnpB-mediated editing in plants
[0109] This Example describes exemplary' experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated coRNA. These vectors will be used to modify plant genomic DNA in protoplast cells, somatic whole plant cells, and germ cells.
[0110] Identification of TnpB target sites
[0111] Four TnpBs capable of targeted mutagenesis in bacteria and human cells were selected for testing in plants: ISDra2, ISYmul, ISAaml, and ISDgelO (Xiang et al. 2023; Karvelis et al. 2021). We will test these TnpBs using their native DNA sequence, maize codon optimized sequence, and Arabidopsis codon optimized DNA sequence (see sequences provided below in informal sequence listing). For each TnpB, a library' of spacer sequences targeting the Arabidopsis PDS3 gene (AT4G14210), will be designed according to their Transposon associated motif (TAM) sequence (Table 3). In Table 3, the gRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U). A similar approach will be used to test for editing capabilities in other species, such as maize.Table 3: TAM and gRNA sequences
[0112] Creating vectors for TnpB-mediated genome editing
[0113] To test whether the TnpBs as described herein may be used to target DNA modifications in the plant genome, a variety of constructs will be created. Based on reports published in human cells, we will test a split system for expression of the TnpB and coRNA sequences (Xiang et al. 2023; Karvelis et al. 2021). Here the TnpB and coRNA will be expressed by a separate promoter and terminator. The TnpB sequence will be cloned into a binary vector with NLS sequence(s) at the 5?and / or 3’ end of the TnpB sequence. The TnpB will be expressed using a Pol II promoter and terminator such as UBQ10 and rbcS E9, respectively (Table 20 below). As we expect the presence of an intron will improve TnpB protein expression, the IV2 intron (Table 20), of various copy numbers, will be cloned into the TnpB coding sequence to test for enhanced editing. The coRNA sequence will be cloned into a separate intermediate vector driven by either a Pol III promoter (such as AtU6) or a Pol II promoter (such as UBQ10) (Table 20). The spacer sequences will then be cloned into theintermediate coRNA vector using a ty pel I restriction enzymes such as PaqCI. RNA processing sequences, such as Hammerhead ribozyme. HDV ribozyme, or tRNA-Gly will be tested for improved editing (Table 20). Once the binary vector containing the TnpB sequence and intermediate vector for each of the coRNA-spacer is created, the coRNA expression cassette from the intermediate vector will get assembled into the TnpB-containing binary' vector (FIG. 6). Of the TnpBs capable of editing, further optimization will be done, such as testing various coRNA sequence lengths and spacer sequence lengths.
[0114] Testing TnpBs for editing in protoplast cells
[0115] The constructs described above will be tested in Arabidopsis mesophyll protoplast cells and maize protoplast cells using previously reported methods (Yoo, Cho, and Sheen 2007; Niu et al. 2019). To test for sensitivity to temperature, protoplast cells will be subjected to varying temperature regimens, such as 26 °C, 32 °C, 37 °C, and various combinations of temperatures including heat shock treatments. The heat shock treatments will be performed at vary ing times-after-transfection, and vary ing lengths and temperatures, such as a 2-hour incubation at 37 °C twelve hours after transfection.
[0116] Testing TnpBs for somatic and germline editing in whole plants
[0117] The ability of TnpBs to edit whole plants and faithfully transmit those edits to the next generation is crucial for the creation of novel stable genoty pes. To evaluate the TnpBs for somatic editing in stable transgenic plants, and transmission of germline edits to transgene-free progenies, the TnpBs along with promoter, terminator, and RNA processing combinations capable of generating targeted DNA modifications in protoplasts will be transformed into Arabidopsis using the floral dip technique (Zhang et al. 2006).
[0118] Following floral dip, the T1 seeds will be harvested, subjected to selection to identify transgenic plants, and quantified for editing with somatic tissue such as leaves of T1 plants using deep amplicon next generation sequencing (NGS). In addition to the ubiquitous promoter in Example 8, we suspect highly active germ-cell promoters may improve germline editing and heritability of mutations. We will test additional pol II promoters such as RPS5A, YAO1 and EC1.2 promoters. Heritability of edits which disrupts the PDS3 gene function can be easily identified due to the bleached phenotype of biallelic mutations in the PDS3 (AT4G14210) gene being targeted for editing in the T2 generation. Albino plants in T2 generation in which the TnpB transgene has already segregated away indicates the heritability of the edits in the PDS3 gene. These plants will also be subjected to Sanger sequencing to confirm the edits in the PDS3 gene.
[0119] Example 9 - Single expression cassettes for improved editing efficiency
[0120] This Example describes exemplary' experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated coRNA in a single expression cassette. These vectors may be used to modify plant genomic DNA in protoplasts, whole plants and to generate heritable edits in genomic sequences using similar methods as described in Example 8. We expect these results to translate to eukaryotic species other than those being tested here.
[0121] In all previous reports of TnpB-mediated genome editing the TnpB has been separated from the coRNA. resulting in separate expression cassettes, similarto what is outlined in Example 8. However, in nature, the coRNA overlaps with the TnpB sequence (FIG. 7A).
[0122] We expect the single expression cassette will work as efficiently or better than the split systems previously described. To test this, we will clone the TnpB and overlapping cuRN A into a single expression cassette (see informal sequence listing below - SEQ ID NOs: 140-147) on a binary vector. As an example, the single expression cassette will contain a promoter (such as UBQ10), TnpB coding sequence, coRNA sequence, spacer sequence, RNA processing sequence, and terminator, as depicted in (FIG. 7B). Further optimization of TnpB and coRNA expression will be performed using various promoters, terminators, and RNA processing sequences such as tRNA-Gly, Hammerhead ribozyme and HDV ribozyme (Table 20).
[0123] Example 10 - TnpBs for multiplexed editing
[0124] This Example describes exemplary experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated coRNA with a tandem repeat of the coRNA in a single expression cassette. These vectors may be used to modify genomic DNA in protoplasts, whole plants and to generate heritable edits in genomic sequences using similar methods as described in Example 8.
[0125] The ability to make multiple edits in a single generation can greatly enhance genome editing possibilities.
[0126] Because TnpBs have been shown to self-process their own coRNA (Nety et al. 2023), we will test tandem repeats of the coRNA for improved coRNA processing, editing, and multiplexing capabilities; an example with one tandem coRNA repeat is depicted in FIG. 8, however the copy number may vary.
[0127] Example 11 - Constructing and testing viral vectors for TnpB-mediated somatic and germ cell editing
[0128] This Example describes exemplary' experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated coRNA in viral vectors such as Tobacco Rattle Virus (TRY).
[0129] Delivery' of viral vectors for genome editing has the potential to expedite the creation of novel genetic diversity, and at the same time avoid the process of generating transgenic plants. However, due to the sequence size limitation, viruses harboring the RNA- guided endonuclease and the gRNA (such as Cas9 and Casl2a) cannot fit in single viral vectors such as TRV.
[0130] We will test and optimize the TnpBs identified in the prior examples as viral vector reagents. We expect these vectors to create somatic, and germline mutations that are transmitted to the following generation. This work will primarily be conducted in Arabidopsis, however due to TRV’s wide host range this technology will be applicable to many other plant species. In addition, we expect the results with TRV to be applicable to a wide range of viral vectors, further broadening the range of plant species to which this technology can apply.
[0131] TRV is a bipartite positive-sense single stranded RNA virus consisting of two components: TRV1 and TRV2 (E. E. Ellison, Chamness, and Voytas 2021). TRV1 contains the RNA-dependent RNA polymerase (RdRP) which replicates both TRV1 and TRV2 and expresses sub genomic RNAs of other coding regions. TRV1 also encodes the movement protein (MP), responsible for enabling cell-to-cell movement, and the viral suppressor of RNA silencing (VSR), responsible for suppression of the host immune response. TRV2 contains the coat protein (CP) and can be modified for expression of heterologous sequences. The native TRV2 contains coding sequences to enable transmission between plants, which are removed in the vector. Heterologous sequences are expressed from a sub genomic promoter.
[0132] Cloning of the viral vectors
[0133] We will take advantage of what we learned in the previous Examples to reengineer TRV2 to express the TnpB and the OJRNA. In one instance we will drive the expression of the single expression cassette, containing the TnpB and oRNA-spacer. using the sub genomic pPEBV promoter (FIG. 9A). To test the split system, where the TnpB and coRNA are dnven by separate promoters, one example will be to express the TnpB protein coding sequence inframe with the CP sequence, separated by a (P or T)2A sequence, with the coRNA being driven by the sub genomic pPEBV promoter (FIG. 9B) (Tian et al. 2014). We will also test the strategy of delivering multiple TRV2 vectors for editing capabilities. For example, placing pPEBV:TnpB on one TRV2 and the pPEBV:coRNA on a second TRV2 vector, and simultaneously delivering them using the methods described above. These construct designs will be tested using the most efficient TnpBs identified in the previous Examples.
[0134] Movement sequences such as tRNA-Ile (SEQ ID NO: 148) will be used to improve the movement of the genome editing reagents (Liu et al. 2022; Nagalakshmi et al. 2022; Evan E. Ellison et al. 2020).
[0135] We expect 3' processing of the coRNA-spacer sequence to greatly influence the editing capabilities of these TnpB systems in viral vectors. As such, we will test a variety of RNA processing sequences downstream of the spacer, such as HDV ribozyme, tRNA-Ile, and the tandem repeat coRNA (FIG. 9C and 9D).
[0136] Evaluating viral vectors for somatic and germ cell editing
[0137] TRV1 and TRV2 vectors will be delivered to Arabi dopsis plants using previously described approaches such as syringe infiltration of leaves, agro-pri eking, and agro-flooding methods (Nagalakshmi et al. 2022).
[0138] Somatic editing will be quantified using methods similar to those described in Example 8. Briefly, two approaches will be used: deep amplicon NGS, and visually via the photobleaching phenotype caused by a bi-allelic frameshift mutation in the AtPDS3 gene (AT4G14210).
[0139] Heritable germline editing will also be assayed in the offspring generations using methods similar to those described in Example 8.
[0140] Example 12 - Utilizing TnpBs for genome editing in plant cells
[0141] This Example describes experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated coRNA. These vectors were used to modify plant genomic DNA m ' Arabidopsis mesophyll protoplast cells. This example is related to the prophetic Example 8, and the experiments were performed as described in that example.
[0142] Creating vectors for TnpB-mediated genome editing
[0143] Three TnpBs capable of targeted mutagenesis in bacteria and human cells were selected for testing in plants: ISDra2, ISYmul, and ISAaml (Xiang et al. 2023; Karvelis et al. 2021). In all previous reports of TnpB-mediated genome editing, the TnpB protein coding sequences have been separated from the DNA sequence encoding the coRNA, resulting in separate expression cassettes. However, in the natural host genomes of the TnpBs, the coRNA overlaps with or is directly downstream of the TnpB sequence, as revealed by metagenomic data in Example 1 and publicly available genomic data (FIG. 7A). Here it is demonstrated that targeted genome editing in plant cells is achieved by TnpB systems encoded in the single expression cassette format. The native DNA sequence encoding the TnpB and the coRNA was cloned into a single expression cassette, driven by the UBQ10 promoter, followed by the desired spacer sequence, HDV ribozyme RNA processing sequence, and rbcS-E9 terminatoron a binary' vector, as depicted in FIG. 7B (Tables 4 and 5 below). For each TnpB, gRNA sequences targeting the Arabidopsis PDS3 gene (AT4G14210) were designed according to their Transposon associated motif (TAM) sequence (Table 6).
[0144] Testing TnpBs for editing in protoplast cells
[0145] The constructs described above were tested in Arabidopsis mesophyll protoplast cells using the previously reported method (Yoo, Cho, and Sheen 2007). Protoplast cells were subjected to a 37°C heat shock treatment for 2 hours at 16 hours post transfection. At 48 hours post transfection, protoplasts were harvested for genomic DNA extraction and analysis for targeted mutagenesis using Next Generation Sequencing (NGS). Editing frequencies of up to 13.34% (experiment_3_g2) for ISYmul, 11.76% (experiment_2_g9) for ISDra2, and 0.31% (gl 7) for ISAaml. (FIG. 10, Table 7) were observed.
[0146] Example 13 - Identifying unpublished TnpBs for genome editing in plant cells
[0147] This Example describes experimental guidelines for testing vectors expressing theTnpB protein and the associated oRNA. These vectors were used to modify plant genomic DNA in protoplast cells.
[0148] Identification of novel TnpBs capable of targeted editing in plant cells
[0149] 21 novel TnpBs were identified using the pipeline described in Example 1. Using the same approach described above, the 21 TnpBs targeting the PDS3 gene (AT4G14210) were cloned into individual binary' vectors (Table 8 (SEQ ID NOs: 150-170) and Table 9 (SEQ ID NOs: 171-358). Each vector was transfected into Arabidopsis mesophyll protoplast cells. Protoplast cells were subjected to a 37°C heat shock treatment for 2 hours at 16 hours post transfection. At 48 hours post transfection, protoplasts were harvested for genomic DNA extraction and analysis for targeted mutagenesis using NGS. Of the 21 TnpBs tested, 14 TnpBs displayed little to no editing activity, with an average single target site editing frequency less than 0.1% (Table 7); Seven TnpBs exhibited higher activity, capable of editing an average greater than 0.1% for at least one target site (FIG. 11, Table 7). Notably, TnpB_20, TnpB_25, TnpB_15 and TnpB_21 reached up to an editing frequency of 0.53%, 0.57%, 0.83, and 1.36% at an individual target site, respectively. (FIG. 11, Table 7).
[0150] Example 14 - mRNA engineering for improved TnpB editing efficiency
[0151] This Example describes exemplary experimental guidelines for testing vectors expressing the TnpB protein and the associated CDRNA.
[0152] roRNA engineering for increased editing efficiency
[0153] It was recently demonstrated that RNA engineering of the ISDra2 OJRNA is a strategy' for improving the TnpBs ability to edit genomic DNA in mammalian cells (Li et al.2024). In the single cassette architecture, TnpB and ©RNA are co-expressed as a single transcript where sometimes there is an overlapping region between the C-terminus of TnpB coding sequences (CDS) and the 5 ’-end of the ©RNA. For example, the ISDra2 ©RNA scaffold sequence overlaps with the TnpB CDS except for 7 nucleotides. Such overlap between the CDS and ©RNA introduces additional complexity and difficulty7for RNA engineering, i.e. changing ©RNA sequences can also interfere with the CDS. To test for ©RNA engineering potential in plant cells. ISBagOl and ISYmul TnpB systems were selected, with 32-nt overlap and no overlap between TnpB CDS and ©RNA, respectively.
[0154] Small RNA-seq and RNA folding analysis on ISYmul and ISBagOl ©RNA suggests a three stem-loop (127-nt) and a five stem-loop (186-nt) ©RNA architecture, respectively (FIG. 12). The three stem-loop architectures of ISYmul ©RNA resembles the ISDra2 ©RNA (Sasnauskas et al. 2023; Nakagawa et al. 2023), in which the apical loop of stem-loop 1 forms pseudoknot interactions w ith the 3’-end of the ©RNA scaffold. Compared to ISDra2 and ISYmul, the ©RNA of ISBagOl bears two additional stem-loops at the 5’-end forming an RNA structure comprised of five stem-loops, where the pseudoknot interactions are still preserved at similar positions. Given the fact that the ©RNA architectures of ISYmul and ISBagOl are similar to ISDra2 and their evolutionary descendant enzyme CRISPR-Casl2f which has been structurally characterized and engineered (Sasnauskas et al. 2023; Nakagawa et al. 2023; Takeda et al. 2021; Hino et al. 2023; Kim et al. 2021), we plan to explore ©RNA engineering of ISYmul and ISBagOl for improved editing capabilities.
[0155] There are two primary strategies for RNA engineering for TnpB and CRISPR- Casl2f: stem truncation and base-pair substitution. For stem truncation, it has been shown that trimming or deleting RNA stems that are in structurally disordered regions and / or not interacting with proteins can improve the editing activities for ISDra2 TnpB and CRISPR- CasI2f (Nakagawa et al. 2023; Li et al. 2024; Kim et al. 2021). For base-pair substitution, it has been shown that replacing certain A-U or G-U base pairs with G-C base pairs at certain stems can stabilize the RNA folding and also improve the editing activities for CRISPR-Casl2f (Kong et al. 2023). Moreover, base-pair substitution has been applied to silent internal premature expression terminator signals (i.e. an internal penta-uridinylate region under U6 promoter) (Kim et al. 2021; Xu et al. 2021). In addition to the above two strategies, additional strategies are sought for ©RNA engineering including stabilizing pseudoknot interactions and RNA internal loop engineering (nucleotide insertions / deletions at the RNA internal loop). For ISYmul, five groups of ©RNA variants were designed: (1) stem-loop 3 truncation (2) stemloop 2 truncation (3) stem-loop 2 internal-loop insert on / deleti on (4) stem-loop 1 base-pairsubstitution and (5) PK base-pair substitution. For ISBagOl. five groups of ®RNA variants were designed: (1) stem-loop 5 truncation (2) stem-loop 4 truncation (3) stem-loop 4 internalloop insert on / deleti on (4) stem-loop 3 base-pair substitution and (5) stem-loop 2 truncation. All sequences for the engineered ©RNA variants are listed in Table 10 (SEQ ID NOs: 359- 430). To further increase editing capabilities using the ®RNA variants, stacking of oRNA mutations that improve editing will also be explored.
[0156] Of the ®RNA variants described above, two of the ISBagOl variants have thus far been tested in protoplasts. For ISBagOl, variants 1.1 and 1.2 were tested by trimming stemloop 5 (FIG. 13A and Table 11 (SEQ ID NOs: 431-436)). These TnpB and «RNA sequences were expressed as described in the above examples, with a single promoter and terminator used to express the overlapping TnpB-oRNA sequence and a 16bp spacer sequence (gl) targeting the PDS3 gene (AT4G14210). A significant increase in editing efficiency was observed for the ISBagOl TnpB with oiRNA variant 1.1 (3.31% editing) compared to the ISBagOl TnpB with the native wRNA sequence (2.17% editing), as depicted in FIG. 13B.
[0157] Example 15 - TnpB-mediated genome editing using TRY viral vectors inArabidopsis protoplast cells
[0158] As outlined in Example 11, viral vectors for TnpB-mediated genome editing of somatic and germ cells in plant species, such as Arabidopsis and major crop species are to be developed. It is expected the 5’ and 3’ processing of the ©RNA-spacer sequence greatly influence the editing capabilities of TnpBs in viral vectors. To test the impact of HDV ribozyme, tRNA-Ile, or the tandem repeat roRNAs on ®RNA processing and ultimately TnpB mediated genome editing, these sequences were encoded along with the ISYmul TnpB, ®RNA, and gRNA 2 as the TRV2 cargo as depicted in FIGS. 9A, 9C, and 9D, and Table 12 (SEQ ID NOs: 437-474).
[0159] Each TRV2 plasmid was co-transfected into Arabidopsis mesophyll protoplast cells together with the TRV 1 plasmid. Protoplast cells were subjected to a 37°C heat shock treatment for 2 hours at 16 hours post transfection. At 48 hours post transfection, protoplasts were harvested for genomic DNA extraction and analysis for targeted mutagenesis using NGS. Detectable editing was observed in all three configurations tested, with a similar trend in both experiments one and two. An average 0.15% and 1.09% editing using the ISYmul-g2- tRNA-Isoleucine configuration, and 0.84% and 1.61% editing using the ISYmul -g2-tandem- 200bp-wRNA-tRNA-Isoleucine configuration (FIG. 14) was observed. The highest editing, 4.93% and 11.61% using the ISYmul -g2 -HD V-tRNA-Isoleucine sequence, likely due to the HDV ribozyme’s highly efficient RNA processing, resulting in a precise 3’ end of the spacersequence (FIG. 14) was observed. Without being bound by any particular theory, it is possible that the use of the HDV ribozyme in TRV delivery to whole plants such as Arabidopsis seedlings may not result in as efficient editing as in the transfected protoplast cells due to the highly efficient RNA processing of the HDV ribozyme sequence, which may hinder the TRV’s ability to replicate and spread throughout the plant.
[0160] Therefore, to further optimize this system for the successful viral delivery in plants, strategies will be explored to precisely process the 3’ spacer and second coRNA junction, without preventing the virus’s ability to replicate and spread from cell-to-cell. The initial experiment using ISYmul-g2-tandem-200bp-wRNA-tRNA-Isoleucine contained a 200 nucleotide second ©RNA. The optimal length of the second coRNA in the ISYmul-g2-tandem coRN A-tRNA-Isoleucine configuration will be identified to create a precisely processed end at the 3’ end of the first spacer sequence. A screen will be performed using the same ISYmul - g2 -tandem ®RN A-tRNA-Isoleucine design, but with varying lengths for the second ©RNA, ranging from 100-250 nucleotides in increments of ten nucleotides (Table 12). Based on ISYmul small RNA-seq data in E.coli, it is suspected that 127 nucleotides may be an optimal length for the second coRN A. which will also be tested (FIG. 12A). Once the range with highest editing efficiencies is found, all ©RNA sequence lengths in that range to identity the optimal length of the second ©RNA sequence will be tested.
[0161] Example 16 - TnpB-mediated editing in transgenic plants
[0162] This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated ©RNA in transgenic plants by Agrobacterium transformation, as well as testing the corresponding editing outcomes in these transgenic plants. This example is related to Examples 5, 8, and 9.
[0163] TnpBs for gene editing in transgenic plants
[0164] Because Agrobacterium mediated transformation is still widely used to insert gene editing systems into plant genomes to generate desired edits, the ability' of TnpBs to edit in transgenic plants is an important approach for the creation of novel genetic materials. To evaluate TnpB-mediated editing in transgenic plants, the two TnpBs, ISDra2, and ISYmul, demonstrating the highest editing efficiency in Arabidopsis mesophyll protoplast cells were selected (FIG. 10). The native DNA sequence encoding the TnpB and the ®RNA were cloned into a single expression cassette (SEQ ID NOs: 140 and 143), driven by the UBQ10 promoter (SEQ ID NO: 132), followed by the desired spacer sequence (SEQ ID NOs: 55, 65. 121, 123), HDV ribozyme RNA processing sequence (SEQ ID NO: 137), and rbcS-E9 terminator (SEQID NO: 135) on a binary vector, as depicted in Figure 7B (Tables 4, 5, and 6) and used floral dipped according to Zhang et al. (Zhang et al. 2006).
[0165] Following floral dip, the T1 seeds were harvested and subjected to hygromycin selection onMS plates to select for transgenic plants. Seedlings that passed selection were then transplanted to soil and grown at room temperature for three weeks. Next, leaf tissue samples were randomly collected from each transgenic plant and editing efficiencies were quantified by amplicon sequencing using Illumina next generation sequencing (NGS). Editing in leaf tissue for ISDRa2 (g9 and gl 2) and ISYmul (g2 and gl 2) was observed. ISDra2 g9 demonstrated an average editing frequency in wild type (WT) of 0.32% and in rdr6 of 0.04% (FIG. 15A). ISDra2 g!2 demonstrated an average editing frequency of 0.28% and 0.37% in wild type and rdr6. respectively (FIG. 15B). For ISYmul g2 we observed an editing frequency in wild type of 0.14%, and in rdr6 of 4.59% (FIG. 15C). ISYmul g!2 demonstrated a very high editing frequency7in wildtype (48.96%) and in rdr6 (79.73%) (FIG. 15D). Furthermore, analysis of the repair profiles for ISDra2 g9 and ISYmul g!2 revealed deletion-dominant repair outcomes (FIGS. 15E-15G).
[0166] Example 17 - TnpB-mediated editing via delivery of a TRY vector to whole plants
[0167] This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated coRNA via the Tobacco Rattle Virus (TRV) delivery7. These vectors were used to edit plant DNA in whole plants. This example is related to the previously filed Examples 6, 11 and 15.
[0168] Here, TRV vectors for TnpB-mediated genome editing of Arctbidopsis were designed. Two different viral vector configurations were cloned into the TRV “cargo” site for testing: 1. ISYmul-g2-tRNA-Isoleucine (SEQ ID NO: 456) and 2. ISYmul-g2- HDV-tRNA-Isoleucine (SEQ ID NO: 457), as shown in FIGS. 9A and 9C, respectively (Table 12).
[0169] Each TRV2 plasmid was co-delivered to wild type or rdr6 Arctbidopsis seedlings together with the TRV1 plasmid at 9 days post germination using the agroflood method (Nagalakshmi et al. 2022). After four days of agroflood co-culture, seedlings were transplanted to soil and grown for 18 days. Next, leaf tissue samples were randomly sampled from each plant for genomic DNA extraction and analysis for targeted mutagenesis using Illumina sequencing. Editing was detected in both vector configurations tested. An average 0.03% and 0.12% editing efficiency using the ISYmul -g2-tRNA~Isoleucine configuration in the wild type and rdr6 samples, respectively (FIG. 16A) was observed. Notably, editing efficiency up to 1.26% in the rdr6 genotype (FIG. 16A) was observed. Additionally, an average of 0.29%and 0.28% using the ISYmul-g2-HDV-tRNA-Isoleucine configuration in wild type and rdr6, respectively (FIG. 16B) was observed. Notably, moderately high editing efficiency up to 5.54% and 8.93% in the wild type and rdr6 samples, respectively (FIG. 16B) was observed. Additionally, a deletion-dominant indel profile, similar to the transgenic T1 data (FIGS. 16C- 16D) was observed. The editing target of these experiments is the PDS3 gene. When both copies of the PDS3 gene are mutated (biallelic edits), the mutated plant cells are white rather than green. On several of the plants infected with either of these TRV vectors white sectors were observed, indicating biallelic editing of PDS3 in these areas of the plant (FIG. 16E). The efficiency of editing observed in these experiments by amplicon sequencing, together with the appearance of white sectors on the plants, suggests that editing of the PDS3 gene will likely occur in some of the germ cells resulting in germ line transmission of edits. This will be confirmed in the next generation by growing plants and looking for white seedlings that contain biallelic edits in the PDS3 gene.
[0170] Example 18 - Investigating the activities of novel TnpBs in bacteria
[0171] This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated coRNA in bacteria. These vectors were used to test the activities of novel TnpB proteins in bacteria.
[0172] Thirty' novel TnpBs were identified using either the DART approach outlined in Examples 1 and 13 or a protein homology -based search, including seven new examples of TnpBs (Table 13 below). A subset of these TnpBs was confirmed for their plasmid interference activities in bacteria. The native DNA sequence encoding the TnpB and the wRNA was cloned into a single expression cassette (bearing a CmR gene) driven by the TetR promoter (FIG. 17A), followed by the desired spacer sequence targeting another vector bearing an AmpR gene (FIG. 17A). Both vectors were co-transformed into E.coli, which were then subjected to serial dilution and grown on plates with either the single antibiotic (Cam) or double antibiotic (Cam and Amp). The experiments were performed at 26°C. The activities of TnpB were calculated by the ratio of colony -forming units (normalized CFU) betw een the single antibiotic plate and the double antibiotic plate. An active TnpB will target the AmpR vector resulting in plasmid interference and reduced normalized CFU. The normalized CFU for each TnpB were also compared under conditions with and without the predicted TAM sequences on the AmpR vector. We tested the plasmid interference activities of five selected TnpBs (TnpB_5, TnpB_15, TnpB_19, TnpB_21 and TnpB_30) and observed the highest activities in TnpB_30 in bacteria (FIG. 17B).
[0173] The TAM of TnpB_30 was further confirmed by a TAM depletion assay as the “AGGAG” motif (FIG. 17C). Small RNA-seq of TnpB_30 (FIG. 17D) revealed a usRNA scaffold size of 138 nucleotides, consisting of a three stem-loop architecture (Fig. 17E).
[0174] Example 19 - Identifying novel TnpBs for genome editing in plant cells
[0175] This Example describes exemplary' experimental guidelines for testing vectors expressing the TnpB protein and the associated coRNA. These vectors were used to modify plant genomic DNA in protoplast cells. It is expected that these results translate to eukaryotic species other than those being tested in this Example.
[0176] Seven novel TnpBs for genome editing capabilities in plant cells (Table 13) were identified and tested. Using the same approach outlined in previous Examples, the native DNA sequence encoding the TnpB and the toRNA was cloned into a single expression cassette, driven by the UBQ10 promoter (SEQ ID NO: 132), followed by the desired spacer sequence, HDV ribozyme RNA processing sequence (SEQ ID NO: 137), and rbcS-E9 terminator (SEQ ID NO: 135) on a binary vector, as depicted in (FIG. 7B, Tables 13 and 14). Each vector was transfected into Arabidopsis mesophyll protoplast cells. Protoplast cells were subjected to a 37°C heat shock treatment for 2 hours at 16 hours post transfection. At 48 hours post transfection, protoplasts were harvested for genomic DNA extraction and targeted mutagenesis analysis using amplicon sequencing. Of the seven TnpBs tested, six TnpBs displayed little to no editing activity, with all target sites averaging less than 0.1% editing efficiency (Table 15). Notably, TnpB 30, the most active novel TnpB in bacteria in Example 18, generated edits at an average 0.45% across nine target sites, reaching up to and average of 1.2% at target site g2 (FIG. 18, Table 15).
[0177] Example 20 - The spacer length for ISYmul, ISDra2, ISBagOl. and TnpB 30 can impact editing efficacy
[0178] This Example describes exemplary experimental guidelines for testing vectors expressing the TnpB protein and the associated wRNA. These vectors were used to modify plant genomic DNA in protoplast cells. It is expected that these results translate to eukaryotic species other than those being tested in this Example.
[0179] It has been shown in Cas9 and Casl2s that the spacer sequence length used for targeted genome editing can have a major impact on nuclease efficacy. Therefore, we tested different spacer lengths for the TnpBs ISYmul (SEQ ID NO: 140), ISDra2 (SEQ ID NO: 143), ISBagOl (SEQ ID NO: 14), and TnpB_30 (SEQ ID NO: 481). Using the same single expression cassette described in previous Examples, we cloned plasmids using a single expression cassette, driven by the UBQ10 promoter (SEQ ID NO: 132), followed by the desired spacersequence length (ranging from 14-20 nucleotides for ISYmul. ISDra2, and ISBagOl; for TnpB_30 we tested spacer lengths ranging from 13-21 nucleotides (Table 16)), HDV ribozyme RNA processing sequence (SEQ ID NO: 137), and rbcS-E9 terminator (SEQ ID NO: 135) as depicted in FIG. 7B.
[0180] Each vector was transfected into Arabidopsis mesophyll protoplast cells. Protoplast cells were subjected to a 37°C heat shock treatment for 2 hours at 16 hours post transfection. At 48 hours post transfection, protoplasts were harvested for genomic DNA extraction and targeted mutagenesis analysis using amplicon sequencing. Analysis of the editing frequencies revealed very little difference between the spacer lengths for ISDra2 (FIG. 19A). ISYmul displayed a preference for spacer lengths of 16 and 19 nucleotides, with 16 being the optimal length (FIG. 19B). ISBagOl appeared to have a strong preference for a 16 nucleotide spacer length (FIG. 19C). TnpB_30 showed increased editing for spacer lengths greater than 15 nucleotides, with 19 nucleotides demonstrating the highest editing efficiency (FIG. 19D). In summary, these data indicate the spacer length can influence TnpB-mediated genome editing efficacy.
[0181] Example 21 - Expression of TnpB-uiRNA using a single promoter for optimalTnpB-mediated genome editing
[0182] This Example describes exemplary experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated wRNA in a single expression cassette, or a split expression cassette design. These vectors may be used to modify plant genomic DNA in protoplasts, whole plants and to generate heritable edits in genomic sequences using similar methods as described in other Examples in this patent. It is expected that these results translate to eukaryotic species, and TnpBs, other than those being tested in this Example. This Example is related to Example 9.
[0183] In all previous reports of TnpB-mediated genome editing the TnpB has been separated from the wRNA, resulting in separate expression cassettes. However, in nature, the OJRNA is directly downstream of, or overlaps with, the TnpB sequence (FIG. 7 A). We hypothesized the single expression cassette will work as efficiently or more efficiently than the split systems described in the literature. To compare the single expression cassette with the split expression cassette we cloned the TnpB and uiRNA into a single expression cassette (SEQ ID NO: 140) on a binary' vector (FIG. 20A). The single expression cassette contained a pol II promoter (such as UBQ 10 (SEQ ID NO: 132)), TnpB coding sequence, wRNA sequence, spacer sequence, HDV ribozyme RNA processing sequence (SEQ ID NO: 137). and terminator (such as rbcS-E9 (SEQ ID NO: 135)) (FIG. 20A). As a comparison, we created split designswhere the TnpB coding sequence is driven by a Pol II promoter (such as UBQ10 (SEQ ID NO: 132)) and the toRNA sequence, spacer sequence, and HDV ribozyme (SEQ ID NO: 137) sequence are under the expression of a pol III U6 promoter (SEQ ID NO: 134) (FIG. 20A). In total 11 split design plasmids were designed, with the wRNA length varying from 100-250 nucleotides (FIG. 20A, Table 17). Based on small RNA-seq performed in E.coli (Example 14 and FIG. 20B), we hypothesized the wRNA is processed at 127 nucleotides. Therefore, we included a split design with 127 nucleotides wRNA (Table 17).
[0184] The single versus split expression configurations using ISYmul g2 (SEQ ID NOs: 140 and 55, Table 17) was tested, however, it is expected these results translate to other TnpBs and target sites. Each vector was transfected into Arabidopsis mesophyll protoplast cells. Protoplast cells were subjected to a 37°C heat shock treatment for 2 hours at 16 hours post transfection. At 48 hours post transfection, protoplasts were harvested for genomic DNA extraction and targeted mutagenesis analysis using amplicon sequencing.
[0185] Analysis of the amplicon sequencing data revealed the single expression cassette to be the top performer at 0.70% editing efficiency (FIG. 20C). By comparison, all of the split designs demonstrated much lower editing efficiency, ranging from 0.01-0.23% for the 150 and 127 nucleotides wRNA, respectively (FIG. 20C). As expected, the 127 nucleotides wRNA in the split design generated the highest editing efficiency out of all of the split designs tested, indicating the proper length and processing of the wRNA is critical for TnpB editing efficacy. In summary, we demonstrated that the single expression cassette is an optimal configuration to express the ISYmul and wRNA sequences for TnpB-mediated genome editing. We expect these results to translate to other TnpBs and eukaryotic species.
[0186] Example 22 - ISYmul wRNA engineering for improved plant genome editing
[0187] This Example describes exemplary experimental guidelines for testing vectors expressing the TnpB protein and the associated wRNA. This Example is related to Example 14. It is expected that these results translate to TnpBs and eukaryotic species other than those being tested in this Example.
[0188] It was recently demonstrated that RNA engineering of the ISDra2 wRNA is a strategy for improving the TnpB’s ability to edit genomic DNA in mammalian cells (Li et al. 2024). In the single cassette architecture, the TnpB and wRNA are co-expressed as a single transcript where sometimes there is an overlapping region between the C-terminus of TnpB coding sequences (CDS) and the 5’-end of the wRNA. For example, the ISDra2 wRNA scaffold sequence overlaps with the TnpB CDS except for 7 nucleotides. Such overlap betweenthe CDS and UJRNA introduces additional complexity and difficulty for RNA engineering, i.e. changing usRNA sequences can also interfere with the CDS.
[0189] To test for wRNA engineering potential in plant cells, 32 unique ISYmul wRNA variants were cloned, as outlined in Example 14, into individual binary vectors using the single expression cassette design (Table 10). Each vector was transfected into Arabidopsis mesophyll protoplast cells. Protoplast cells were subjected to a 37°C heat shock treatment for 2 hours at 16 hours post-transfection. At 48 hours post-transfection, protoplasts were harvested for genomic DNA extraction and targeted mutagenesis analysis using amplicon sequencing. We observed three trends in the wRNA variants relative to the wild type (WT) wRNA: (1) increased editing frequency (wRNA variants v2.1, v3.2, v3.3, v3.4, v3.7, v3.10, v3.11, v3.16, vl. l, and v5.2), (2) decreased editing frequency (ci RNA variants v2.2, v3.1, v3.6, v3.9, v3.12, v3.13, v3.14, v3.15, vl.2, vl.3, vl.4, vl.5, vl.6, v2.4, v4.1, v4.2, v5.1, v5.3, and v5.4), or (3) similar editing frequency (v2.3, v3.5, v3.7, and v3.8) (FIG. 21). Notably we observed 2.1, 1.63, 1.38-fold increased editing for wRNA variants v3.2, v3.16, and v2. 1, respectively (FIG. 21).
[0190] RNA secondary structure analysis revealed that v3.2 features an SISI rather than an S2S1 two-way junction topology at lower stem-loop 2; v3.16 features an S2S0 rather than an SI SO two-way junction topology at the middle of stem-loop 2; and v2. 1 features stem truncation at upper stem-loop 2 (FIG. 22). Since modulation of RNA two-way junctions in stem-loop 2, as seen with v3.2, enhances genome editing activities the most, we hypothesize that maintaining similar topologies in stem-loop 2 are critical for ISYmul TnpB activity. Consequently, 9 additional variants based on the ISYmul oiRNA variant v3.2 with a new mutation (Table 18) were designed. Furthermore, as indicated in Example 14, we plan to combine coRNA variant mutations to test for improved editing. Stacked coRNA variants using combinations of v3.2, v3.16, v3.4, v5.2, and / or v2.1 to evaluate enhanced ISYmul editing capabilities (Table 18) have been designed. It is anticipated that further exploration of coRNA engineering variants and stacking will significantly improve ISYmul genome editing efficacy.
[0191] Example 23 - Highly efficient ISYmul-mediated somatic editing in transgenic Arabidopsis
[0192] This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated oiRNA in transgenic plants by Agrobacterium transformation, as well as testing the corresponding editing outcomes in these transgenic plants. This Example is related to Example 16.
[0193] Because Agrobacterium mediated transformation is still widely used to insert gene editing systems into plant genomes to generate targeted edits, the ability’ of TnpBs to edit intransgenic plants is an important approach for the creation of novel genetic materials. To evaluate TnpB-mediated editing in transgenic plants, ISYmul g2 and ISYmul gl2 were selected. The native DNA sequence encoding the TnpB and the rnRNA were cloned into a single expression cassette (SEQ ID NO: 140), driven by the UBQ 10 promoter (SEQ ID NO: 132), followed by the desired spacer sequence (SEQ ID NOs: 55 and 65), HDV ribozy me RNA processing sequence (SEQ ID NO: 137), and rbcS-E9 terminator (SEQ ID NO: 135) on a binary vector, as depicted in FIG. 7B, and used for floral dipped according to Zhang et al. 2006. Floral dip was performed using wild type (WT) and RNA-DEPENDENT RNA POLYMERASE 6 (rdr6) genotypes, as rdr6 has been shown to have reduced transgene silencing.
[0194] Following floral dip. the T1 seeds were harvested and subjected to hygromycin selection on !6 MS plates to select for transgenic plants. Seedlings that passed selection were then transplanted to soil and grown at room temperature for one week. After one week, plants continued to grow at room temperature for two more weeks; however, plants that underwent heat shock treatment were exposed to 8 hours of heat exposure at 37°C every7day for 5 days, followed by 2 days of recovery at room temperature. This heat shock regime lasted for two weeks. Next, leaf tissue samples were randomly collected from each transgenic plant and editing efficiencies were quantified by amplicon sequencing using Illumina next generation sequencing (NGS). Editing in leaf tissue for ISYmul g2 and ISYmul g!2 was detected. For ISYmul g2 we observed an average editing frequency in WT of 1.56%. and 2.49% in rdr6 for the room temperature samples (FIG. 23 A). ISYmul gl2 demonstrated a high editing frequency in WT (44.86%) and in rdr6 (75.55%) for the room temperature grown plants (FIG. 23B). Analysis of editing efficiency in the heat shock treatment samples revealed increased editing at ISYmul g2 target site in both WT and rdr6, averaging 9.81% and 32.47%. respectively (FIG. 23 A). The ISYmul g!2 plants that underwent the heat shock treatment showed increased editing in WT, averaging 63.73%; whereas the rdr6 samples saw no difference in editing efficiency, averaging 75.43% (FIG. 23B).
[0195] Example 24 - Efficient ISDra2-mediated somatic editing in transgenic Arabidopsis requires heat shock treatment.
[0196] This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated wRNA in transgenic plants by Agrobacterium transformation, as well as testing the corresponding editing outcomes in these transgenic plants. This Example is related to Examples 16.
[0197] Because Agrobacterium mediated transformation is still widely used to insert gene editing systems into plant genomes to generate targeted edits, the ability of TnpBs to edit intransgenic plants is an important approach for the creation of novel genetic materials. To evaluate TnpB-mediated editing in transgenic plants, ISDra2 g9 and ISDra2 gl2 were selected. The native DNA sequence encoding the TnpB and the wRNA were cloned into a single expression cassette (SEQ ID NO: 143), driven by the UBQ IO promoter (SEQ ID NO: 132), followed by the desired spacer sequence (SEQ ID NOs: 121 and 123), EIDV ribozy me RNA processing sequence (SEQ ID NO: 137), and rbcS-E9 terminator (SEQ ID NO: 135) on a binary vector, as depicted in FIG. 7B 7B, and used for floral dipped according to Zhang et al. 2006. Floral dip was performed using wild type (WT) and RNA-DEPENDENT RNA POLYMERASE 6 (rdrti) genotypes, as rdr6 has been shown to have reduced transgene silencing.
[0198] Following floral dip. the T1 seeds were harvested and subjected to hygromycin selection on Vz MS plates to select for transgenic plants. Seedlings that passed selection were then transplanted to soil and grown at room temperature for one week. After one week, room temperature plants continued to grow at room temperature for two more weeks; however, plants that underwent heat shock treatment were exposed to 8 hours of heat exposure at 37°C every day for 5 days, followed by 2 days of recovery at room temperature. This heat shock regime lasted for two weeks. Next, leaf tissue samples were randomly collected from each transgenic plant and editing efficiencies were quantified by amplicon sequencing using Illumina next generation sequencing (NGS). Editing in leaf tissue for ISDra2 g9 and ISDra2 gl2 was detected. For ISDra2 g9 we observed an average editing frequency in WT of 0. 16%, and in rdr6 of 0.10% for the room temperature samples (FIG. 24A). For ISDra2 g!2 we observed an average editing frequency in WT of 7.62%, and in rdr6 of 1. 14% for the room temperature samples (FIG. 24B). Comparison of editing efficiency between room temperature and heat shock treatment revealed a substantial difference for both ISDra2 g9 and ISDra2 g!2 (FIG. 24A-24B). For ISDra2 g9 we observed an average editing frequency in WT of 12.40%, and 41.20% in rdr6 for the heat shock treatment samples (FIG. 24A). For ISDra2 g!2 we observed an average editing frequency in WT of 18.92%, and 58.40% in rdr6 for the heat shock treatment plants (FIG. 24B). These data indicate an elevated temperature treatment is critical for efficient plant gene editing using ISDra2.
[0199] Example 25 - Somatic editing of whole plants using TRV vectors expressing theISYmul TnpB and tuRNA
[0200] This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated wRNA via Tobacco Rattle Virus (TRV)delivery'. These vectors were used to edit plant DNA in whole plants. This Example is related to the previously filed Example 17.
[0201] First, TRV vectors for TnpB-mediated genome editing in Arabidopsis were designed targeting ISYmul g2 target site. Three different viral vector configurations were cloned into the TRV "cargo" site for testing: (1) ISYmul -g2-tRNA-Isoleucine (SEQ ID NO: 437), (2) ISYmul-g2-HDV-tRNA-Isoleucine (SEQ ID NO: 438), and (3) ISYmul -g2-tandem- 200bp-wRNA-tRNA-Isoleucine (SEQ ID NO: 439) (FIG. 25A).
[0202] Each TRV2 plasmid was co-delivered with the TRV1 plasmid to wild type (WT), ku70, or RNA-DEPENDENT RNA POLYMERASE 6 (rdr 6) Arabidopsis seedlings at 9 days post germination using the agroflood method (Nagalakshmi et al. 2022). Ku70 is a nonhomologous end joining DNA repair mutant. We hypothesized the ku 70 genetic background may show increased editing efficiency.
[0203] After four days of agroflood co-culture, seedlings were transplanted to soil and grown for 18-22 days. For the heat shock treatment, one week after seedlings were transplanted from agroflood plates to soil, they were exposed to 8 hours of heat exposure at 37°C every day for 5 days, followed by 2 days of recovery at room temperature. This heat shock regime lasted for two w eeks (or until the plants started flowering). Next, three leaf tissue samples distal from the TRV deliver}' location were randomly collected and pooled from each individual plant for genomic DNA extraction and analysis for targeted mutagenesis using Illumina amplicon sequencing. Editing was detected in all three vector configurations tested.
[0204] An average 0.06%, 0.09%, and 0.02% editing efficiency using the ISYmul-g2- tRNA-Isoleucine configuration in the WT, rdr6. and ku70 samples, respectively, was observed for the room temperature treatment samples (FIG. 25B). An average 0.36%, 0.25%, and 0.68% editing efficiency using the ISYmul -g2-tRNA-Isoleucine configuration in the WT, rdr6, and ku70 samples, respectively, was observed for the heat shock treatment samples (FIG. 25B). Additionally, an average of 0.56%, 0.81%, and 1.96% editing efficiency using the ISYmul- g2-HDV-tRNA-Isoleucine configuration in WT, rdr6. and ku70, respectively, was observed for the room temperature samples (FIG. 25B). An average of 3.29%. 0.16%. and 8.93% editing efficiency using the ISYmul-g2-HDV-tRNA-Isoleucine configuration in WT, rdr6. and ku70. respectively, was observed for the heat shock treatment samples (FIG. 25B). Additionally, we detected editing using the ISYmul -g2-tandem-200bp-wRNA-tRNA-Isoleucine TRV vector design. An average of 0.01% using the ISYmul -g2-tandem-200bp-wRNA-tRNA-Isoleucine configuration in all three genotypes was observed for the room temperature samples (FIG. 25B). An average of 0.14%, 5.26%, and 0.17% using the ISYmul-g2-tandem-200bp-wRNA-tRNA-Isoleucine configuration in wild type, rdr6, and ku70, respectively, was observed for the heat shock treatment samples (Figure 26B).
[0205] Next we designed TRV vectors to target the ISYmul gl2 target site. Here, the ISYmul-gl2-HDV -tRNA-Isoleucine TRV vector was used (Table 19). The TRV2 plasmid was co-delivered with TRV1 to WT or rdr6 Arabidopsis seedlings using the same agroflood methods described above (Nagalakshmi et al. 2022). Additionally, the plant growth and heat shock treatment was the same as described above. Next, three leaf tissue samples distal from the TRV delivery location were randomly collected and pooled from each individual plant for genomic DNA extraction and analysis for targeted mutagenesis using Illumina amplicon sequencing. Editing was detected in both genotypes tested. An average 8.51% and 13.54% editing efficiency using the ISYmul -g!2-HDV -tRNA-Isoleucine configuration in the WT and rdr6 samples, respectively, was observed for the room temperature grown plants (FIG. 25C). Additionally, an average of 4.27% and 1.34% using the ISYmul-gl2-HDV-tRNA-Isoleucine configuration in WT and rdr6, respectively, was observ ed for the plants that underwent heat shock treatment (FIG. 25C). Further, 6 out of 57 WT plants displayed editing greater than 40%. with four of those being greater than 75%, for the lSYmuI-gl2-HDV-tRNA-lsoleucine vector with room temperature treatment (FIG. 25C). A deletion dominant editing pattern for ISYmul gl2 was observed, similar to the mutation profile generated using TRV with ISYmul g2 (FIG. 25D).
[0206] These results indicate that both ISYmul g2 and gl2 are capable of somatic editing when delivered to whole plants using the TRV vector. Further, these data demonstrate that all three TRV vector designs are capable of targeted genome editing in plants. Additionally, we demonstrate the use of genetic mutants for improved TRV-mediated genome editing efficiency using ISYmul. It is expected that these results translate to TnpBs, viruses, and eukaryotic species other than those tested in this Example.
[0207] Example 26 - Heritable editing using TRV vectors expressing the ISYmul TnpB and uiRNA in Arabidopsis
[0208] This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated wRNA via Tobacco Rattle Virus (TRV) delivery. These vectors were used to edit plant DNA in whole plants. This Example is related to the previously filed Example 17.
[0209] White sectors were observed on several of the plants infected with the TRV vectors in Example 17 and Example 25, indicating biallelic editing of the PHYTOENE DESATURASES gene (PDS3) in these areas of the plant (FIG. 16E, FIG. 26A). To test for heritability of editedalleles to the next generation, seeds were collected from a WT plant showing 54.54% somatic editing using the ISYmul-g2-HDV-tRNA-Isoleucine TRV design with the heat shock treatment (FIG. 26A). In total, 2318 seeds were sown on !6 MS plates containing 3% sucrose. After 8-10 days 68 albino seedlings were observed, indicating biallelic mutations in the PDS3 gene (FIG. 26B). To confirm that the PDS3 gene is mutated we performed Sanger sequencing on the two white seedlings shown in FIG. 26B. Sanger sequencing revealed that both plants are homozygous for a 4bp deletion (FIG. 26C). To further characterize transmission of edited alleles, NGS was performed on 58 seedlings, of which 31 were albino. Analysis of NGS revealed biallelic mutations for all of the albino seedlings, with the majority of mutations being the 4bp deletion observed in FIG. 26C. Of the 27 green seedlings, two of them were heterozygous (WT / 4bp del). These data indicate that TRV -mediated edits using ISYmul are heritable. We anticipate higher heritability rates for the highly edited plants targeting ISYmul gl2 from Example 25. Further, w e expect the approaches described in this Example will extend to other TnpBs, viral vectors, and plant species.
[0210] Example 27 - Efficient plasmid interference activities of TnpBs with matured O)RNA
[0211] This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated wRNA in bacteria. These vectors were used to test the activities of TnpB proteins with various mechanisms of wRNA expression and processing.
[0212] Different expression cassettes for TnpB and wRNA in bacteria were created as depicted in (FIG. 27 A): (a) A single expression cassette for TnpB and wRN A with the 3’-end of the guide uncapped; (b) A single expression cassette for TnpB and wRNA with a 16-nt guide region with the 3’-end capped by HDV ribozyme; (c) A split expression cassette for TnpB and UJRNA, with the wRNA scaffold length inferred from small RNA-seq experiments, a 16- nucleotide guide region with the 3’ -end capped by HDV ribozyme; (d) A split expression cassette for TnpB and uiRNA, with the wRNA scaffold truncated by 52 nucleotides at the 5’- end, a 16-nucleotide guide region with the 3'-end capped by HDV ribozyme; (e) A split expression cassette for TnpB and wRN A, with the wRN A scaffold extended by 73 nucleotides at the 5’-end, a 16-nucleotide guide region with the 3’-end capped by HDV ribozyme; (f) A single expression cassette for TnpB and wRN A, with the 3’-end of the guide capped by another u’RNA scaffold of a size inferred from small RNA-seq experiments; and (g) A single expression cassette for TnpB and wRNA, with the 3 ’-end of the guide capped by another OJRNA scaffold extended by 73 nucleotides at the 5’-end.
[0213] First, the importance of 3 ’-end processing of coRNA for TnpB activity was assessed. TnpB was expressed from ISDra2, ISBagOl, ISYmul, and ISAaml using cassette (a). Strong plasmid interference activities were observed for ISDra2 and ISBagOl but not for ISYmul and ISAaml (FIG. 27B). When the HDV ribozyme was introduced at the 3’-end of the guide in cassette (b), the activities of ISYmul and ISAaml were restored (FIG. 27C), indicating that the lack of 3 '-end processing was responsible for their weak activities in cassette (a).
[0214] Next, the importance of 5’-end processing of coRNA for TnpB’s activities was examined. ISBagOl and ISYmul were expressed using the split cassette (c), where the 3’-end of the coRNA was capped by the HDV ribozyme, resulting in efficient plasmid interference activities for both TnpBs (FIG. 27D and 27E). When the 5’-end of the coRNA for ISBagOl was truncated by 52 nucleotides, plasmid interference activity diminished (FIG. 27D). Extending the 5’-end of the coRNA for ISYmul by 73 nucleotides did not affect its activity (FIG. 27E). These findings suggest that the proper size of the coRNA scaffold is crucial for TnpB activity and that 5 ’-end extension does not impact TnpB function, likely because the 5 ’-end can be processed into mature coRNA by TnpB or by other host nucleases.
[0215] Utilizing these insights into coRNA processing, a single cassette (f) was designed where the 16-nucleotide guide region was capped by an intact 127-nucleotide coRNA scaffold from ISYmul. This design demonstrated efficient plasmid interference activity (FIG. 27F). However, capping the 16-nucleotide guide region with an extended coRNA scaffold resulted in diminished activity (FIG. 27F), suggesting inefficient 3 ’-end processing. The design of cassette (I) can potentially be leveraged for multiplex editing by alternating the intact 127-nucleotide coRNA scaffold and guide region.
[0216] References for Examples 1-7
[0217] Al-Shayeb, B., Sachdeva. R., Chen, L. X., Ward, F., Munk. P., Devoto. A., Castelle, C. J., Olm, M. R., Bouma-Gregson, K., Amano, Y., He, C., Meheust, R., Brooks, B., Thomas,A., Lavy, A., Matheus-Camevali, P., Sun, C., Goltsman, D. S. A., Borton, M. A., . . . Banfield, J. F. (2020). Clades of huge phages from across Earth's ecosystems. Nature, 578(1195), 425- 431.
[0218] Al-Shayeb. B.. Skopintsev, P., Soczek, K. M., Stahl, E. C., Li, Z., Groover, E., Smock, D., Eggers, A. R., Pausch, P., Cress, B. F., Huang, C. J., Staskawicz, B., Savage, D. F., Jacobsen, S. E., Banfield, J. F., & Doudna, J. A. (2022). Diverse virus-encoded CRISPR-Cas systems include streamlined genome editors. Cell. 185(24), 4574-4586 e4516.
[0219] Aliaga Goltsman. D. S., Alexander. L. M., Lin, J. L., Fregoso Ocampo, R., Freeman,B., Lamothe, R. C., Perez Rivas, A., Temoche-Diaz, M. M., Chadha, S., NordenfelL N., Janson,O. P., Barr, I., Devoto, A. E., Cost, G. J., Butterfield, C. N., Thomas, B. C., & Brown, C. T. (2022). Compact Cas9d and HEARO enzymes for genome editing discovered from uncultivated microbes. Nat Commun, 13(1), 7602.
[0220] Altae-Tran, H., Kannan, S., Demircioglu, F. E., Oshiro, R., Nety, S. P., McKay, L. J., Dlakic, M., Inskeep, W. P., Makarova, K. S., Macrae, R. K., Koonin, E. V., & Zhang, F. (2021). The widespread IS200 / IS605 transposon family encodes diverse programmable RNA- guided endonucleases. Science, 374(6563), 57-65.
[0221] Altpeter, F., Springer, N. M., Bartley, L. E., Blechl, A. E., Brutnell, T. P., Citovsky, V., Conrad, L. J., Gelvin, S. B., Jackson, D. P., Kausch, A. P., Lemaux, P. G., Medford, J. I., Orozco-Cardenas, M. L., Tricoli, D. M., Van Eck, J., Voytas, D. F., Walbot, V., Wang, K., Zhang, Z. J., & Stewart, C. N., Jr. (2016). Advancing Crop Transformation in the Era of Genome Editing. Plant Cell, 28(7), 1510-1520.
[0222] Bigelyte, G., Young, J. K., Karvelis, T., Budre, K., Zedaveinyte, R., Djukanovic, V., Van Ginkel, E., Paulraj, S., Gasior, S., Jones, S., Feigenbutz, L., Clair, G. S., Barone, P., Bohn, J., Acharya, A., Zastrow-Hayes, G., Henkel-Heinecke, S., Silanskas, A., Seidel, R., & Siksnys, V. (2021). Miniature type V-F CRISPR-Cas nucleases enable targeted DNA modification in cells. Nat Commun, 12(1), 6191.
[0223] Bouma-Gregson, K., Crits-Christoph, A., Olm, M. R., Power, M. E., & Banfield, J. F. (2022). Microcoleus (Cyanobacteria) form watershed-wide populations without strong gradients in population structure. Mol Ecol, 31(1). 86-103.
[0224] Bouma-Gregson, K., Olm, M. R., Probst, A. J., Anantharaman, K., Power, M. E., & Banfield, J. F. (2019). Impacts of microbial assemblage and environmental conditions on the distribution of anatoxin-a producing cyanobacteria within a river network. ISME J, 13(6), 1618-1634.
[0225] Bradamante, G., Mittelsten Scheid. O.. & Incarbone. M. (2021). Under siege: virus control in plant meristems and progeny. Plant Cell, 33(8), 2523-2537.
[0226] Burch-Smith, T. M., Anderson, J. C., Martin, G. B., & Dinesh-Kumar, S. P. (2004). Applications and advantages of virus-induced gene silencing for gene function studies in plants. Plant J. 39(5), 734-746.
[0227] Burstein, D., Harrington, L. B., Strutt, S. C., Probst, A. J., Anantharaman, K., Thomas, B. C., Doudna, J. A., & Banfield, J. F. (2017). New CRISPR-Cas systems from uncultivated microbes. Nature, 542(7640), 237-241.
[0228] Chen, J. S., Dagdas, Y. S., Kleinstiver, B. P., Welch, M. M., Sousa, A. A., Harrington, L. B., Sternberg. S. H., Joung, J. K., Yildiz, A., & Doudna, J. A. (2017). Enhanced proofreading governs CRISPR-Cas9 targeting accuracy. Nature, 550(7676), 407-410.
[0229] Chen, J. S., Ma, E., Harrington, L. B., Da Costa, M., Tian, X., Palefsky, J. M., & Doudna, J. A. (2018). CRISPR-Casl2a target binding unleashes indiscriminate single-stranded DNase activity. Science, 360(6387), 436-439.
[0230] Chen. K.. Wang, Y„ Zhang, R.. Zhang. H„ & Gao, C. (2019). CRISPR / Cas Genome Editing and Precision Plant Breeding in Agriculture. Annu Rev Plant Riol, 70, 667-697.
[0231] Crits-Christoph, A., Olm, M. R., Diamond, S., Bouma-Gregson, K., & Banfield, J. F. (2020). Soil bacterial populations are shaped by recombination and gene-specific selection across a grassland meadow. ISMEJ. 14(7), 1834-1846.
[0232] East-Seletsky, A., O'Connell, M. R., Knight, S. C., Burstein, D., Cate, J. H., Tjian, R., & Doudna, J. A. (2016). Two distinct RNase activities of CRISPR-C2c2 enable guide-RNA processing and RNA detection. Nature, 538(7624), 270-273.
[0233] Ellison. E. E., Nagalakshmi, U., Gamo, M. E., Huang, P. J., Dinesh-Kumar, S., & Voytas, D. F. (2020). Multiplexed heritable gene editing using RNA viruses and mobile single guide RNAs. Nat Plants, 6(6), 620-624.
[0234] Ghoshal, B., Vong, B., Picard, C. L., Feng, S., Tam, J. M., & Jacobsen, S. E. (2020). A viral guide RNA delivery system for CRISPR-based transcriptional activation and heritable targeted DNA demethylation in Arabidopsis thaliana. PLoS Genet, 16(\2), el008983.
[0235] Hand, T. H., Das, A., & Li, H. (2019). Directed evolution studies of a thermophilic Type II-C Cas9. Methods Enzymol, 616, 265-288.
[0236] Harrington, L. B., Ma, E., Chen, J. S., Witte, I. P., Gertz, D., Paez-Espino, D., Al- Shayeb, B.. Kyrpides, N. C., Burstein, D., Banfield, J. F., & Doudna, J. A. (2020). A scoutRNA Is Required for Some Type V CRISPR-Cas Systems. Mol Cell. 79(3), 416-424 e415.
[0237] Haun, W., Coffman, A., Clasen, B. M., Demorest, Z. L., Lowy, A., Ray, E., Retterath, A., Stoddard, T., Juillerat, A., Cedrone, F., Mathis, L., Voytas, D. F., & Zhang, F. (2014). Improved soybean oil quality by targeted mutagenesis of the fatty' acid desaturase 2 gene family. Plant Biotechnol J. 12(7), 934-940.
[0238] Jinek, M., Chylinski, K., Fonfara, I., Hauer, M., Doudna, J. A., & Charpentier, E. (2012). A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity'. Science. 337(6096), 816-821.
[0239] Jinek. M„ East, A., Cheng. A., Lin, S„ Ma, E., & Doudna, J. (2013). RNA- programmed genome editing in human cells. Elife, 2, e00471.
[0240] Karvelis, T., Druteika, G., Bigelyte, G., Budre, K., Zedaveinyte, R., Silanskas, A., Kazlauskas, D., Venclovas. C., & Siksnys, V. (2021). Transposon-associated TnpB is a programmable RNA-guided DNA endonuclease. Nature, 599(7886), 692-696.
[0241] Kim, D. Y„ Lee, J. M„ Moon, S. B., Chin, H. J., Park, S„ Lim, Y , Kim, D., Koo, T., Ko, J. H., & Kim, Y. S. (2022). Efficient CRISPR editing with a hypercompact Casl2fl and engineered guide RNAs delivered by adeno-associated virus. Nat Biotechnol, 40(1), 94- 102.
[0242] Lemmon, Z. H., Reem, N. T., Dalrymple, J., Soyk, S., Swartwood, K. E., Rodriguez-Leal, D., Van Eck, J., & Lippman, Z. B. (2018). Rapid improvement of domestication traits in an orphan crop by genome editing. Nat Plants. 4(10), 766-770.
[0243] Li, Z., Zhong, Z., Wu, Z., Pausch, P., Al-Shayeb, B., Amerasekera, J., Doudna, J. A., & Jacobsen, S. E. (2023). Genome editing in plants using the compact editor CasPhi. Proc Natl Acad Set USA, 120(4), e2216822120.
[0244] Liu, D., Xuan, S., Prichard, L. E., Donahue, L. I., Pan, C., Nagalakshmi, U., Ellison, E. E., Starker, C. G., Dinesh-Kumar, S. P.. Qi. Y.. & Voytas, D. F. (2022). Heritable baseediting in Arabidopsis using RNA viral vectors. Plant Physiol, 189(4), 1920-1924.
[0245] Liu, L., Gallagher, J., Arevalo, E. D , Chen, R., Skopelitis, T., Wu, Q., Bartlett, M., & Jackson, D. (2021). Enhancing grain-yield-related traits by CRISPR-Cas9 promoter editing of maize CLE genes. Nat Plants, 7(3), 287-294.
[0246] Liu, Y.. Schiff, M., & Dinesh-Kumar, S. P. (2002). Virus-induced gene silencing in tomato. Plant J, 31(6), T11-T&6.
[0247] Liu, Y., Schiff, M., Marathe, R., & Dinesh-Kumar, S. P. (2002). Tobacco Rarl, EDS1 and NPR1 / NIM1 like genes are required for N-mediated resistance to tobacco mosaic virus. Plant J, 30(4), 415-429.
[0248] Ma. E.. Chen, K., Shi, H., Stahl. E. C., Adler. B.. Trinidad. M., Liu. J., Zhou, K., Ye, J., & Doudna, J. A. (2022). Improved genome editing by an engineered CRISPR-Casl2a. Nucleic Acids Res, 50(22), 12689-12701.
[0249] Macfarlane, S. A. (2010). Tobraviruses— plant pathogens and tools for biotechnology . Mol Plant Pathol, 11(4). 577-583.
[0250] Maher, M. F., Nash, R. A., Vollbrecht, M., Starker, C. G., Clark, M. D., & Voytas, D. F. (2020). Plant gene editing through de novo induction of meristems. Nat Biotechnol, 3S(1), 84-89.
[0251] Nagalakshmi. U., Meier. N., Liu, J. Y., Voytas, D. F., & Dinesh-Kumar, S. P. (2022). High-efficiency multiplex biallelic heritable editing in Arabidopsis using an RNA virus. Plant Physiol, 189(3), 1241-1245.
[0252] Nakagawa, R., Hirano, H., Omura, S. N., Nety, S., Kannan, S., Altae-Tran, H., Yao, X., Sakaguchi, Y., Ohira, T., Wu, W. Y., Nakayama, H., Shuto, Y., Tanaka, T., Sano, F. K., Kusakizako, T., Kise, Y., Itoh, Y., Dohmae, N., van der Oost. J., . . . Nureki, O. (2023). Cryo- EM structure of the transposon-associated TnpB enzyme. Nature. 616(7956). 390-397.
[0253] Nasti, R. A., & Voytas, D. F. (2021). Attaining the promise of plant gene editing at scale. Proc Natl Acad Sci USA, 118(22).
[0254] Olm, M. R., Crits-Christoph, A., Bouma-Gregson, K., Firek, B. A., Morowitz, M. J., & Banfield, J. F. (2021). inStrain profiles population microdiversity from metagenomic data and sensitively detects shared microbial strains. Nat Biotechnol, 39(6), 727-736.
[0255] Pausch, P., Al-Shayeb, B., Bisom-Rapp, E., Tsuchida, C. A., Li, Z., Cress, B. F., Knott, G. J., Jacobsen, S. E., Banfield, J. F., & Doudna, J. A. (2020). CRISPR-CasPhi from huge phages is a hypercompact genome editor. Science, 369(6501), 333-337.
[0256] Pausch. P., Soczek. K. M., Herbst, D. A.. Tsuchida, C. A., Al-Shayeb. B., Banfield, J. F., Nogales, E., & Doudna, J. A. (2021). DNA interference states of the hypercompact CRISPR-CasPhi effector. Nat Struct Mol Biol, 28(S), 652-661.
[0257] Ray, D. K„ Mueller, N. D„ West, P. C.. & Foley, J. A. (2013). Yield Trends Are Insufficient to Double Global Crop Production by 2050. PLoS One. 8(6), e66428.
[0258] Rodriguez-Leal, D., Lemmon, Z. H., Man, J., Bartlett, M. E., & Lippman, Z. B. (2017). Engineering Quantitative Trait Variation for Crop Improvement by Genome Editing. Cell, 171(2), 470-480 e478.
[0259] Rossner, C., Lotz. D., & Becker, A. (2022). VIGS Goes Viral: How VIGS Transforms Our Understanding of Plant Science. Annu Rev Plant Biol, 73, 703-728.
[0260] Sasnauskas, G., Tamulaitiene, G., Druteika, G., Carabias, A., Silanskas, A., Kazlauskas, D., Venclovas, C., Montoya, G., Karvelis, T., & Siksnys, V. (2023). TnpB structure reveals minimal functional core of Casl2 nuclease family. Nature, 616(7956), 384- 389.
[0261] Schuler, G., Hu, C., & Ke, A. (2022). Structural basis for RNA-guided DNA cleavage by IscB-omegaRNA and mechanistic comparison with Cas9. Science, 376(6600), 1476-1481.
[0262] Shi. G., Hao, M„ Tian, B„ Cao, G„ Wei, F„ & Xie, Z. (2021). A Methodological Advance of Tobacco Ratle Virus-Induced Gene Silencing for Functional Genomics in Plants. Front Plant Sci, 12, 671091.
[0263] Tsuchida, C. A., Zhang, S., Doost, M. S., Zhao, Y., Wang, J., O'Brien, E., Fang, H., Li, C. P., Li, D., Hai, Z. Y., Chuck, J., Brotzmann, J., Vartoumian, A., Burstein, D., Chen, X. W.. Nogales. E., Doudna, J. A., & Liu, J. G. (2022). Chimeric CRISPR-CasX enzymes and guide RNAs for improved genome editing activity. Mol Cell, 82(6), 1199-1209 el 196.
[0264] Wang, J. Y., Hoel, C. M., Al-Shayeb, B., Banfield, J. F., Brohawn, S. G., & Doudna, J. A. (2021). Structural coordination between active sites of a CRISPR reverse transcriptaseintegrase complex. Nat Commun, 12(\), 2571.
[0265] Wu, Z.. Zhang. Y.„ Yu, H.. Pan, D , Wang, Y., Wang. Y.. Li, F., Liu. C.„ Nan, H., Chen, W., & Ji, Q. (2021). Programmed genome editing by a miniature CRISPR-Casl2f nuclease. Nat Chem Biol, 77(11), 1132-1138.
[0266] Xu, X., Chemparathy, A., Zeng, L., Kempton, H. R., Shang, S., Nakamura, M., & Qi, L. S. (2021). Engineered miniature CRISPR-Cas system for mammalian genome regulation and editing. Mol Cell, 81(20), 4333-4345 e4334.
[0267] Yang, L., Machin, F., Wang, S., Saplaoura, E., & Kragler, F. (2023). Heritable transgene-free genome editing in plants by grafting of wild-type shoots to transgenic donor rootstocks. Nat Biotechnol.
[0268] Zsogon, A., Cermak, T., Naves, E. R.. Notini, M. M., Edel, K. H., Weinl, S., Freschi, L., Voytas, D. F., Kudla, J., & Peres, L. E. P. (2018). De novo domestication of wild tomato using genome editing. Nat Biotechnol.
[0269] References for Examples 8-11
[0270] Ellison. E. E., J. C. Chamness, and D. F. Voytas. 2021. "‘Viruses as Vectors for the Delivery of Gene-Editing Reagents.” htps: / / library.oapen.org / handle / 20.500.12657 / 61482.
[0271] Ellison, Evan E., Ugrappa Nagalakshmi, Maria Elena Gamo, Pin-Jui Huang, Savithramma Dinesh-Kumar, and Daniel F. Voytas. 2020. “Multiplexed Heritable Gene Editing Using RNA Viruses and Mobile Single Guide RNAs.” Nature Plants 6 (6): 620-24.
[0272] Karvelis, Tautvydas. Gytis Druteika, Greta Bigelyte, Karolina Budre, Rimante Zedaveinyte, Arunas Silanskas, Darius Kazlauskas, Ceslovas Venclovas, and Virginijus Siksnys. 2021. “Transposon- Associated TnpB Is a Programmable RNA-Guided DNA Endonuclease.” Nature 599 (7886): 692-96.
[0273] Liu. Degao, Shuya Xuan, Lynn E. Prichard, Lilee I. Donahue, Changtian Pan, Ugrappa Nagalakshmi, Evan E. Ellison, et al. 2022. “Heritable Base-Editing in Arabidopsis Using RNA Viral Vectors.’’ Plant Physiology 189 (4): 1920-24.
[0274] Nagalakshmi, Ugrappa, Nathan Meier, Jau-Yi Liu, Daniel F. Voytas, and Savithramma P. Dinesh-Kumar. 2022. “High-Efficiency Multiplex Biallelic Heritable Editing in Arabidopsis Using an RNA Virus.’’ Plant Physiology 189 (3): 1241-45.
[0275] Nety. Suchita P., Han Altae-Tran, Soumya Kannan, F. Esra Demircioglu. Guilhem Faure, Seiichi Hirano, Kepler Mears, Yugang Zhang, Rhiannon K. Macrae, and Feng Zhang. 2023. “The Transposon-Encoded Protein TnpB Processes Its Own mRNA into coRNA for Guided Nuclease Activity.” The CRISPR Journal 6 (3): 232-42.
[0276] Niu, Yajie, Randall William Shultz, Maria Margarita D. Unson, Michael Andreas Kock, and John P. Casey Jr. 2019. Novel plant cells, plants, and seeds. USPTO 20190352655:Al. US Patent, filed January 29, 2018, and issued November 21, 2019.
[0277] Tian, Ji, Haixia Pei, Shuai Zhang, Jiwei Chen, Wen Chen, Ruoyun Yang, Yonglu Meng, Jie You, Junping Gao, and Nan Ma. 2014. “TRV-GFP: A Modified Tobacco Rattle Virus Vector for Efficient and Visualizable Analysis of Gene Function.” Journal of Experimental Botany 65 (1): 311-22.
[0278] Xiang, Guanghai, Yuanqing Li, Jing Sun. Yongyuan Huo, Shiwei Cao, Yuanwei Cao, Yanyan Guo, et al. 2023. “Evolutionary Mining and Functional Characterization of TnpB Nucleases Identify Efficient Miniature Genome Editors.” Nature Biotechnology, June, 1-13.
[0279] Yoo, Sang-Dong, Young-Hee Cho, and Jen Sheen. 2007. “Arabidopsis Mesophyll Protoplasts: A Versatile Cell System for Transient Gene Expression Analysis.” Nature Protocols 2 (7): 1565-72.
[0280] Zhang, Xiuren, Rossana Henriques, Shih-Shun Lin, Qi-Wen Niu, and Nam-Hai Chua. 2006. “Agrobacterium-Mediated Transformation of Arabidopsis Thaliana Using the Floral Dip Method.” Nature Protocols 1 (2): 641-46.
[0281] References for Examples 12-15
[0282] Gruber, Andreas R., Ronny Lorenz, Stephan H. Bemhart, Richard Neubock, and Ivo L. Hofacker. 2008. “The Vienna RNA Websuite.” Nucleic Acids Research 36 (Web Server issue): W70-74.
[0283] Hino, Tomohiro, Satoshi N. Omura, RyoyaNakagawa, Tomoki Togashi, Satoru N. Takeda, Takafumi Hiramoto, Satoshi Tasaka, et al. 2023. “An AsCasl2f-Based CompactGenome-Editing Tool Derived by Deep Mutational Scanning and Structural Analysis.’' Cell 186 (22): 4920-35. e23.
[0284] Karvelis, Tautvydas, Gytis Druteika, Greta Bigelyte, Karolina Budre, Rimante Zedaveinyte, Arunas Silanskas, Darius Kazlauskas, Ceslovas Venclovas, and Virginijus Siksnys. 2021. “Transposon- Associated TnpB Is a Programmable RNA-Guided DNA Endonuclease."’ Nature 599 (7886): 692-96.
[0285] Kim, Do Yon, Jeong Mi Lee, Su Bin Moon. Hyun Jung Chin. Seyeon Park. Youjung Lim, Daesik Kim, Taeyoung Koo, Jeong-Heon Ko, and Yong-Sam Kim. 2021. “Efficient CRISPR Editing with a Hypercompact Casl2H and Engineered Guide RNAs Delivered by Adeno-Associated Virus. ’’ Nature Biotechnology' 40 (1): 94-102.
[0286] Kong, Xiangfeng, Hainan Zhang, Guoling Li, Zikang Wang, Xuqiang Kong, Lecong Wang, Mingxing Xue, et al. 2023. “Engineered CRISPR-OsCasl2fl and RhCasl2fl with Robust Activities and Expanded Target Range for Genome Editing.” Nature Communications 14 (1): 1-13.
[0287] Li, Zhifang, Ruochen Guo, Xiaozhi Sun, Guoling Li, Zhuang Shao, Xiaona Huo, Rongrong Yang, et al. 2024. “Engineering a Transposon-Associated TnpB-coRNA System for Efficient Gene Editing and Phenotypic Correction of a Tyrosinaemia Mouse Model.” Nature Communications 15 (1): 1-11.
[0288] Nakagawa, Ryoya, Hisato Hirano, Satoshi N. Omura, Suchita Nety, Soumya Karman, Han Altae-Tran, Xiao Yao, et al. 2023. “Cryo-EM Structure of the Transposon- Associated TnpB Enzyme.” Nature 616 (7956): 390-97.
[0289] Sasnauskas, Giedrius, Giedre Tamulaitiene, Gytis Druteika, Arturo Carabias, Arunas Silanskas, Darius Kazlauskas, Ceslovas Venclovas, Guillermo Montoya, Tautvydas Karvelis, and Virginijus Siksnys. 2023. “TnpB Structure Reveals Minimal Functional Core of Casl2 Nuclease Family.” Nature 616 (7956): 384-89.
[0290] Takeda, Satoru N. , Ryoya Nakagawa, Sae Okazaki, Hisato Hirano, Kan Kobayashi, Tsukasa Kusakizako, Tomohiro Nishizawa, Keitaro Yamashita, Hiroshi Nishimasu, and Osamu Nureki. 2021. “Structure of the Miniature Type V-F CRISPR-Cas Effector Enzy me.” Molecular Cell 81 (3): 558-70.e3.
[0291] Xiang, Guanghai, Yuanqing Li, Jing Sun, Yongyuan Huo, Shiwei Cao, Yuanwei Cao, Yanyan Guo, et al. 2023. “Evolutionary' Mining and Functional Characterization of TnpB Nucleases Identify Efficient Miniature Genome Editors.” Nature Biotechnology', June, 1-13.
[0292] Xu, Xiaoshu, Augustine Chemparathy, Leiping Zeng, Hannah R. Kempton, Stephen Shang, Muneaki Nakamura, and Lei S. Qi. 2021. "‘Engineered Miniature CRISPR-Cas System for Mammalian Genome Regulation and Editing. ” Molecular Cell 81 (20): 4333-45. e4.
[0293] Yoo, Sang-Dong, Young-Hee Cho, and Jen Sheen. 2007. '‘Arabidopsis Mesophyll Protoplasts: A Versatile Cell System for Transient Gene Expression Analysis.” Nature Protocols 2 (7): 1565-72.
[0294] References for Examples 16 and 17
[0295] Nagalakshmi, Ugrappa, Nathan Meier, Jau-Yi Liu, Daniel F. Voytas, and Savithramma P. Dinesh-Kumar. 2022. “High-Efficiency Multiplex Biallelic Heritable Editing in Arabidopsis Using an RNA Virus.” Plant Physiology 189 (3): 1241-45.
[0296] Zhang, Xiuren. Rossana Henriques. Shih-Shun Lin, Qi-Wen Niu. and Nam-Hai Chua. 2006. “Agrobacterium-Mediated Transformation of Arabidopsis Thaliana Using the Floral Dip Method.” Nature Protocols 1 (2): 641-46.Table 4: Vectors for TnpB-mediated genome editingTable 5: Selected sequences for Example 12Table 6: Additional Sequences for Example 12 (In Table 6. the gRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U)).Table 7: TnpB editing frequency associated with Example 13Table 8: Various TnpB sequences associated with Example 13Table 9: TAM Sequences and gRNA Sequences for Selected TnpBs (In Table 9, the gRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U)Table 10: Sequences Associated with Example 14 (In Table 10, the coRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U))Table 11 : Additional Sequences Associated with Example 14Table 12: Sequences associated with Example 15Table 13: Example TnpBs associated with Example 18Table 14: Selected TAM and gRNA sequences (In Table 14, the gRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U)Table 15: Editing Efficeincy of Selected TnpBs - Example 19Table 16: Various spacer sequences for selected TnpBs associated with Example 20Table 17: associated withTable 18: Selected coRNA sequences associated with Example 22 (In Table 18, the coRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U))Table 19: Vector and gRNA sequence associated with Example 25Table 20: Selected sequences associated with Example 8
[0297] Informal Sequence Listing:
[0298] ISBagOl - TnpB DNA sequence without codon optimization (SEO ID NO: 1):
[0299] atgtcgcagaccatcacagtcaaagtgaaattgcttccaacaaaagaacaggcttctattttgcgtgagatgagtcaaa cgtacctctccactatcaatactctcgtttccgaaatggttgctgaaaagaaaagcacaaagaagtcgagtaaggatatttctgcttctttg ccaagtgccgttaaaaatcaagctatcaaggacgcgaaaagtgtgtttaagaaagcgaagaaaagcaaattcactactgttcctgttttg aagaaacctgtctgcatttggaacaatcaaaactattcgtttgatttcacgcacatttccatgccgattatgattgacggcaaggtaactaa aacacctatccgtgctttgttagttgataaagacaaccgtaactttgctttgctaaaacacaaattaggtacgcttcgaatcacacagaaag
[0318] DNA sequence (250 nt) encoding for the coRNA sequence for ISBagOl TnpB (SEQ ID NO: 11):
[0322] DNA sequence (144 nt) encoding for the coRNA sequence for ISXfaOl TnpB (SEO ID NO: 13):
[0323] acatggccgtgagttccacggtatcagcctgtggagaggaaggcactggccctcgtcgcaagacggcggtgaaac cggcctcggtgaagcaggaagtcagctttgaatctgtttaaacagatttgagtaggtatgacagaacgg
[0324] ISYmul TnpB Native DNA Sequence (SEQ ID NO: 34)
[0328] ISDgel 0 TnpB Native DNA Sequence (SEQ ID NO: 36)A T
[0332] ISYmul TnpB Arabidopsis codon optimized DNA Sequence (SEQ ID NO: 38)
[0340] ISYmul TnpB Maize codon optimized DNA Sequence (SEO ID NO: 42)
[0341] TTGCTcCAgCAcAAGGCCTACGAGTAcAGAATCTATCCAGATAAGAAgC AGGAGACTCTcATCGCTAAgACCATcGGCTCTTCCAGGTAcGTTTACAACCACTTT CTCGAgTTGTGGAATAAGGAgTACGAGGAgACTGGGAAgGGACTGACTTACTAcG CATGCTCGAAgCTCCTcACCAAgTTGAAgAGAGACCCGGAgACCGTGTGGCTGTGTGAGGTTGACAAgTTCTCTCTCCAgAAcTCACTGAGGAACCTCTCTGACGCCTTcTC ACGCTTCTTcAAGGGGCAgAACGAGCATCCGCAGTTcAAgTCCAAGAAGTCACCCC GGCAgTCCTACACAACGCAgTATACTAAcAAcAATATcGCCGTGAGCGGTAACTGC CTTAAgCTcCCGAAGCTcGGcTTGGTGAAGTTCGCCGACTCACGGGAgATGAAgGG CCGGATcCTCAAcGCCACTGTCAGACGGAAgAGCAGCGGGAAgTTCTTCGTgTCGA TcTTGTGTGAGGAgGAgATCTGTGAGCTCCCGAAgACTGACTCATCTGTCGGTATC GATCTcGGGATcATcGACTTcGCAGTGATGTCCGACGGTTCACGGCATGATAAcAA CCACTTcACTCGGCAGATGGAgGAgCGCCTGCGGCGGGAGCAGAGgAAgCTCGCC CGCAGAGCCCTcGCAGCCGAgAAGCGCGGGATCTCGCTCTCAGAGGCGCGGAAcTACCAgAAGCAGCGGCGGAAGGTCGCGAGATTGTAcGAgAAgGTGGCTAAcCAgCG C A AGGAgT AcCTcAAc AAgCTc AGC AC GGAg ATcGTTAAgAAcC AC GATATcATcTGT ATCGAGGACCTCAATGTTAAGGGGATGATGCGGAAcCACAAGTTGGCCAAgTCAA TCAGCGACGTGTCCTGGACTTCTCTCGTgTCCAAGCTGCAATAcAAgGCGTCGTGG TACGGcAAgGAgGTCATcAGAATcTCGCGGTGGTTCCCCTCTTCGCAGATcTGCTCC GAGTGCGGTCACAAgGACCGCAAGAAGCCTCTCCATGTTAGGGAGTGGACTTGC CCCGTGTGCCACGCGCATCATGACAGAGATGTTAACGCAGCAAGGAACATcTTGG CTGAgGGCCTGAGAATcCGGGCACTcACACCAGGAAGCTAG
[0342] ISAaml TnpB Maize codon optimized DNA Sequence (SEQ ID NO: 43)
[0346] ISDra2 TnpB Maize codon optimized DNA Sequence (SEQ ID NO: 45)
[0348] ISYmul TnpB amino acid sequence (SEQ ID NO: 46)
[0350] ISAaml TnpB amino acid sequence (SEQ ID NO: 47)
[0352] ISDgelO TnpB amino acid sequence (SEQ ID NO: 48)
[0354] ISDra2 TnpB amino acid sequence (SEQ ID NO: 49)
[0356] ISYmul coRNA DNA sequence (SEQ ID NO: 50)
[0358] ISAaml coRNA DNA sequence (SEQ ID NO: 51)
[0360] ISDgelO coRNA DNA sequence (SEQ ID NO: 52)
[0362] ISDra2 coRNA DNA sequence (SEO ID NO: 53)
[0363] gatcaagaatcccgaagtgaagaatctgccgtccgtacatggactgcccgaactgtggggaaacccatgaccga gaCgagaacgctgcgctgaacatcggcgtgaagcgttggtggctgcgggaatctcagacacctaaacgctcatggaggctatgtc agacctgcttcggcgggcaatggtctgcgaagtgagaatcacgcgactttagtcgtgtgaGGTTCAA
[0364] ISYmul TnpB with coRNA overlap native DNA sequence (single) (SEQ ID NO: 140)ACGTCTTGCTCGTCTCCGAGTTTGAGATGCCCGATGTACCGCTGCCCCCCGTCCCT GAGTCCGAGGTGGTCGGGATTGACCTCGGCCTGAAGGATTTCTACGTGTTGTCCG
[0372] ISYmul TnpB with wRNA overlap Arabidopsis codon optimized DNA sequence (single) (SEQ ID NO: 144)
[0374] ISAaml TnpB with wRNA overlap Arabidopsis codon optimized DNA sequence (single) (SEQ ID NO: 145)
[0376] ISDgel O TnpB with wRNA overlap Arabidopsis codon optimized DNA sequence(single) (SEP ID NO: 146)
[0380] tRNA-Ile (SEQ ID NO: 148)
[0382] pPEBV (SEQ ID NO: 149)
[0384] TnpB30 amino acid sequence (SEQ ID NO: 702)
[0386] In the foregoing description, it will be readily apparent to one skilled in the art that varying substitutions and modifications may be made to the invention disclosed herein without departing from the scope and spirit of the invention. The invention illustratively described herein suitably may be practiced in the absence of any element or elements, limitation or limitations which is not specifically disclosed herein. The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention that in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention. Thus, it should be understood that although the presentinvention has been illustrated by specific embodiments and optional features, modification and / or variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention.
[0387] All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples provided herein, is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention unless otherw ise claimed. No language in the specification should be constmed as indicating any non-claimed element as essential to the practice of the invention.
[0388] Citations to a number of patent and non-patent references are made herein. The cited references are incorporated by reference herein in their entireties. In the event that there is an inconsistency between a definition of a term in the specification as compared to a definition of the term in a cited reference, the term should be interpreted based on the definition in the specification.
Claims
CLAIMS1. A composition, comprising: a TnpB bacterial enzyme; and a polynucleotide comprising a first portion adapted to bind to at least a portion of the TnpB bacterial enzyme and a second portion for binding to a plant genomic target.
2. The composition of claim 1, wherein the polynucleotide is RNA.
3. The composition of claim 1 or 2, wherein the TnpB bacterial enzyme comprises less than about 500 amino acids, less than about 450 amino acids, less than about 420 amino acids, or less than about 400 amino acids.
4. The composition of any one of claims 1-3, wherein the TnpB bacterial enzyme is from, or derived from, tnpB from Brevibacillus agri.
5. The composition of any one of claims 1-3, wherein the TnpB bacterial enzyme comprises ISBagOl, ISYMul, or TnpB30.
6. The composition of any one of claims 1-5, wherein the TnpB bacterial enzyme comprises a peptide tag.
7. The composition of claim 6, wherein the peptide tag comprises a FLAG tag.
8. A construct comprising a plant promoter operably connected to a polynucleotide encoding for a TnpB bacterial enzyme.
9. The construct of claim 8, wherein the polynucleotide further encodes for an coRNA having a first portion adapted to bind to at least a portion of the TnpB bacterial enzyme and a second portion for binding to a plant genomic target.
10. The construct of claim 9, wherein the coRNA is about 100-400 nucleotides in length.
11. The construct of any one of claims 8-10, wherein the polynucleotide encoding for a TnpB bacterial enzyme is codon optimized for expression in plant cells.
12. The construct of any one of claims 8-11, wherein the plant promotor comprises a UBQ10 gene promoter.
13. The construct of any one of claims 9-12, wherein a coding sequence of the coRNA and a coding sequence of the TnpB bacterial enzyme at least partly overlap.
14. The construct of any one of claims 8-13, wherein the TnpB bacterial enzyme comprises less than about 500 amino acids, less than about 450 amino acids, less than about 420 amino acids, or less than about 400 amino acids.
15. The construct of any one of claims 8-14, wherein the TnpB bacterial enzyme is from, or derived from, tnpB from Brevibacillus agri.
16. The construct of any one of claims 8-14, wherein the TnpB bacterial enzyme comprises ISBagOl, ISYMul, or TnpB30.
17. The construct of any one of claims 8-16, wherein the TnpB bacterial enzyme comprises a peptide tag.
18. The construct of claim 17, wherein the peptide tag comprises a FLAG tag.
19. A plant viral vector comprising the construct of any one of claims 8-18, 26, 27, 29, 31, or 32.
20. The plant viral vector of claim 19, wherein the plant viral vector is derived from the Tobacco rattle virus (TRV).
21. A plant cell comprising the construct of any one of claims 8-18, 26, 27, 29, 31, or 32 or the plant viral vector of claims 19 or 20.
22. A method for gene editing a plant, comprising: introducing the construct of any one of claims 8-18, 26, 27, 29, 31, or 32 and the plant viral vector of claims 19 or 20 into one or more plant cells.
23. The method of claim 22, wherein the one or more plant cells are Arabidopsis cells.
24. The composition of any one of claims 1-4, 6, and 7, wherein the TnpB bacterial enzyme comprises an amino acid sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the amino acid sequence of SEQ ID NO:
525. The composition of any one of claims 1-3 and 5-7, wherein the TnpB bacterial enzyme comprises an amino acid sequence that is 70 % or more. 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the amino acid sequence of SEQ ID NO: 46.
26. The composition of any one of claims 1-3 and 5-7. wherein the TnpB bacterial enzyme comprises an amino acid sequence that is 70 % or more. 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the amino acid sequence of SEQ ID NO: 702.
27. The construct of any one of claims 8-15.
17. and 18, wherein the polynucleotide encoding for the TnpB bacterial enzyme comprises a polynucleotide sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the polynucleotide sequence of one or more of SEQ ID NOs: 1-4.
28. The construct of any one of claims 8-14 and 16-18, wherein the polynucleotide encoding for the TnpB bacterial enzyme comprises a polynucleotide sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more. 90 % or more, 95 % or more, or 99 % more identical to the polynucleotide sequence of one or more of SEQ ID NOs: 34, 38, or 42.
29. The composition of any one of claims 1-7, 24, 25, and 26 wherein the polynucleotide is complementary to a sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the polynucleotide sequence of one or more of SEQ ID NOs: 11-13.
30. The construct of any one of claims 9-18, 26, and 27, wherein the polynucleotide comprises a sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the polynucleotide sequence of one or more of SEQ ID NOs: 11-13.
31. The construct of any one of claims 8-14, 17, or 18 wherein the polynucleotide encoding for the TnpB bacterial enzyme comprises a polynucleotide sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the polynucleotide sequence of one or more of SEQ ID NOs: 34-45.
32. The construct of claim 13, wherein the polynucleotide encoding for the TnpB bacterial enzyme and the ®RNA comprises a polynucleotide sequence that is 70 % or more. 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the polynucleotide sequence of one or more of SEQ ID NOs: 140-147, 150-170, 431-433, or 475-481.