Transposase-assisted homology-independent targeted integration
The engineered nucleic acid modification system using a nuclease-deficient transposase and CRISPR/Cas nuclease addresses the challenge of precise genome integration by enabling targeted insertion of transgenes in eukaryotic cells, improving transgene stability and expression.
Patent Information
- Application Number
- PCT/US2025/031305
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-25
- Filing Date
- 2025-05-29
- Publication Date
- 2025-12-04
AI Technical Summary
Current genome engineering techniques face challenges in achieving accurate and efficient insertion of transgenes into specific locations in the genome, particularly in organisms with low frequencies of Homologous Recombination (HR) and Homology-Directed Repair (HDR), leading to random insertions, mutations, and inconsistent gene expression.
An engineered nucleic acid modification system using a nuclease-deficient transposase and a programmable targeting nuclease, such as a CRISPR/Cas system, to excise and insert donor polynucleotides at user-defined target loci, guided by transposition sequences and programmable nucleases.
Enables precise and targeted integration of transposable nucleic acid constructs into eukaryotic cells, reducing off-target insertions and ensuring stable transgene expression, as demonstrated in plant cells and seeds.
Smart Images

Figure US2025031305_04122025_PF_FP_ABST
Abstract
Description
Danforth Ref. DDPSC0158-401 -PCT Via Patent CenterTRANSPOSASE-ASSISTED HOMOLOGY-INDEPENDENT TARGETED INTEGRATIONGOVERNMENTAL RIGHTS
[0001] This invention was made with government support under Award No. 24311 I IOS- 2149964 awarded by the National Science Foundation. The government has certain rights in the invention.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority from Provisional Application number 63 / 654,366, filed May 31 , 2024 and Provisional Application number 63 / 724,584, filed November 25, 2024. The contents of each of the aforementioned applications are hereby incorporated by reference in their entirety.INCORPORATION OF SEQUENCE LISTING
[0003] The present application contains a Sequence Listing that has been submitted in .XML format via Patentcenter and is hereby incorporated herein by reference in its entirety. Said XMI was created on May 23, 2025, is named DDPSC0158_sequence_listing, and is 370 kilobytes in size.FIELD OF THE INVENTION
[0004] The present disclosure provides systems and methods of generating genetically modified cells comprising no off-target insertions.BACKGROUND OF THE INVENTION
[0005] Genome engineering is a revolutionary technology that promises the ability to improve or overcome current deficiencies in the genetic code as well as to introduce novel functionality. However, some applications of the technology do not always generate completely reliable results. For instance, transgene integration of foreign DNA into or near genes can generate new mutations or alter the regulation of nearby genes, while insertions into heterochromatic regions are often not permissive to the desired high levels of transgene expression or do not provide stable expression over multiplegenerations. Further, in most instances, when performing transgenesis, the transgene frequently inserts into the nuclear genome in a random location. This can lead to new mutations at the insertion locus and at unintended insertion points, gene silencing, and general inconsistencies in experiments or products. For instance, in plants, where the frequency of homologous recombination is less than 1%, efficient and accurate insertion of transgenes is possible only in theory and is often associated with uncontrolled deletions of neighboring regions, as well as rearrangement of the transgene sequences. In fact, in a typical scenario, it simply is not possible to obtain the optimal, desired change. Additionally, although recently developed tools such as CRISPR systems have allowed biologists to target random genetic modifications to specific regions of genomes, accurate nucleic insertions in target loci is still a major challenge. In plants, this is because Homologous Recombination (HR) and Homology-Directed Repair (HDR) of donor sequences into the targeted locus occurs at a very low frequency.
[0006] Therefore, a long-felt need exists for improved and effective means of inserting polynucleotides into a user-defined location in the genome, especially in organisms where the frequency of HR and HDR are low, including plants.SUMMARY OF THE INVENTION
[0007] One aspect of the instant disclosure encompasses an engineered nucleic acid modification system for generation of a genetically modified cell. The system comprises one or more nucleic acid constructs for expressing a nuclease-deficient transposase and a programmable targeting nuclease and a donor polynucleotide comprising transposition sequences compatible with the nuclease-deficient transposase. The donor polynucleotide is optionally comprised in a source polynucleotide. The one or more nucleic acid constructs comprise an expression construct for expressing the nuclease-deficient transposase, the expression construct comprising a promoter operably linked to a nucleic acid sequence encoding the nuclease-deficient transposase; and an expression construct for expressing the programmable targeting nuclease, wherein the expression construct comprises a promoter operably linked to a nucleic acid sequence encoding the programmable targeting nuclease. The programmable targeting nuclease is engineered to excise the donor polynucleotide from the source polynucleotide,wherein the nuclease-deficient transposase recognizes and binds the transposition sequences of the donor polynucleotide, and wherein the programmable targeting nuclease is engineered to introduce a cut in a target nucleic acid locus in the cell thereby guiding insertion of the excised donor polynucleotide at the target nucleic acid locus to generate a genetically modified cell comprising the transposable nucleic acid construct inserted at the target nucleic acid locus.
[0008] In some aspects, the programmable nuclease-deficient transposase is linked to the programmable targeting nuclease. In some aspects, the nuclease-deficient transposase is linked to the programmable targeting nuclease. The nuclease- deficient transposase can be a nuclease-deficient split transposase. The nuclease- deficient split transposase can be a Pong or Pong-like split transposase comprising a Pong ORF1 protein and a Pong ORF2 protein, and the nuclease-deficient Pong ORF2 protein can comprise a modification that inactivates the nuclease function of the Pong ORF2 protein. The nuclease-deficient transposase can be a Pong or Ponglike split transposase comprising a Pong ORF1 protein and a nuclease-deficient Pong ORF2 protein, and the nuclease-deficient Pong ORF2 protein can comprise a modification in a nuclease catalytic site of the Pong ORF2 protein that inactivates the nuclease function of the Pong ORF2 protein. The nuclease-deficient transposase can be a Pong or Pong-like split transposase comprising a Pong ORF1 protein and a nuclease-deficient Pong ORF2 protein, and the nuclease-deficient Pong ORF2 protein can comprise a modification in a DDE catalytic site of the nuclease-deficient Pong ORF2 protein that inactivates the nuclease function of the Pong ORF2 protein.
[0009] The Pong ORF1 protein can comprise an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 1. The nuclease-deficient Pong ORF2 protein can comprise an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 121 . The expression construct for expressing the nuclease- deficient transposase can comprise an expression construct for expressing the Pong ORF1 protein, wherein the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100; and anexpression construct for expressing the nuclease-deficient Pong ORF2 protein, wherein the expression construct for expressing the nuclease-deficient Pong ORF2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 120.
[0010] The nuclease-deficient split transposase can be a Pong or Pong-like split transposase comprising a Pong ORF1 protein, and the Pong ORF2 protein can be excluded from the system. The engineered system can comprise an expression construct for expressing the Pong ORF1 protein, wherein the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100, and the Pong ORF2 protein can be excluded from the system.
[0011] The donor polynucleotide can comprise a cargo polynucleotide flanked by the transposition sequences compatible with the nuclease-deficient transposase of the system. The transposition sequences can be transposition sequences of a miniature inverted-repeat transposable element (MITE). The MITE can be an mPing MITE. Transposition sequences of the mPing MITE can comprise mPing inverted repeat 1 and mPing inverted repeat 2, wherein mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 7, SEQ ID NO: 111 , or SEQ ID NO: 108, and mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 8, SEQ ID NO: 112, or SEQ ID NO: 109. Transposition sequences of the mPing MITE can comprise mPing inverted repeat 1 and mPing inverted repeat 2, wherein mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 108, and mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 109.
[0012] The programmable targeting nuclease can comprise a programmable, sequence-specific nucleic acid-binding domain and a nuclease domain. Theprogrammable targeting nuclease can be an RNA-guided clustered regularly interspersed short palindromic repeats (CRISPR) nuclease system, a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, a ssDNA-guided Argonaute endonuclease, a meganuclease, a rare- cutting endonuclease, or any combination thereof. The programmable targeting nuclease can be a CRISPR-associated (Cas) (CRISPR / Cas) nuclease system comprising a nuclease and a guide RNA (gRNA). The programmable targeting nuclease can comprise a Cas9 nuclease and a gRNA. The Cas9 nuclease can comprise an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 5.
[0013] The Cas9 nuclease can be linked to the nuclease-deficient Pong ORF2 protein. The nuclease-deficient Pong ORF2 protein can be linked to the Cas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64. The nuclease-deficient Pong ORF2 protein can be linked to the Cas9 nuclease by three copies of a G4S linker of SEQ ID NO: 64. In some aspects, the Cas9 nuclease is not linked to the nuclease-deficient Pong ORF2 protein. The engineered system can comprise a nucleic acid expression construct for expressing a Cas9 nuclease, wherein the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94.
[0014] The engineered system can comprise one or more nucleic acid expression constructs for expressing one or more excision gRNAs that guide the Cas9 nuclease to nucleic acid sequences in or flanking the transposition sequences of the donor polynucleotide to guide excision of the donor polynucleotide from the source polynucleotide and for expressing a targeting gRNA that guides the Cas9 nuclease to the target nucleic acid locus to introduce a cut at the target nucleic acid locus, thereby guiding insertion of the excised donor polynucleotide at the target nucleic acid locus to generate a genetically modified cell. The one or more excision gRNAs can comprise a nucleic acid sequence of SEQ ID NO: 118. The targeting gRNA that guides the Cas9 nuclease to the target nucleic acid locus to introduce a cut at thetarget nucleic acid locus can comprise a nucleic acid sequence of SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 80, SEQ ID NO: 113, SEQ ID NO: 67 and SEQ ID NO: 113, or any combination thereof. In some aspects, the one or more nucleic acid expression constructs for expressing the excision and targeting gRNAs comprise a nucleic acid sequence comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 122 and wherein a nucleic acid sequence starting at base 425 to base 444 is replaced with a nucleic acid sequence of a gRNA for targeting the transposase and nuclease to a target nucleic acid locus.
[0015] The engineered system comprises: an expression construct for expressing the Pong ORF1 protein; an expression construct for expressing the Cas9 nuclease;
[0016] one or more expression constructs for expressing an excision gRNAs and a targeting gRNA; and a source polynucleotide comprising the donor polynucleotide. In some aspects, the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100; the expression construct for expressing a Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94; the excision gRNA comprises a nucleic acid sequence of SEQ ID NO: 118; and the donor polynucleotide comprises mPing inverted repeat 1 and mPing inverted repeat 2, wherein mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 108, and mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 109.
[0017] In some aspects, the engineered system comprises: an expression construct for expressing the Pong ORF1 protein; an expression construct for expressing the modified nuclease deficient Pong ORF2 protein; an expression construct for expressing the Cas9 nuclease; one or more expression constructs for expressing anexcision gRNAs and a targeting gRNA; and a source polynucleotide comprising the donor polynucleotide. In some aspects, the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100; the expression construct for expressing the nuclease-deficient Pong ORF2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 120; the expression construct for expressing a Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94; the excision gRNA comprises a nucleic acid sequence of SEQ ID NO: 118; and the donor polynucleotide comprises mPing inverted repeat 1 and mPing inverted repeat 2, wherein mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 108, and mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 109.
[0018] The donor polynucleotide can comprise a cargo polynucleotide. In some aspects, the cargo polynucleotide comprises HSEs. In other aspects, the cargo polynucleotide comprises an expression construct for expressing a herbicide resistance function. When the cargo polynucleotide comprises an expression construct for expressing a herbicide resistance function, the herbicide resistance function is resistance to bialaphos herbicide, resistance to glyphosate herbicides, or both. In some aspects, the cargo polynucleotide comprises an expression construct for expressing resistance to bialophos herbicide, wherein the expression construct comprises a promoter operably linked to a polynucleotide encoding resistance to bialaphos, and wherein the donor polynucleotide comprising the expression construct for expressing resistance to bialophos herbicide comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequenceof SEQ ID NO: 97 or SEQ ID NO: 99. In some aspects, the cargo polynucleotide comprises an expression construct for expressing resistance to bialophos herbicide, wherein the expression construct comprises a promoter operably linked to a polynucleotide encoding resistance to bialaphos. In some aspects, the donor polynucleotide comprising the expression construct for expressing resistance to bialophos herbicide comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97.
[0019] The cargo polynucleotide can also comprise an expression construct for expressing resistance to glyphosate herbicide, wherein the expression construct comprises a promoter operably linked to a polynucleotide encoding resistance to glyphosate. The donor polynucleotide comprising the expression construct for expressing resistance to glyphosate herbicide can comprise a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 124The engineered system of any one of the preceding claims, wherein the target nucleic acid locus is in a nuclear, organellar, or extrachromosomal nucleic acid sequence.
[0020] The target nucleic acid locus can be in a protein-coding gene, an RNA coding gene, or an intergenic region. The cell can be a eukaryotic cell. In some aspects, the cell is a plant cell, a plant or part thereof, or seed. In some aspects, the plant is an Arabidopsis sp. or a soybean plant cell, a plant or part thereof, or seed.
[0021] Another aspect of the instant disclosure encompasses an engineered nucleic acid modification system for generation of a genetically modified cell. The system comprises: one or more nucleic acid expression constructs for expressing a nuclease-deficient transposase and a programmable targeting nuclease, wherein the nuclease-deficient transposase is a Pong or Pong-like split transposase comprising a Pong ORF1 protein and a Pong ORF2 protein and wherein the modified nuclease- deficient Pong ORF2 protein comprises an absence of a Pong ORF2 protein. The one or more nucleic acid constructs comprise: an expression construct for expressing the Pong ORF1 protein, the expression construct comprising a promoter operably linked to a nucleic acid sequence encoding the ORF1 protein; and anexpression construct for expressing the programmable targeting nuclease, wherein the expression construct comprises a promoter operably linked to a nucleic acid sequence encoding the programmable targeting nuclease. The one or more nucleic acid constructs also comprises a donor polynucleotide comprising transposition sequences compatible with the nuclease-deficient transposase, wherein the donor polynucleotide is optionally comprised in a source polynucleotide. The programmable targeting nuclease is engineered to excise the donor polynucleotide from the source polynucleotide, wherein the nuclease-deficient transposase recognizes and binds the transposition sequences of the donor polynucleotide, and the programmable targeting nuclease is engineered to introduce a cut in a target nucleic acid locus in the cell thereby guiding insertion of the excised donor polynucleotide at the target nucleic acid locus to generate a genetically modified cell comprising the transposable nucleic acid construct inserted at the target nucleic acid locus. The transposition sequences can be transposition sequences of a miniature inverted-repeat transposable element (MITE). In some aspects, the MITE is an mPing MITE.
[0022] The programmable targeting nuclease can be a CRISPR-associated (Cas) (CRISPR / Cas) nuclease system comprising a nuclease and a guide RNA (gRNA).
[0023] In some aspects, the programmable targeting nuclease comprises a Cas9 nuclease and a gRNA.
[0024] Yet another aspect of the instant disclosure encompasses one or more nucleic acid constructs encoding an engineered nucleic acid modification system described herein above.
[0025] One aspect of the instant disclosure encompasses a cell comprising an engineered system described herein above or one or more nucleic acid constructs described herein above. The cell can be a eukaryotic cell. In some aspects, the cell is a plant cell, a plant or part thereof, or seed.
[0026] An additional aspect cell method of generating a genetically modified cell comprising no off-target insertion. The method comprises introducing an engineered nucleic acid modification system for generation of a genetically modified or one or more nucleic acid constructs encoding the engineered nucleic acid modification system into the cell; maintaining the cell under conditions and for a time sufficient forthe one or more nucleic acid constructs for expressing the nuclease-deficient transposase, expressing the programmable targeting nuclease, for the programmable targeting nuclease to excise the donor polynucleotide from a source polynucleotide, and for the donor polynucleotide to be inserted in the target locus to generate a genetically modified cell comprising the donor polynucleotide inserted in the target locus; optionally identifying an insertion of the donor polynucleotide in the nucleic acid locus in the cell; and optionally confirming an absence of off-target insertions of the transposable nucleic acid construct. The engineered nucleic acid modification system and the one or more constructs encoding the engineered nucleic acid modification system can be as described herein above. In some aspects, an inserted donor polynucleotide replaces a target locus. In some aspects, the cell is a eukaryotic cell. The eukaryotic cell can be a plant cell, a plant or part thereof, or seed. In some aspects, the cell is ex vivo.
[0027] Another aspect of the instant disclosure encompasses a kit for generating a genetically modified cell comprising no off-target insertion. The kit comprises one or more engineered nucleic acid modification systems or one or more nucleic acid constructs encoding the one or more engineered nucleic acid modification systems. The one or more engineered nucleic acid modification systems and the constructs encoding the one or more engineered nucleic acid modification systems can be as described herein above. Each of the engineered systems generates a genetically modified cell comprising an accurate insertion of the donor polynucleotide into the target nucleic acid locus. In some aspects, the kit comprises one or more cells comprising one or more engineered systems, one or more nucleic acid constructs, or any combination thereof. In some aspects, the one or more cells are eukaryotic. The eukaryotic cell can be a plant cell, a plant or part thereof, or seed.BRIEF DESCRIPTION OF THE FIGURES
[0028] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0029] FIG. 1 is a diagram depicting an engineered system excising a donor polynucleotide from a donor site in a plant and inserting the excised donor polynucleotide into a locus in the Arabidopsis PDS3 gene.
[0030] FIG. 2 depicts a schematic overview of twelve different transgenes comprising Cas9 and derivative proteins linked either to the N- or C-terminus of Pong transposase ORF1 (blue) or to the N- or C-terminus of Pong ORF2 (orange) protein coding regions. Three different versions of Cas9 were used: double-strand cleavage Cas9, the single stranded nickase deCas9, and the catalytically dead dCas9.
[0031] FIG. 3A. The functional verification of ORF1 / 2 and Cas9 fusion proteins. GFP fluorescence was detected for all 12 fusion proteins as well as the ORF1 / ORF2 positive control, since mPing excision from the GFP donor site restores the GFP expression. The negative control without ORF1 / ORF2 (-ORF1 -ORF2) was not able to excise mPing.
[0032] FIG. 3B. The functional verification of ORF1 / 2 and Cas9 fusion proteins. A functional CRISPR / Cas9 system when linked to ORF1 / 2 was verified through the observation of white seedlings and sectors in plants generated from the Cas9 targeting of the Arabidopsis PDS3 gene with all four Cas9 fusion proteins. Three examples of individual plants are shown.
[0033] FIG. 4A. Screening insertions. PCR strategy to detect targeted insertions into the PDS3 gene. mPing can insert in the forward or reverse orientation relative to PDS3.
[0034] FIG. 4B. Screening insertions. PCR with negative controls: a line lacking the ORF1 / ORF2 proteins (mPing only), lacking Cas9 (mPing+ ORF1 / ORF2) and a no template PCR (-). The expected amplification sizes are indicated by black arrowheads. The correct PCR products validated by Sanger sequencing are marked with red arrows.
[0035] FIG. 4C. Screening insertions. Replicate of the PCR from clone #2 in FIG. 4B. This PCR displays the correct sized and sequenced bands (red arrows) in each reaction.
[0036] FIG. 5 depicts nucleic acid sequences at insertion sites of 9 unique transposition events. The sequence of the mPing transposable element is green. The target site duplication sequence is red. The guide RNA target site is grey highlighted. The PDS gene is unhighlighted black. For simplicity, only the mPing / PDS3 junction of these sequences are shown.
[0037] FIG. 6A. PCR strategy to determine if any transgenic DNA would insert at a Cas9 cleavage site. The PCR shows no bands of expected size (black arrowheads), which demonstrates that mPing insertion from FIG. 4 is a product of transposition, and not random.
[0038] FIG. 6B. Testing if the single components of the system could recapitulate the results. No Cas9 and ORF1 / 2 (mPing only), no Cas9 (+ORF1 / 2), and no ORF1 / 2 (+Cas9) controls each failed to produce the expected band and therefore cannot generate targeted insertions. Having Cas9 and ORF1 / 2, but in an un-linked configuration, produced targeted insertion. The lane to the far right is clone #2 from FIG. 4A, which is used as a positive control in this experiment. The four gels represent the same four PCR assays from FIG 4B. Black arrowheads denote the expected size of the targeted insertion in each PCR.
[0039] FIG. 7A is a diagram showing the three systems designed with gRNAs targeted to three different target loci: the PDS3 gene, the ADH1 gene, and the promoter of ACT8 gene.
[0040] FIG. 7B are the Sanger sequencing results of junctions of target insertions into the PDS3 gene, the ADH1 gene, and the promoter of ACT8 gene. The sequence below mPing is the expected sequence of a perfect “seamless” insertion. The chromatograms above the sequence show the sequences at the insertion sites. The highlighted bases are 1-2 nucleotide insertions or deletions.
[0041] FIG. 8A depicts a PCR strategy to detect targeted insertions into the PDS3 gene. mPing can insert in either the forward direction (above the PDS3 region) or reverse direction (below the PSD3 region). The location of 4 PCR primers (R,L,U,D) are shown for orientation.
[0042] FIG. 8B depicts an agarose gel run of PCR products using primers from FIG. 8A from systems comprising ORF1 and 2 linked or unlinked to Cas9 nuclease. Arrowheads denote the correct size of the PCR products for each set of primers. No Cas9 and ORF1 / 2 (“mPing only”), no Cas9 (“+ORF1 / 2”), and no ORF1 / 2 (“+Cas9”) are negative controls and showed no bands.
[0043] FIG. 9A is a diagram of a vector that contains the CRISPR / Cas9 system (including gRNA), the mPing donor element, and ORF1 and ORF2 transposase proteins.
[0044] FIG. 9B depicts a PCR strategy to detect targeted insertions into the PDS3 gene using the vector of FIG. 9A. mPing can insert in either the forward direction (above the PDS3 region) or reverse direction (below the PSD3 region). The location of 4 PCR primers (R,L,U,D) are shown for orientation.
[0045] FIG. 9C depicts PCR detection of mPing targeted insertion in the Arabidopsis genome using the vector in FIG. 9A. PCR detection used primer sets from FIG. 9B.
[0046] FIG. 10 depicts targeted insertion based on the Pong / mPing transposon system. Fusion of the Pong transposase ORFs with Cas9 provides the transposase sequence specificity for the insertion of the non-autonomous mPing element. The mPing element is excised out of a donor site provided on the transgene, generating fluorescence. mPing insertion at the target site is screened for by PCR.
[0047] FIG. 11 depicts the Experimental Design of Protein Fusions and Testing. Twelve different transgenes where created and transformed into Arabidopsis. Cas9 and derivative proteins where linked either to the Pong transposase ORF1 (blue) or ORF2 (orange) protein coding regions. Both N- and C-terminal fusions were created. Three different versions of Cas9 were used: double-strand cleavage Cas9, the single stranded nickase deCas9, and the catalytically dead dCas9. When a functional transposase protein is generated by expression of ORF1 and ORF2, it excises the mPing transposable element out of the 35S-GFP donor location, producing fluorescence. The goal of this project was to demonstrate user-defined targeted insertion of the mPing transposable element by programming the CRISPR-Cas9 system with a custom guide RNA.
[0048] FIG. 12A depicts photographs showing fluorescence generated upon excision of mPing from the 35S:GFP donor site. mPing only transposes in the presence of both ORF1 and ORF2 transposase proteins, and fusing ORF2 to Cas9 still results in mPing excision.
[0049] FIG. 12B depicts a PCR gel showing excision as in FIG. 12A assayed by PCR using primers at the 35S:GFP donor site. A smaller sized band is generated upon mPing excision.
[0050] FIG. 12C depicts a PCR assay to detect targeted insertion of mPing at PDS3 gene. Primer names (U,L,R,D) and locations are listed above. Targeted insertion is detected via PCR in plants that have all three proteins: ORF1 , ORF2 and Cas9.Targeted insertions are detected when ORF2 and Cas9 are physically linked, or when unlinked but present in the same cells.
[0051] FIG. 12D depicts a cartoon of mPing excision and targeted insertion when ORF2 is linked to Cas9.
[0052] FIG. 12E depicts an example of a Sanger sequence read of the junction between the PDS3 gene and the targeted insertion of mPing.
[0053] FIG. 12F depicts sequence analysis of 17 distinct insertion events of mPing at PDS3. mPing sequences are shown in yellow, and the target site duplication of TTA / TAA from the donor site is shown in red. Within the PDS3 target site, the gRNA targeted sequence is shown in grey. The mPing is inserted between the third and fourth base of the gRNA target sequence (black arrowhead). The variation of the sequence found on either end of the insertion site is shown.
[0054] FIG. 12G depicts a plot showing the number of SNPs at the insertion site identified by Sanger sequencing targeted insertion events.
[0055] FIG. 13A depicts photographs showing the functional verification of ORF1 / 2 and Cas9 fusion proteins. GFP fluorescence was detected for all 12 fusion proteins as well as the ORF1 / ORF2 positive control, since mPing excision from the GFP donor site restores the GFP expression. The negative control without ORF1 / ORF2 (-ORF1 -ORF2) was not able to excise mPing.
[0056] FIG. 13B depicts the functional verification of ORF1 / 2 and Cas9 fusion proteins. A functional CRISPR / Cas9 system when linked to ORF1 / 2 was verified through the observation of white seedlings and sectors in plants with all four Cas9 fusion proteins. Three examples of individual plants are shown.
[0057] FIG. 14A depicts a PCR strategy to detect targeted insertions into the PDS3 gene. mPing can insert in the forward or reverse orientation relative to PDS3.
[0058] FIG. 14B depicts an electrophoresis gel of PCR products with negative controls: a line lacking the ORF1 / ORF2 proteins (mPing only), lacking Cas9 (mPing+ORF1 / ORF2) and a no template PCR (-). The expected amplification sizes are indicated by black arrowheads. The correct PCR products are marked with red arrows.
[0059] FIG. 14C depicts screening insertions. Replicate of the PCR from clone #2. This PCR displays the correct sized bands (red arrows) in each reaction.
[0060] FIG. 15 depicts the comparison of the number of base deletions (left of zero on the X-axis) and insertions (right of zero on the X-axis) for two configurations of Cas9 and ORF2: linked and unlinked. Insertions of mPing (red) into PDS3 (blue) were subject to amplicon deep sequencing and each junction analyzed separately. Since mPing can insert in either orientation (black arrows within red mPing elements), four distinct junction points are analyzed. The size of the black filled circle represents the percentage of deep sequenced reads.
[0061] FIG. 16A depicts additional controls. PCR strategy to determine if any transgenic DNA would insert at a Cas9 cleavage site. The PCR shows no bands, which demonstrates that mPing insertion from FIGs. 12A-13B is a product of transposition, and not random.
[0062] FIG. 16B depicts additional controls. Testing if the single components of our system could recapitulate our results. No Cas9 and ORF1 / 2 (mPing only), no Cas9 (+ORF1 / 2), and no ORF1 / 2 (+Cas9) controls each failed to produce the expected band and therefore cannot generate targeted insertions. Having Cas9 and ORF1 / 2, but in an un-linked configuration, produced targeted insertion. The lane to the far right is clone #2 from FIGs. 12-12G, which is used as a positive control in this experiment. The four gels represent the same four PCR assays from FIG. 12C. Black arrowheads denote the expected size of the targeted insertion in each PCR.
[0063] FIG. 17A depicts an overview of targeted insertion at 3 distinct loci. By switching the CRISPR gRNA, distinct regions of the genome are targeted for mPing insertion.
[0064] FIG. 17B depicts how mPing can insert into DNA for both directions. Arrows indicate primers used to detect target insertions: U, upstream of target gene; D, downstream of target gene; R, right end of mPing; L, left end of mPing. PCR products were then purified and sequenced.
[0065] FIG. 17C depicts sanger sequencing chromatograms for junctions of target insertions into an additional target besides PDS3: ADH1.
[0066] FIG. 17D depicts sanger sequencing chromatograms for junctions of target insertions into an additional target besides PDS3: ACT8 promoter.
[0067] FIG. 18 depicts analysis of the left and right junctions of mPing targeted insertions upstream of the ACT8 gene in T2 plants with Cas9 linked to ORF2. Singleindividual T2 plants were assayed one-by-one, and 8 plants were confirmed by Sanger sequencing to have targeted insertions of mPing.
[0068] FIG. 19A. Addition of 6 heat shock element (HSE) sequences originally upstream of a heat-shock responsive gene into mPing and cartoon of attempted targeted insertion upstream of the ACT8 gene. The individual HSEs are shown as red bars in the mPing- HSE element.
[0069] FIG. 19B. PCR gel of mPing element excision from the donor location demonstrating that the modified mPing-HSE element could excise properly. The Sspl digest is performed to improve the assay’s sensitivity. AtADHI is shown as a PCR control.
[0070] FIG. 19C PCR gel detecting targeted insertions. Both a pool of T2 plants was assayed, as well as four individual T2 generation plants. Bands with red arrow heads are the correct size and were Sanger sequenced to demonstrate the correct targeted insertion into the promoter region of the ACT8 gene. AtADHI is shown as a PCR control.
[0071] FIG 19D Sanger sequencing results of the junction of mPing-HSE inserted at its target site upstream of the ACT8 gene. The red highlighted two bases are deleted compared to the predicted seamless insertion.
[0072] FIG 19E Sanger sequencing through the mPing-HSE element inserted upstream of ACT8 as in FIG19D. The PCR primers used to generate this amplicon are shown above. Below, all 6 delivered HSEs are shown as red arrows and in this example a 11 base deletion is detected at the junction between mPing-HSE and the upstream region of ACT8.
[0073] FIG. 20 depicts experimental design to use targeted transposition of a modified mPing element in order to transcriptionally rewire the ACT8 gene. The goal is to engineer the ACT8 gene have transcriptional activation during heat stress.
[0074] FIG. 21 A depicts a map of the vector testing the ability of unlinked Cas9 Nickase to direct targeted insertions of mPing. Targeted insertion into ADH1 has been detected at a low frequency and sequenced. This insertion shows the left junction of mPing at ADH1 with a 14 bp deletion.
[0075] FIG. 21 B depicts further experimentation demonstrating that dCas9 can participate in targeted insertion when two gRNAs are used. In this case, thetransposase is inserting mPing at a TTA site nearby the gRNA target sites. The Sanger sequencing of one end of mPing is shown.
[0076] FIG. 21 C depicts the experimental design to use of two gRNAs and a catalytically active Cas9 protein. In this example, a region of DNA is cut out of the genome with two gRNAs and replaced with mPing.
[0077] FIG. 21 D PCR primer placement for screening mPing targeted insertion.
[0078] FIG. 21 E shows targeted insertion screening assay. Red arrowheads are PCR products that were Sanger sequenced and verified targeted insertions.
[0079] FIG. 21 F shows one end of a targeted insertion that replaces the DNA between the two gRNAs used.
[0080] FIG. 22A Vector maps of TDNAs used for a two-step (two-component) transformation. The donor vector was transformed into Arabidopsis first, and a stable transgenic line was used for a second transformation using the helper vector.
[0081] FIG. 22B The one-component vector containing both donor TE (mPing) and helpers (ORF1 , ORF2-Cas9) was also tested to be able to direct targeted insertion. Blue triangles are LB and RB ends of the T-DNA. Arrows denote promoters, and black boxes are terminators. The mPing donor TE is shown in red.
[0082] FIG. 23A depicts the vector for transposase-mediated targeted insertion of mPing into the soybean (Glycine max) crop genome. Soybean transformation vector with a gRNA that targets the “DD20” non-protein coding region of the soybean genome, using an unlinked ORF2 and Cas9 configuration.
[0083] FIG. 23B depicts the vector for transposase-mediated targeted insertion of mPing into the soybean (Glycine max) crop genome. Similar vector as in FIG. 23A, but with a linked ORF2 and Cas9.
[0084] FIG. 23C depicts the transposase-mediated targeted insertion of mPing into the soybean (Glycine max) crop genome. The overall goal of targeted insertion of mPing into the DD20 non-protein coding region of the soybean genome without previously integrating and new sequences such as a landing pad for targeted insertion.
[0085] FIG. 23D depicts the transposase-mediated targeted insertion of mPing into the soybean (Glycine max) crop genome. PCR primer strategy to detect targeted insertion (top) and PCR gel (bottom). Bands with red arrowheads are the correct size and werevalidated by Sanger sequencing. Two out of nine transgenic soybean plants showed targeted insertion of mPing.
[0086] FIG. 23E depicts the transposase-mediated targeted insertion of mPing into the soybean (Glycine max) crop genome. Top is the Sanger sequence example of a targeted insertion into the soybean genome (plant RO #8 from FIG. 23D). Bottom is an example of mPing-HSE inserted into DD20 in the soybean genome.
[0087] FIG 23F depicts the constructs used for transposase-mediated targeted insertion of mPing into the soybean (Glycine max) crop genome. The seven mPing constructs test how to functionally fuse ORF2 to Cas9 in soybean, and if the mPing-HSE and m Ping-bar cargos can be delivered to specific sites in the soybean genome.
[0088] FIG23G depicts the transposase-mediated targeted insertion of mPing into the soybean (Glycine max) crop genome. The percent of plants tested with excision of mPing (top left), mutagenesis of the target location by Cas9 (top right), plants with combined excision and mutagenesis (bottom left), and targeted insertion of mPing at the DD20 location in the soybean genome (bottom right).
[0089] FIG. 24A depicts the four mPing constructs used to determine mPing sequences required for transposition and to test longer cargo sequences. Each of these has the tested capability to excise from the genome and participate in targeted integration.
[0090] FIG. 24B depicts an electrophoresis gel of PCR products testing the ability of the mPing constructs from FIG. 24A to excise out of the donor position. Blue triangle denote the size of the mPing constructs at the donor site, and the smaller band the same position after successful mPing excision. The mPing element with only the TIRs (mPing TIR_bar gene) does not excise efficiently.
[0091] FIG. 24C depicts an electrophoresis gel of PCR products targeted insertion of mPing and the mPing_bar CDS to the non-coding region upstream of the ACTIN8 gene. Red triangles denote the correct PCR product for a targeted insertion.
[0092] FIG. 25A depicts an electrophoresis gel of PCR products showing the excision of each of the mPing derived constructs mPing_bar CDS and mPing_bar gene from the donor position. Each pool of plants displays mPing excision.
[0093] FIG. 25B depicts the PCR strategy and primer placement for screening targeted insertion events. The m Ping-bar CDS and m Ping-bar versions of mPing can insert into the targeted location in either orientation.
[0094] FIG. 25C depicts an electrophoresis gel of PCR products showing the targeted insertion of mPing_bar CDS and mPing_bar gene upstream of the ACTIN8 gene. Red triangles denote PCR products of the correct size for a targeted insertion event.
[0095] FIG. 25D depicts the rate of mPing element excision (left) and targeted insertion (right) for different mPing versions in T1 Arabidopsis plants.
[0096] FIG. 26A depicts a map of the construct comprising the bar CDS in mPing inserted into the ACT8 gene. This insertion shows the right junction of mPing_bar CDS at ACT8 with a 2 bp deletion.
[0097] FIG. 26B shows Sanger sequencing results of bar CDS in mPing inserted into the ACT8 gene of FIG. 26A aligned to the expected sequence of targeted insertion showing the 2 bp deletion. Red regions are mPing sequence, grey highlighted are the bar gene coding region, and green is the promoter region upstream of A CT8.
[0098] FIG. 26C is a sequence alignment.
[0099] FIG. 27A depicts a map of the construct comprising the bar gene with the bar promoter and terminator elements in mPing inserted into the ACT8 gene. This insertion shows the right junction of mPing_bar gene at ACT8 with a 2 bp deletion.
[0100] FIG. 27B shows Sanger sequencing results of bar in mPing inserted into the ACT8 gene of FIG. 27A aligned to the expected sequence of targeted insertion showing the 2 bp deletion. Red regions are mPing sequence, grey highlighted are the Nos promoter+bar gene+Nos terminator, and green is the promoter region upstream of ACT8.
[0101] FIG. 28A shows that the m Ping-bar targeted insertion confers the herbicide resistance trait. Amplicons “PCR1” to “PCR6” are used to genotype for the presence of the mPing-bar transgene in R0 transformed soybean plants.
[0102] FIG. 28B shows PCR results of the PCR targets in FIG 28A. GmLel is a control gene.
[0103] FIG. 28C shows PCR primer placement in order to assay for the mPing- bar targeted insertion.
[0104] FIG. 28D shows the PCR assay for targeted insertion in the DD20 targeted location in the soybean genome. Red arrowheads denotes targeted insertions that were verified by Sanger sequencing.
[0105] FIG. 29A is a diagrammatic depiction of sequential transformation of DD45::Cas9 plants with mPing construct containing all components of the system, except Cas9.
[0106] FIG. 29B is the excision assay of mPing out of the donor transgene.
[0107] FIG. 29C is the PCR to detect targeted insertions.
[0108] FIG. 29D is the Sanger sequencing of a targeted insertion of mPing into the ACT8 region of the Arabidopsis genome.
[0109] FIG. 29E is a diagram of the measurement of the rate of excision and targeted insertion in the DD45::Cas9 line.
[0110] FIG. 30A is a diagram depicting experimental design of a system used to replace a section of genomic DNA in the non-coding region upstream of the ACT8 gene of Arabidopsis.
[0111] FIG. 30B is a diagram depicting the PCR strategy used to detect the DNA replacement.
[0112] FIG. 30C is an image of an electrophoresis gel showing the results of the PCRs of FIG. 30B.
[0113] FIG. 30D shows the Sanger sequencing of one of the bands in FIG. 30C confirming replacement of the genomic DNA with mPing at the ACT8 region of the Arabidopsis genome.
[0114] FIG. 31 A shows diagrams illustrating Transposase-Assisted Target Site Integration (TATSI), Homology-Independent Targeted Insertion (HITI), TATSI-DDE wherein the amino acid residues at the “DDE” catalytic site of the ORF2 transposase are mutated, and when TATSI-DDE and HITI are combined (TATSI-DDE + HITI).
[0115] FIG. 31 B are bar plots showing the rate of excision of mPing donor DNA (left panel) and rate of targeted insertion of the mPing donor DNA (right panel).
[0116] FIG. 32A are bar plots showing the rate of targeted insertion of the mPing donor DNA using the TATSI (unfused), TATSI (fused), HITI, and TAHITI systems.
[0117] FIG. 32B are bar plots showing the number of insertions of the mPing donor DNA using the TATSI (unfused), TATSI (fused), HITI, and TAHITI systems.
[0118] FIG. 32C is a plot showing the location of the single targeted insertion obtained when using the TAHITI system at the ACT8 promoter in Arabidopsis.
[0119] FIG. 33A shows a diagram of donor DNA comprising mPing flanked by a target site of a gRNA for guiding Cas9 to excise mPing from the donor DNA in HITI, TAHITI-ORF1 / dORF2, and TAHITI-ORF1 systems shown in FIG. 33B.
[0120] FIG. 33B shows diagrams illustrating constructs used for Transposase- Assisted Target Site Integration wherein the ORF2 protein is not fused to the Cas9 protein (TATSI (unfused)), Homology-Independent Targeted Insertion (HITI), TATSI- DDE wherein the amino acid residues at the “DDE” catalytic site of the ORF2 transposase are mutated (ORF1 / dORF2), when TATSI-DDE and HITI are combined (TAHITI-ORF1 / dORF2), and when TATSI-DDE and HITI are combined and the ORF2 expression construct is absent (TAHITI-ORF1 ).
[0121] FIG. 33C shows bar plots showing the rate of excision of mPing donor DNA (left panel) and rate of targeted insertion of the mPing donor DNA (right panel) using the TATSI (unfused), HITI, ORF1 / dORF2, and TAHITI-ORF1 systems.
[0122] FIG. 33D shows map diagrams representing the two possible insertions of mPing into the ACT8 gene.
[0123] FIG. 33E shows the Sanger sequencing of mPing insertions in FIG. 33D confirming insertion of mPing at the ACT8 region of the Arabidopsis genome.
[0124] FIG. 34A shows bar plots showing the rate of excision of mPing donor DNA (left panel) and rate of targeted insertion of the mPing donor DNA (right panel) using the TATSI (unfused) and ORF1 / dORF2 systems in soybean.
[0125] FIG. 34B shows the Sanger sequencing of mPing insertions in FIG. 34A confirming insertion of mPing in the soybean genome.
[0126] FIG. 35. mPing insertion sites in pooled seedlings. The Arabidopsis genome is displayed on the x-axis. The upstream ACT8 target site is shown with an arrow and red data point. The scale of each y-axis is determined by the maximum data point. A dashed line at 10,000 RPM is shown for each sample. There is one biological replicate of control line (top left panel), two distinct biological replicates of TATSI lines (top center and right panels), three distinct biological replicates of HITI lines (center panels), and three distinct biological replicates of TAHITI lines (bottom panels).
[0127] FIG. 36A. Nucleotide variation at the junction of mPing insertions into upstream ACT8. The precision of each nucleotide at the insertion site was determined on the 5’ junction or 3’ junction. Four amplicons capturing the junctions of mPinginsertions were analyzed for TATSI. The size of the circle represents the percentage of reads where that nucleotide is as expected (y-axis = 0), has a deletion (y-axis <-1).
[0128] FIG. 36B. Nucleotide variation at the junction of mPing insertions into upstream ACT8. The precision of each nucleotide at the insertion site was determined on the 5’ junction or 3’ junction. Four amplicons capturing the junctions of mPing insertions were analyzed for HITI. The size of the circle represents the percentage of reads where that nucleotide is as expected (y-axis = 0), has a deletion (y-axis <-1).
[0129] FIG. 36C. Nucleotide variation at the junction of mPing insertions into upstream ACT8. The precision of each nucleotide at the insertion site was determined on the 5’ junction or 3’ junction. Four amplicons capturing the junctions of mPing insertions were analyzed for TAHITI_ORF1 / dORF2. The size of the circle represents the percentage of reads where that nucleotide is as expected (y-axis = 0), has a deletion (y-axis <-1 ).
[0130] FIG. 36D. Nucleotide variation at the junction of mPing insertions into upstream ACT8. The precision of each nucleotide at the insertion site was determined on the 5’ junction or 3’ junction. Four amplicons capturing the junctions of mPing insertions were analyzed for TAHITI_ORF1 . The size of the circle represents the percentage of reads where that nucleotide is as expected (y-axis = 0), has a deletion (y- axis <-1 ).
[0131] FIG. 36E. Calculated seamlessness of insertions using TATSI, HITI, and TAHITI. First panel shows seamlessness as calculated using the following formula:Second panel shows the average seamlessness calculated from data shown in the first panel.
[0132] FIG. 37A. P rog ram m ability of the cargo. The cargo of different mPing versions demonstrated to excise and undergo targeted insertion in the Arabidopsis genome. NOS P:: bar:: NOS T is an expression cassette that expresses an herbicide resistance gene, and EPSPS_TIPA is a mutated Arabidopsis gene that grants herbicide resistance. mPing versions are not drawn to scale and the size of each is indicated.
[0133] FIG. 37B. P rog ram m ability of the cargo. Measurement of the excision frequency of mPing from the donor site (left) and rate of targeted insertion (right). n= the number of T1 transgenic Arabidopsis plants analyzed. The color of each data bar corresponds to the mPing cargo color code in FIG. 37A.DETAILED DESCRIPTION
[0134] The present disclosure encompasses engineered nucleic acid modification systems and methods for generating genetically modified cells and organisms. The disclosed systems and methods efficiently mediate targeted insertion of polynucleotides, even in organisms where genetic manipulation relying on homologous recombination or homology-directed repair is problematic, including plants. Unlike existing insertion systems that depend on homologous recombination or homology- directed repair to insert or replace nucleic acid sequences, the disclosed systems and methods enable controlled and high-efficiency targeted insertion of a polynucleotide of choice to generate a genetically modified cell with the polynucleotide inserted at a target nucleic acid locus in a gene of interest. In some aspects, the insertion replaces a nucleic acid sequence in the cell. Furthermore, the compositions and methods insert polynucleotides without introducing unwanted mutations in the transferred polynucleotide or the nucleic acid sequences at the target locus.
[0135] The present disclosure provides an engineered nucleic acid modification system for generating genetically modified cells with no or reduced off-target insertions, addressing a common challenge with transposable element systems that cause unintended transposition events, including off-target insertions catalyzed by the transposase's nuclease activity and activation of endogenous transposable elements. The engineered system achieves this through a transposase-assisted homologydependent targeted integration (TAHITI) system of the instant disclosure, which combines the targeting capabilities of a programmable targeting nuclease with a nuclease-deficient transposase. In this system, the nuclease activity of the transposase is replaced by the nuclease activity of the programmable targeting nuclease, significantly reducing or eliminating off-target insertions. The TAHITI system also generates precise junction sequences at insertion sites, ensuring high accuracy and stability of genetic modifications. This improved precision and lack of off-targetinsertions is accomplished simultaneously with a higher overall efficiency rate of targeted insertion. This approach bypasses host-encoded homologous recombination or damage repair pathways typically used for polynucleotide introduction, while simultaneously minimizing off-target effects and enhancing the precision of targeted genetic modifications.I. Composition
[0136] One aspect of the present disclosure encompasses an engineered nucleic acid modification system (the “engineered system”) for generating a genetically modified cell. The engineered system is a transposon-assisted homology-independent targeted integration (TAHITI) system. The TAHITI system comprises a nuclease-deficient transposase, a programmable targeting nuclease, and a transposable nucleic acid construct. The nuclease-deficient transposase retains its DNA-binding function, enabling it to recognize and bind transposition sequences of a transposable nucleic acid construct to form a transposition complex or "transpososome" during a transposition event. In a system of the instant disclosure, the nuclease functions of the transposase that functions in excising a transposable element (or a donor polynucleotide) and cutting a genomic sequence for insertion of the transposable element in a new location of the genome, are replaced by the nuclease function of a programmable targeting nuclease. The programmable targeting nuclease is programmed to excise the donor polynucleotide from a source polynucleotide and introduce a cut at a target nucleic acid locus, thereby guiding the insertion of the excised transposable nucleic acid construct at the target locus. This process results in the insertion of the donor polynucleotide at the target nucleic acid locus to generate a genetically modified cell (See fourth panel of FIG. 31 A). Importantly, the inventors discovered that the absence of the nuclease activity of the transposase significantly reduces or eliminates off-target insertions catalyzed by the transposase's nuclease activity. Additionally, the absence of nuclease activity minimizes activation of endogenous transposable elements.
[0137] The programmable targeting nuclease, the transposase, and the donor polynucleotide are described in further detail below. In some aspects, the cell is a plant cell, a plant or part thereof, or seed.(a) Transposase
[0138] The engineered system of the instant disclosure comprises a transposase. As used herein, the term “transposase” refers to a protein or a protein fragment derived from any transposable element (TE), wherein the transposase is capable of cutting or copying a donor polynucleotide from a nucleic acid sequence comprising the donor polynucleotide, protecting the donor polynucleotide from degradation by binding to transposable element sequences in the donor polynucleotide and forming a transpososome comprising the donor polynucleotide and the transposase protein, inserting the donor polynucleotide at a target locus using the nuclease function of the transposase, or any combination thereof. TEs can be assigned to any one of two classes according to their mechanism of transposition, which can be described as either copy and paste (Class I TEs) or cut and paste (Class II TEs).
[0139] Class I TEs are retrotransposons that copy and paste themselves into different genomic locations in two stages: first, TE nucleic acid sequences are transcribed from DNA to RNA, and the RNA produced is then reverse transcribed to DNA. This copied DNA is then inserted back into the genome at a new position. The reverse transcription step is catalyzed by a reverse transcriptase activity, which is often encoded by the TE itself. Non-limiting examples of Class I TEs include Tnt1 , Opie, Huck, and BARE1.
[0140] The transposition mechanism of Class II TEs does not involve an RNA intermediate. The transpositions are catalyzed by a transposase enzyme that binds transposition sequences of the TE, cuts out the transposon, and positions it for ligation into the target site. Non-limiting examples of Class II TEs include P Instability Factor (PIF), Ac / Ds, Pong TE or Pong-like TEs, Spm / dSpm, Harbinger, P-elements, Tn5 and Mutator.
[0141] Transposases generally recognize and interact with compatible transposition sequences at the ends of the TE to mediate transposition of the TE. For instance, in Class II TEs, the transposase can have nucleic acid binding sequences that bind the transposition sequences at the terminal ends of the TE and can cleave the DNA, removing the TE from the excision / donor site, can protect the TE ends from degradation while it is outside the chromosome in a protein-DNA complex termed transpososome, and can cleave the insertion site at a new location in the genome of acell and integration of the TE at the insertion site. For Class I TEs, the transposases of some TEs recognize the terminal transposition sequences at the ends of an RNA transcript of the TE, reverse transcribe the transcript into DNA, then cleave and integrate the TE at the insertion site. Accordingly, a transposase of the instant disclosure can be any transposase or a fragment or derivative thereof, provided the transposase recognizes the compatible terminal transposition sequences of the donor polynucleotide and mediates insertion of the polynucleotide at the target locus. Transposition sequences compatible with the transposase can be as described in Section 1(b) below.
[0142] In an engineered system of the instant disclosure, a transposase recognizes the transposition sequences of the donor polynucleotide and can facilitate formation of the transpososome (transposition complex). When the transposase is derived from a Class I TE, the transposase first transcribes the donor polynucleotide into an RNA transcript and reverse transcribes the RNA transcript to DNA for insertion at the target locus. When the transposase is derived from a Class II TE, the nuclease function of the transposase first cleaves the donor polynucleotide from a source nucleic acid sequence such as a nucleic acid construct comprising the donor polynucleotide for insertion at the target locus. The transposase remains bound to the polynucleotide using the DNA-binding function of the transposase and can facilitate formation of a transposition complex while it is outside the chromosome among other functions.
[0143] As described herein above, the system of the instant disclosure comprises a nuclease-deficient transposases, wherein the nuclease functions of the transposase of the system are replaced by the nuclease function of the programmable targeting nuclease. As used herein, the term nuclease-deficient transposase refers to a transposase protein or protein complex that retains the DNA binding function of a transposase but lacks nuclease activity responsible for excising the transposable element from a source polynucleotide and cleaving DNA at transposition sites. Nonlimiting methods for generating a nuclease-free transposase include inactivating the nuclease function by mutating the catalytic residues of the nuclease domain, deleting the nuclease domain of the transposase, expressing only the DNA binding domain of the transposase, or in the case of a split transposase, omitting the subunit responsible for nuclease activity. However a nuclease deficient transposase is generated, thenuclease-deficient transposase of a system of the instant disclosure retains its ability to bind transposition sequences and facilitate the transposition process, relying on an external programmable targeting nuclease to perform the DNA cleavage required for excision from a source polynucleotide and insertion. As explained herein above, the inventors demonstrated that removing the nuclease function of the transposase reduces off-target transposition events and improves the precision of targeted genetic modifications.
[0144] In some aspects, the transposase is derived from a Class II TE. Class II transposable elements (TEs), also known as DNA transposons, can be grouped into numerous superfamilies based on sequence homology, transposase structure, and mechanism of transposition. Non-limiting examples of Class II TEs include the Tcl / mariner superfamily of TEs, the hAT superfamily of TEs, the Mutator superfamily of TEs, the CACTA superfamily of TEs, the PIF / Harbinger superfamily of TEs, the Kolobok superfamily of TEs, the Zator superfamily of TEs, the PiggyBac superfamily of TEs, the Helitron superfamily of TEs, the Polinton / Maverick superfamily of TEs, the Transib superfamily of TEs, the Merlin superfamily of TEs, the Rehavkus superfamily of TEs, the Novosib superfamily of TEs, the Ginger superfamily of TEs. and the Academ superfamily of TEs.
[0145] In some aspects, a transposase of the instant disclosure is a Class II TE comprising a split or bipartite transposase. In a Class II TE comprising a split or bipartite transposase, the transposase is split to comprise two proteins. A first protein provides the DNA binding function of the transposase and the formation of the transpososome function, and a second protein comprises the nuclease function of the transposase. Non-limiting examples of Class II TEs comprising a split transposase include transposases from Kolobok transposons, Zator Transposons, ISL2EU transposons, P Instability Factor (PIF) / Harbinger Superfamily of TEs, including the PIF or PIF-like family of TEs, Pong or Pong-like family of TEs, Tourist or Tourist-like family of TEs. In some aspects, the transposase is derived from the P Instability Factor (PIF) family of TEs or P / F-like family of TEs.
[0146] In some aspects, the transposase is a Pong or Pong-like transposase. The transposases of the Pong and Pong-like TEs are split transposases comprising a first protein encoded by open reading frame 1 (ORF1 protein) and a second proteinencoded by open reading frame 2 (ORF2 protein) of the TE. ORF1 provides the DNA binding ftranspososome formation functions of the transposes© and ORF2 comprises the nuclease function of the transposase. As explained herein above, the nuclease function of the transposase is engineered to be absent. Accordingly, as explained herein further below, in some aspects, when the transposase of the system is a Pong transposase, the system can comprise ORF1 of a Pong transposase and a mutated nuclease-deficient ORF2. Alternatively, in some aspects, when the transposase of the system is a Pong transposase, the system can comprise only the ORF1 of a Pong transposase; i.e,, ORF2 is absent or excluded. Put another way, when the nuclease- deficient nuclease of the system of the instant disclosure is a Pong or Pong-like transposase, the system can be engineered to omit a Pong ORF2 protein.
[0147] ORF1 of a Pang or Pong-like transposase comprises a Myb / SANT DNA- binding domain. ORF1 can also comprise residues or domains for formation of the transpososome which can include domains or residues for mediating protein-protein interactions. Accordingly, a nuclease-deficient transposase of a system of the instant disclosure can also comprise any protein, natural or manmade, that comprises a Myb / SANT DNA-binding domain, and can comprise domains for formation of the transpososome, provided the protein can bind transposition sequences of a transposable nucleic acid construct of a system of the instant disclosure. Myb / SANT DNA-binding domains comprise tandem repeats of Myb motifs — each repeat is -50 amino acids and includes an HTH-like fold. Outside of transposable elements, Myb-like (or Myb / SANT) domains are found in a diverse range of eukaryotic transcription factors and chromatin-associated proteins, where they play central roles in gene regulation, chromatin remodeling, and development. Non-limiting examples of DNA binding protein comprising a Myb / SANT DNA-binding domain include MYB Transcription Factors (animals and plants), R2R3-MYB Proteins, SANT Domain-Containing Proteins, TRF Family (Telomeric Repeat-binding Factors), and MYT Transcription Factors (e.g., MYT1 ).
[0148] When a transposase of the instant disclosure is a Pong or Pong-like transposase, the engineered system can comprise a nuclease-deficient transposase comprising only ORF1. In other words, a nuclease deficient nuclease of the instant disclosure can be lacking Orf2 of the Pong transposase. In some aspects, the Pong0RF1 protein comprises an amino acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 1. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 1. In some aspects, a nucleic acid sequence encoding the Pong ORF1 protein comprises about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 2. In some aspects, a nucleic acid sequence encoding the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 2.
[0149] When a transposase of the instant disclosure is a Pong or Pong-like transposase, the engineered system can also comprise a nuclease-deficient transposase comprising both ORF1 and a nuclease-deficient ORF2 proteins. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 1. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 1 . In some aspects, a nucleic acid sequence encoding the Pong ORF1 protein comprises about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 2. In some aspects, a nucleic acid sequence encoding the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 2. In some aspects, the nuclease-deficient Pong ORF2 protein comprises an amino acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%,98%, 99%, or 100% sequence identity with the amino sequence of SEQ ID NO: 121 . In some aspects, the nuclease-deficient Pong ORF2 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 121. In some aspects, a nucleic acid sequence encoding the nuclease-deficient Pong ORF2 protein comprises about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 123. In some aspects, a nucleic acid sequence encoding the nuclease-deficient Pong ORF2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 123.(b) Donor polynucleotide
[0150] Engineered systems of the disclosure also comprise a donor transposable nucleic acid construct (donor polynucleotide). In the presence of the nuclease-deficient transposase of the system and the programmable targeting nuclease of the engineered system of the instant disclosure, the donor polynucleotide is cut or copied from a nucleic acid sequence comprising the donor polynucleotide and targeted by the programmable targeting nuclease to a target nucleic acid locus to thereby mediate insertion of the donor polynucleotide into the target nucleic acid locus. A donor polynucleotide of the instant disclosure comprises transposition sequences compatible with the nuclease- deficient transposase of an engineered system of the instant disclosure. As used herein, the term “compatible” when referring to transposition sequences refers to nucleic acid sequences that can be recognized by a DNA-binding function of the transposase of the instant disclosure for transposition of the donor polynucleotide in the cell.
[0151] The donor polynucleotide can be comprised in a source polynucleotide from which the donor polynucleotide can be excised. A source polynucleotide can be a plasmid or vector comprising the donor polynucleotide. Alternatively, the source polynucleotide can be genomic DNA if the donor polynucleotide is integrated in the genomic DNA of the cell.
[0152] In some aspects, transposition sequences comprise a first transposition sequence at a first end of the donor polynucleotide, and a second transposition sequence at a second end of the donor polynucleotide.
[0153] In some aspects, the transposition sequences are derived from the TE from which the transposase is derived. However, the transposition sequences can also be derived from TEs other than the TE from which the transposases are derived, provided the transposition sequences are compatible with the transposon of the engineered system. Transposition sequences of the instant disclosure can be derived from autonomous or non-autonomous TEs. Non-autonomous TEs have short internal sequences devoid of open reading frames (ORF) that encode a defective transposase, or do not encode any transposase. Non-autonomous elements transpose through transposases encoded by autonomous TEs. The transposition sequences of the donor polynucleotide can each have about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with transposition sequences of the TE from which they are derived.
[0154] As explained in Section l(a) herein above, the nuclease-deficient transposase recognizes the transposition sequences and mediates the insertion of the donor polynucleotide at a target site recognized and cut by a programmable nuclease of the instant disclosure in the desired target locus. A donor polynucleotide can be an RNA polynucleotide or a DNA polynucleotide.
[0155] The donor polynucleotide can further comprise a cargo nucleic acid sequences of interest flanked by the transposition sequences, and insertion of the donor polynucleotide can result in the insertion of the cargo nucleic acid sequences of interest into the desired target locus. Non-limiting examples of cargo nucleic acid sequences that can be of interest for inserting in a target locus can be as described in Section IV herein below.
[0156] Further, insertion of the donor polynucleotide in a target locus can alter the function of the target locus. For instance, insertion of a donor polynucleotide in a nucleic acid sequence encoding a reporter can inactivate the reporter, thereby indicating a successful integration event. Conversely, excision of a donor polynucleotide from a source nucleic acid sequence encoding a reporter can re-activate the reporter, therebyindicating a successful excision event. Other uses of insertion and excision of a donor polynucleotide can be recognized by individuals of skill in the art.
[0157] In some aspects, the engineered system further comprises a reporter nucleic acid construct for expressing a reporter, wherein the reporter nucleic acid construct comprises a promoter operably linked to a polynucleotide sequence encoding the reporter, wherein the donor polynucleotide is inserted in the reporter nucleic acid construct thereby inactivating expression of the reporter, and wherein expression of the reporter is activated by excision of the inserted donor polynucleotide from the reporter nucleic acid construct by the transposase. The reporter can be a GFP reporter.
[0158] In some aspects, the transposase of the instant disclosure is derived from a PIF or P / F-like TE, and the transposition sequences compatible with the transposase are derived from a PIF or a P / F-like TE from which the transposase is derived, or can be derived from a tour / sMike miniature inverted-repeat transposable element (MITE). In some aspects, the transposase is derived from a Pong, a Pong-like, Ping, or a P / 7?g-like TE, and the transposition sequences compatible with the transposase can be derived from a stowaway-like MITE. In some aspects, the transposase is derived from a Pong, a Pong-like, a Ping, or a P / ng-like TE, and the transposition sequences compatible with the transposase are derived from an mPing or mP / ng-like MITE.
[0159] In some aspects, the transposition sequences are a first and second transposition sequences of a miniature inverted-repeat transposable element (MITE). In some aspects, the MITE is an mPing MITE. In some aspects, mPing comprises a nucleotide sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 96. In some aspects, mPing comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 96.
[0160] Importantly, it is noted that the inventors discovered that including mPing MITE first and second transposition sequences longer than the inverted repeats which was recognized by the art as being sufficient for transposition, significantly enhanced efficiency of transposition in a engineered system of the instant disclosure. Accordingly, transposition sequences of the instant disclosure can comprise the mPing invertedrepeat 1 and inverted repeat 2 and further comprise mPing sequences flanked (internal to) by the mPing inverted repeat 1 and inverted repeat 2. For instance, transposition sequences of the mPing MITE can comprise the mPing inverted repeat 1 , and further comprise any number of nucleotides of mPing downstream of inverted repeat 1 and any number of nucleotides of mPing downstream of inverted repeat 2.
[0161] In some aspects, transposition sequences of the mPing MITE comprise mPing inverted repeat 1 and inverted repeat 2. In some aspects, mPing inverted repeat 1 comprises a nucleotide sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 7. In some aspects, mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 7.
[0162] In some aspects, mPing inverted repeat 2 comprises a nucleotide sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 8. In some aspects, mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 8.
[0163] In some aspects, transposition sequences of the mPing MITE comprise the mPing inverted repeat 1 and inverted repeat 2 and further comprise mPing sequences flanked (internal to) by the mPing inverted repeat 1 and inverted repeat 2. In some aspects, transposition sequences of the instant disclosure comprise a first mPing transposition sequence comprising a nucleotide sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 111. In some aspects, transposition sequences of the instant disclosure comprise a first mPing transposition sequence comprising a nucleotide sequence comprising about 75% or more, at least about 85% or more, atleast about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 111.
[0164] In some aspects, transposition sequences of the instant disclosure comprise a second mPing transposition sequence comprising a nucleotide sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 112. In some aspects, transposition sequences of the instant disclosure comprise a second mPing transposition sequence comprising a nucleotide sequence comprising about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 112.
[0165] In some aspects, transposition sequences of the instant disclosure comprise a first mPing transposition sequence comprising a nucleotide sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 108. In some aspects, transposition sequences of the instant disclosure comprise a first mPing transposition sequence comprising a nucleotide sequence comprising about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 108.
[0166] In some aspects, transposition sequences of the instant disclosure comprise a second mPing transposition sequence comprising a nucleotide sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 109. In some aspects, transposition sequences of the instant disclosure comprise a second mPing transposition sequence comprising a nucleotide sequence comprising about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 109.
[0167] In some aspects, the donor polynucleotide comprises a cargo nucleic acid sequence comprising heat shock element (HSE) sequences flanked by mPing first and second transposition sequences. In some aspects, the donor polynucleotide comprisesa nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 69 to base 512 of SEQ ID NO: 81 or the nucleic acid sequence starting at base 69 to base 512 of SEQ ID NO: 93. In some aspects, the donor polynucleotide comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 69 to base 512 of SEQ ID NO: 81 or the nucleic acid sequence starting at base 69 to base 512 of SEQ ID NO: 93.
[0168] In some aspects, the donor polynucleotide comprises an expression construct for expressing a herbicide resistance function flanked by mPing first and second transposition sequences. In some aspects, the herbicide resistance function is resistance to bialaphos herbicide. In some aspects, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding resistance to bialaphos wherein the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97 or SEQ ID NO: 99.
[0148] In some aspects, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding resistance to bialaphos wherein the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97 or SEQ ID NO: 99.
[0169] In some aspects, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding resistance to bialaphos wherein the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97. In some aspects, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding resistance tobialaphos wherein the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97.
[0170] In some aspects, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding resistance to glyphosate wherein the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 124. In some aspects, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding resistance to glyphosate wherein the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 124.
[0171] The engineered system can further comprise a nucleic acid expression construct comprising a promoter operably linked to a polynucleotide sequence encoding a GFP reporter, wherein the donor polynucleotide is inserted in the nucleic acid expression construct. In some aspects, the nucleic acid expression construct comprises about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at nucleotide 2414 to nucleotide 23460 and nucleotide 1 to nucleotide 42 of SEQ ID NO: 74. In some aspects, the nucleic acid expression construct comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at nucleotide 2414 to nucleotide 23460 and nucleotide 1 to nucleotide 42 of SEQ ID NO: 74.(c) Programmable targeting system
[0172] The engineered system comprises a programmable targeting system. A programmable targeting system can be any single or group of components capable of targeting components of the engineered targeting system to a target nucleic acid locus, to introduce a cut in the target nucleic acid locus, or both, to thereby accomplishinsertion of the donor polynucleotide into the target locus. As explained herein above, in a system of the instant disclosure, the nuclease functions of the transposase that functions in excising a transposable element (or a donor polynucleotide) from a source polynucleotide and cutting a genomic sequence for insertion of the transposable element in a new location of the genome, are replaced by the nuclease function of a programmable targeting nuclease. Accordingly, a programmable targeting system of the instant disclosure can be any single or group of components capable of targeting components of the engineered targeting system to a donor polynucleotide, to excise the donor polynucleotide from the source polynucleotide, targeting components of the engineered targeting system to a target nucleic acid locus, and to introduce a cut in the target nucleic acid locus, to thereby accomplish insertion of the donor polynucleotide into the target locus.
[0173] The target nucleic acid locus can be in a coding or regulatory region of interest or can be in any other location in a nucleic acid sequence of interest. A gene can be a protein-coding gene, an RNA coding gene, or an intergenic region. The target nucleic acid locus can be in a nuclear, organellar, or extrachromosomal nucleic acid sequence. The cell can be a eukaryotic cell. In some aspects, the cell is a plant cell. In some aspects, the plant is a soybean plant.
[0174] A programmable targeting system generally comprises a programmable, sequence-specific nucleic acid-binding domain. In some aspects, the programmable targeting system further comprises a nuclease function. Non-limiting examples of programmable targeting systems include, without limit, an RNA-guided clustered regularly interspersed short palindromic repeats (CRISPR) / CRISPR- associated (Cas) (CRISPR / Cas) nuclease system, a CRISPR / Cpf1 nuclease system, a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, a ribozyme, or a programmable DNA binding domain that can be linked to a nuclease domain. Other suitable programmable targeting systems will be recognized by individuals skilled in the art.
[0175] In some aspects, the programmable targeting system is a programmable nucleic acid editing system. Such editing systems can be engineered to edit specific DNA or RNA sequences to repress transcription or translation of an mRNA encoded by the gene, and / or produce mutant proteins with reduced activity or stability. Non-limitingexamples of programmable targeting nucleases include, without limit, an RNA-guided clustered regularly interspersed short palindromic repeats (CRISPR) system, such as a CRISPR- associated (Cas) (CRISPR / Cas) nuclease system, a CRISPR / Cpf1 nuclease system, a zinc finger nuclease (ZFN) system, a transcription activator-like effector nuclease (TALEN) system, a MegaTAL, a homing endonuclease (HE), a meganuclease, a ribozyme, or a programmable DNA binding domain linked to a nuclease domain. Other suitable programmable targeting nucleases will be recognized by individuals skilled in the art. Such systems rely for specificity on the delivery of exogenous protein(s), and / or a guide RNA (gRNA) or single guide RNA (sgRNA) having a sequence which binds specifically to a target nucleic acid sequence of interest. When the programmable targeting nuclease comprises more than one component, such as a protein and a guide nucleic acid, the engineered system can be modular, in that the different components can optionally be distributed among two or more nucleic acid constructs as described herein. The components can be delivered by a plasmid or viral vector or as a synthetic oligonucleotide. In some aspects, the components can be delivered by one or more non-integrating nucleic acid constructs. More detailed descriptions of programmable nucleic acid editing systems can be as described in Section II further below.
[0176] The programmable nucleic acid-binding domain can be designed or engineered to recognize and bind different nucleic acid sequences. In some aspects, the nucleic acid-binding domain is mediated by interaction between a protein and the target nucleic acid sequence. Thus, the nucleic acid-binding domain can be programmed to bind a nucleic acid sequence of interest by protein engineering. Methods of programming a nucleic acid domain are well recognized in the art.
[0177] In other targeting systems, the nucleic acid-binding domain is mediated by a guide nucleic acid that interacts with a protein of the targeting system and the target nucleic acid sequence. In such instances, the programmable nucleic acid-binding domain can be targeted to a nucleic acid sequence of interest by designing the appropriate guide nucleic acid. Methods of designing guide nucleic acids are recognized in the art when provided with a target sequence using available tools that are capable of designing functional guide nucleic acids. It will be recognized that gRNA sequences and design of guide nucleic acids can and will vary at least depending on the particularprogrammable targeting system used. By way of non-limiting example, guide nucleic acids optimized by sequence for use with a Cas9 nuclease are likely to differ from guide nucleic acids optimized for use with a CPF1 nuclease, though it is also recognized that the target site location is a key factor in determining guide RNA sequences.
[0178] When a programmable targeting system comprises more than one component, such as a protein and a guide nucleic acid, the multi-component programmable targeting system can be modular, in that expression of the different components can optionally be distributed among two or more nucleic acid constructs as described herein.
[0179] In some aspects, the programmable targeting system is a CRISPR / Cas nuclease system comprising a nuclease protein and one or more guide RNA (gRNA). In some aspects, the targeting nuclease comprises an active nuclease domain. In other aspects, the nuclease activity of the targeting nuclease is altered to only nick or cut a single strand of the double stranded nucleic acid sequence. In some aspects, the programmable targeting system is a CRISPR / Cas nuclease system comprising a nuclease protein, one or more excision gRNAs, and a targeting gRNA. The one or more excision gRNAs guide the Cas9 nuclease to nucleic acid sequences in or flanking the transposition sequences of the donor polynucleotide to guide excision of the donor polynucleotide from the source polynucleotide and for expressing a targeting gRNA that guides the Cas9 nuclease to the target nucleic acid locus to introduce a cut at the target nucleic acid locus, thereby guiding insertion of the excised donor polynucleotide at the target nucleic acid locus to generate a genetically modified cell.
[0180] In some aspects, the Cas9 protein comprises an amino acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 5. In some aspects, the Cas9 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with amino acid sequence of SEQ ID NO: 5.
[0181] In some aspects, a nucleic acid sequence encoding the Cas9 protein comprises about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%sequence identity with the nucleic acid sequence of SEQ ID NO: 6. In some aspects, a nucleic acid sequence encoding the Cas9 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 6.
[0182] In some aspects, a nucleic acid sequence encoding the Cas9 nuclease is a deCas9 nickase, and a nucleic acid expression construct for expressing the deCas9 nickase comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 89. In some aspects, a nucleic acid sequence encoding the Cas9 nuclease is a deCas9 nickase, and a nucleic acid expression construct for expressing the deCas9 nickase comprises a nucleic acid sequence comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at nucleotide 8218 to nucleotide 13856 of SEQ ID NO: 89.
[0183] In some aspects, the gRNA comprises a nucleic acid sequence of SEQ IDNO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 80, SEQ ID NO: 113, SEQ ID NO: 67 and SEQ ID NO: 113, or any combination thereof.
[0184] In some aspects, the targeting nuclease is not linked to the transposase. In some aspects, the engineered system comprises one or more nucleic acid expression constructs for expressing a Pong ORF1 protein and a Pong ORF2 protein, and a nucleic acid expression construct for expressing a Cas9 nuclease protein. Pong ORF1 protein, Pong ORF2 protein can be as described in Section l(a) herein above, and expression constructs for expressing Pong ORF1 and ORF2 proteins can be as described in Section II herein below.
[0185] In other aspects, a transposase of the instant disclosure is linked to the programmable targeting nuclease. In some aspects, the engineered system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein and a nucleic acid expression construct for expressing a Pong ORF2 protein linked to Cas9 nuclease.
[0186] Multiple useful methods of linking proteins are known in the art and included herein. For instance, the targeting nuclease can be linked to the transposase by at least one peptide linker. Protein linkers aid fusion protein design by providingappropriate spacing between domains, supporting correct protein folding in the case that N or C termini interactions are crucial to folding. Commonly, protein linkers permit important domain interactions, reinforce stability, and reduce steric hindrance, making them preferred for use in fusion protein design even when N and C termini can be linked. Linkers can be flexible (e.g., comprising small, non-polar (e.g., Gly) or polar (e.g., Ser, Thr) amino acids). Rigid linkers can be formed of large, cyclic proline residues, which can be helpful when highly specific spacing between domains must be maintained. In vivo cleavable linkers are designed to allow the release of one or more linked domains under certain reaction conditions, such as a specific pH gradient, or when coming in contact with another biomolecule in the cell. Examples of suitable linkers are well known in the art, and programs to design linkers are readily available (Crasto et al., Protein Eng., 2000, 13(5):3096-312), the disclosure of which is incorporated herein in its entirety. Non-limiting examples of suitable linkers include GGSGGGSG (SEQ ID NO: 68), GSSSS (G4S; SEQ ID NO: 64) and (GGGGS)1-4 (SEQ ID NO: 69). One or more copies of this linker can be used sequentially to create longer linkers between the tethered proteins. In some aspects, the linker is three GSSSS (SEQ ID NO: 64) linkers used sequentially to create a longer linker. Alternatively, the linker can be rigid, such as AEAAAKEAAAKA (SEQ ID NO: 70), AEAAAKEAAAKEAAAKA (SEQ ID NO: 71), PAPAP (AP)6-8 (SEQ ID NO: 72), GIHGVPAA (SEQ ID NO: 73), EAAAK (SEQ ID NO: 76), EAAAKEAAAK (SEQ ID NO: 77), EAAAK EAAAK EAAAK (SEQ ID NO: 78), and EAAAKEAAAKEAAAKEAAAK (SEQ ID NO: 79). Other examples of suitable linkers are well known in the art, and programs to design linkers are readily available (Crasto et al., Protein Eng., 2000, 13(5):3096-312). In alternate aspects, the targeting nuclease and the transposase can be linked directly.
[0187] In some aspects, a transposase of the instant disclosure is linked to the programmable targeting nuclease by linking a Pong ORF2 protein to a Cas9 targeting nuclease. In some aspects, the Pong ORF2 protein is linked to a Cas9 targeting nuclease by one or more copies of a G4S linker. In some aspects, the Pong ORF2 protein is linked to a Cas9 targeting nuclease by one copy of a G4S linker. In some aspects, the Pong ORF2 protein linked to a Cas9 targeting nuclease by one copy of a G4S linker comprises an amino acid sequence encoded by a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%,87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 106. In some aspects, the Pong ORF2 protein linked to a Cas9 targeting nuclease by one copy of a G4S linker comprises an amino acid sequence encoded by a nucleic acid sequence comprising about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 106.
[0188] In some aspects, the Pong ORF2 protein is linked to a Cas9 targeting nuclease by three copies of a G4S linker. In some aspects, the Pong ORF2 protein linked to a Cas9 targeting nuclease by three copies of a G4S linker comprises an amino acid sequence encoded by a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 107. In some aspects, the Pong ORF2 protein linked to a Cas9 targeting nuclease by three copies of a G4S linker comprises an amino acid sequence encoded by a nucleic acid sequence comprising about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 107. / . CRISPR nuclease systems.
[0189] The programmable targeting nuclease can be an RNA-guided CRISPR endonuclease system. The CRISPR system comprises a guide RNA or sgRNA to a target sequence at which a protein of the system introduces a double-stranded break in a target nucleic acid sequence, and a CRISPR-associated endonuclease. The gRNA is a short synthetic RNA comprising a sequence necessary for endonuclease binding, and a preselected ~20 nucleotide spacer sequence targeting the sequence of interest in a genomic target. Non-limiting examples of endonucleases include Cas1 , Cas1 B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas12, Cas100, Csy1 , Csy2, Csy3, Cse1 , Cse2, Csc1 , Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1 , Cmr3, Cmr4, Cmr5, Cmr6, Csb1 , Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1 , Csx15, Csf1 , Csf2, Csf3, Csf4, or Cpf1 endonuclease, or a homolog thereof, a recombination of the naturally occurringmolecule thereof, a codon-optimized version thereof, or a modified version thereof, or any combination thereof.
[0190] The CRISPR nuclease system can be derived from any type of CRISPR system, including a type I (i.e. , IA, IB, IC, ID, IE, or IF), type II (i.e. , HA, IIB, or IIC), type III (i.e., I HA or 11 IB), or type V CRISPR system. The CRISPR / Cas system can be from Streptococcus sp. (e.g., Streptococcus pyogenes), Campylobacter sp. (e.g., Campylobacter jejuni), Francisella sp. (e.g., Francisella novicida), Acaryochloris sp., Acetohalobium sp., Acida mi nococcus sp., Acidithiobacillus sp., Alicyclobacillus sp., Allochromatium sp., Ammonifex sp., Anabaena sp., Arthrospira sp., Bacillus sp., Burkholderiales sp., Caldicelulosiruptor sp., Candidatus sp., Clostridium sp., Crocosphaera sp., Cyanothece sp., Exiguobacterium sp., Finegoldia sp., Ktedonobacter sp., Lactobacillus sp., Lyngbya sp., Marinobacter sp., Methanohalobium sp., Microscilla sp., Microcoleus sp., Microcystis sp., Natranaerobius sp., Neisseria sp., Nitrosococcus sp., Nocardiopsis sp., Nodularia sp., Nostoc sp., Oscillatoria sp., Polaromonas sp., Pelotomaculum sp., Pseudoalteromonas sp., Petrotoga sp., Prevotella sp., Staphylococcus sp., Streptomyces sp., Streptosporangium sp., Synechococcus sp., or Thermosipho sp.
[0191] Non-limiting examples of suitable CRISPR systems include CRISPR / Cas systems, CRISPR / Cpf systems, CRISPR / Cmr systems, CRISPR / Csa systems, CRISPR / Csb systems, CRISPR / Csc systems, CRISPR / Cse systems, CRISPR / Csf systems, CRISPR / Csm systems, CRISPR / Csn systems, CRISPR / Csx systems, CRISPR / Csy systems, CRISPR / Csz systems, and derivatives or variants thereof. Preferably, the CRISPR system can be a type II Cas9 protein, a type V Cpf1 protein, or a derivative thereof. In some aspects, the CRISPR / Cas nuclease is Streptococcus pyogenes Cas9 (SpCas9), Streptococcus thermophilus Cas9 (StCas9), Campylobacter jejuni Cas9 (CjCas9), Francisella novicida Cas9 (FnCas9), or Francisella novicida Cpf1 (FnCpfl).
[0192] In general, a protein of the CRISPR system comprises a RNA recognition and / or RNA binding domain, which interacts with the guide RNA. A protein of the CRISPR system also comprises at least one nuclease domain having endonuclease activity. For example, a Cas9 protein can comprise a RuvC-like nuclease domain and an HNH-like nuclease domain, and a Cpf1 protein can comprise a RuvC-like domain. Aprotein of the CRISPR system can also comprise DNA binding domains, helicase domains, RNase domains, protein-protein interaction domains, dimerization domains, as well as other domains.
[0193] A protein of the CRISPR system can be associated with guide RNAs (gRNA). The guide RNA can be a single guide RNA (i.e. , sgRNA), or can comprise two RNA molecules (i.e., crRNA and tracrRNA). The guide RNA interacts with a protein of the CRISPR system to guide it to a target site in the DNA. The target site has no sequence limitation except that the sequence is bordered by a protospacer adjacent motif (PAM). For example, PAM sequences for Cas9 include 3'-NGG, 3'-NGGNG, 3'- NNAGAAW, and 3'-ACAY, and PAM sequences for Cpf1 include 5'-TTN (wherein N is defined as any nucleotide, W is defined as either A or T, and Y is defined as either C or T). Each gRNA comprises a sequence that is complementary to the target sequence (e.g., a Cas9 gRNA can comprise GN17-20GG). The gRNA can also comprise a scaffold sequence that forms a stem loop structure and a single-stranded region. The scaffold region can be the same in every gRNA. In some aspects, the gRNA can be a single molecule (i.e., sgRNA). In other aspects, the gRNA can be two separate molecules. Those skilled in the art are familiar with gRNA design and construction, e.g., gRNA design tools are available on the internet or from commercial sources.
[0194] A CRISPR system can comprise one or more nucleic acid binding domains associated with one or more, or two or more selected guide RNAs used to direct the CRISPR system to one or more, or two or more selected target nucleic acid loci. For instance, a nucleic acid binding domain can be associated with one or more, or two or more selected guide RNAs, each selected guide RNA, when complexed with a nucleic acid binding domain, causing the CRISPR system to localize to the target of the guide RNA.
[0195] A nuclease of a CRISPR nuclease system can be inactivated to obtain a programmable targeting protein. For instance, a CRISPR / Cas system can comprise a nuclease-deficient dead CAS9 protein (dCAS9) and a guide RNA (gRNA). / / . CRISPR nickase systems.
[0196] The programmable targeting nuclease can also be a CRISPR nickase system. CRISPR nickase systems are similar to the CRISPR nuclease systems described above except that a CRISPR nuclease of the system is modified to cleaveonly one strand of a double-stranded nucleic acid sequence. Thus, a CRISPR nickase, in combination with a guide RNA of the system, can create a single-stranded break or nick in the target nucleic acid sequence. Alternatively, a CRISPR nickase in combination with a pair of offset gRNAs can create a double-stranded break in the nucleic acid sequence.
[0197] A CRISPR nuclease of the system can be converted to a nickase by one or more mutations and / or deletions. For example, a Cas9 nickase can comprise one or more mutations in one of the nuclease domains, wherein the one or more mutations can be D10A, E762A, and / or D986A in the RuvC-like domain, or the one or more mutations can be H840A (or H839A), N854A and / or N863A in the HNH-like domain.Hi. ssDNA-guided Argonaute systems.
[0198] Alternatively, the programmable targeting nuclease can comprise a singlestranded DNA-guided Argonaute endonuclease. Argonautes (Agos) are a family of endonucleases that use 5'-phosphorylated short single-stranded nucleic acids as guides to cleave nucleic acid targets. Some prokaryotic Agos use single-stranded guide DNAs and create double-stranded breaks in nucleic acid sequences. The ssDNA-guided Ago endonuclease can be associated with a single-stranded guide DNA.
[0199] The Ago endonuclease can be derived from Alistipes sp., Aquifex sp., Archaeoglobus sp., Bacteriodes sp., Bradyrhizobium sp., Burkholderia sp., Cellvibrio sp., Chlorobium sp., Geobacter sp., Mariprofundus sp., Natronobacterium sp., Parabacteriodes sp., Parvularcula sp., Planctomyces sp., Pseudomonas sp., Pyrococcus sp., Thermus sp., or Xanthomonas sp. For instance, the Ago endonuclease can be Natronobacterium gregoryi Ago (NgAgo). Alternatively, the Ago endonuclease can be Thermus thermophilus Ago (TtAgo). The Ago endonuclease can also be Pyrococcus furiosus (PfAgo).
[0200] The single-stranded guide DNA (gDNA) of an ssDNA-guided Argonaute system is complementary to the target site in the nucleic acid sequence. The target site has no sequence limitations and does not require a PAM. The gDNA generally ranges in length from about 15-30 nucleotides. The gDNA can comprise a 5' phosphate group. Those skilled in the art are familiar with ssDNA oligonucleotide design and construction.iv. Zinc finger nucleases.
[0201] The programmable targeting nuclease can be a zinc finger nuclease (ZFN). A ZFN comprises a DNA-binding zinc finger region and a nuclease domain. The zinc finger region can comprise from about two to seven zinc fingers, for example, about four to six zinc fingers, wherein each zinc finger binds three nucleotides. The zinc finger region can be engineered to recognize and bind to any DNA sequence. Zinc finger design tools or algorithms are available on the internet or from commercial sources. The zinc fingers can be linked together using suitable linker sequences.
[0202] A ZFN also comprises a nuclease domain, which can be obtained from any endonuclease or exonuclease. Non-limiting examples of endonucleases from which a nuclease domain can be derived include, but are not limited to, restriction endonucleases and homing endonucleases. The nuclease domain can be derived from a type ll-S restriction endonuclease. Type ll-S endonucleases cleave DNA at sites that are typically several base pairs away from the recognition / binding site and, as such, have separable binding and cleavage domains. These enzymes generally are monomers that transiently associate to form dimers to cleave each strand of DNA at staggered locations. Non-limiting examples of suitable type ll-S endonucleases include Bfil, Bpml, Bsal, Bsgl, BsmBI, Bsml, BspMI, Fokl, Mboll, and Sapl. The type ll-S nuclease domain can be modified to facilitate dimerization of two different nuclease domains. For example, the cleavage domain of Fokl can be modified by mutating certain amino acid residues. By way of non-limiting example, amino acid residues at positions 446, 447, 479, 483, 484, 486, 487, 490, 491 , 496, 498, 499, 500, 531 , 534, 537, and 538 of Fokl nuclease domains are targets for modification. For example, one modified Fokl domain can comprise Q486E, I499L, and / or N496D mutations, and the other modified Fokl domain can comprise E490K, I538K, and / or H537R mutations. v. Transcription activator-like effector nuclease systems.
[0203] The programmable targeting nuclease can also be a transcription activator-like effector nuclease (TALEN) or the like. TALENs comprise a DNA-binding domain composed of highly conserved repeats derived from transcription activator-like effectors (TALEs) that are linked to a nuclease domain. TALEs are proteins secreted by plant pathogen Xanthomonas to alter transcription of genes in host plant cells. TALErepeat arrays can be engineered via modular protein design to target any DNA sequence of interest. Other transcription activator-like effector nuclease systems can comprise, but are not limited to, the repetitive sequence, transcription activator like effector (RipTAL) system from the bacterial plant pathogenic Ralstonia solanacearum species complex (Rssc). The nuclease domain of TALEs can be any nuclease domain as described above in Section (l)(c)(i). vi. Meganucleases or rare-cutting endonuclease systems.
[0204] The programmable targeting nuclease can also be a meganuclease or derivative thereof. Meganucleases are endodeoxyribonucleases characterized by long recognition sequences, i.e. , the recognition sequence generally ranges from about 12 base pairs to about 45 base pairs. As a consequence of this requirement, the recognition sequence generally occurs only once in any given genome. Among meganucleases, the family of homing endonucleases named LAGLIDADG has become a valuable tool for the study of genomes and genome engineering. Non-limiting examples of meganucleases that can be suitable for the instant disclosure include I- Scel, l-Crel , l-Dmol, or variants and combinations thereof. A meganuclease can be targeted to a specific nucleic acid sequence by modifying its recognition sequence using techniques well known to those skilled in the art.
[0205] The programmable targeting nuclease can be a rare-cutting endonuclease or derivative thereof. Rare-cutting endonucleases are site-specific endonucleases whose recognition sequence occurs rarely in a genome, such as only once in a genome. The rare-cutting endonuclease can recognize a 7-nucleotide sequence, an 8- nucleotide sequence, or longer recognition sequence. Non-limiting examples of rare- cutting endonucleases include Notl, Asci, Pad, AsiSI, Sbfl, and Fsel. vii. Optional additional domains.
[0206] The programmable targeting nuclease can further comprise at least one nuclear localization signal (NLS), at least one cell-penetrating domain, at least one reporter domain, and / or at least one linker.
[0207] In general, an NLS comprises a stretch of basic amino acids. Nuclear localization signals are known in the art (see, e.g., Lange et al., J. Biol. Chem., 2007,282:5101-5105). The NLS can be located at the N-terminus, the C-terminal, or in an internal location of the fusion protein.
[0208] A cell-penetrating domain can be a cell-penetrating peptide sequence derived from the HIV-1 TAT protein. The cell-penetrating domain can be located at the N-terminus, the C-terminal, or in an internal location of the fusion protein.
[0209] A programmable targeting nuclease can further comprise at least one linker. For example, the programmable targeting nuclease, the nuclease domain of the targeting nuclease, and other optional domains can be linked via one or more linkers. The linker can be flexible (e.g., comprising small, non-polar (e.g., Gly) or polar (e.g., Ser, Thr) amino acids). Examples of suitable linkers are well known in the art, and programs to design linkers are readily available (Crasto et al., Protein Eng., 2000, 13(5):3096-312). In alternate aspects, the programmable targeting nuclease, the cell cycle regulated protein, and other optional domains can be linked directly.
[0210] A programmable targeting nuclease can further comprise an organelle localization or targeting signal that directs a molecule to a specific organelle. A signal can be polynucleotide or polypeptide signal, or can be an organic or inorganic compound sufficient to direct an attached molecule to a desired organelle. Organelle localization signals can be as described in U.S. Patent Publication No. 20070196334, the disclosure of which is incorporated herein in its entirety.(d) Engineered system
[0211] An engineered system of the instant disclosure generally comprises a nucleic acid expression construct for expressing a transposase, wherein the expression construct comprises a promoter operably linked to a nucleic acid sequence encoding a transposase. The engineered system also comprises a donor polynucleotide comprising nucleic acid transposition sequences compatible with the transposase and a nucleic acid expression construct for expressing a programmable targeting system, wherein the expression construct comprises a promoter operably linked to a nucleic acid sequence encoding a programmable targeting system. The programmable targeting system is programmed to target the transposase and the donor polynucleotide to a target nucleic acid locus in the cell, thereby accomplishing insertion of the donor polynucleotide at thetarget nucleic acid locus to generate a genetically modified cell comprising the donor polynucleotide inserted at the target nucleic acid locus.
[0212] In some aspects, the targeting system comprises a targeting nuclease and is engineered to introduce a cut in a target nucleic acid locus. In other aspects, the targeting system does not comprise a nuclease function. The transposase can be linked to the targeting system. Alternatively, the transposase is not linked to the targeting nuclease.
[0213] The system can further comprise a nucleic acid expression construct comprising a promoter operably linked to a polynucleotide sequence encoding a reporter, wherein the donor polynucleotide is inserted in the nucleic acid expression construct, wherein the reporter is inactivated by the inserted nucleic acid construct comprising the donor polynucleotide, and wherein the reporter is activated by excision of the inserted nucleic acid construct comprising the donor polynucleotide from the expression construct comprising a promoter operably linked to a polynucleotide sequence encoding a reporter by the transposase. In some aspects, the reporter can be GFP, and the GFP expression construct, wherein the donor polynucleotide is inserted in the nucleic acid expression construct, comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at nucleotide 2414 to nucleotide 23460 and nucleotide 1 to nucleotide 42 of SEQ ID NO: 74. In some aspects, the reporter can be GFP, and the GFP expression construct, wherein the donor polynucleotide is inserted in the nucleic acid expression construct, comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at nucleotide 2414 to nucleotide 23460 and nucleotide 1 to nucleotide 42 of SEQ ID NO: 74.
[0214] The transposase can be a split transposase. When the transposase is a split transposase, the transposase can be a Pong or Pong-like transposase comprising a Pong ORF1 protein and a Pong ORF2 protein. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%,94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 1. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 1. A nucleic acid sequence encoding the Pong ORF1 protein can comprise about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 2. A nucleic acid sequence encoding the Pong ORF1 protein can comprise at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 2.
[0215] In some aspects, the Pong ORF2 protein comprises an amino acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 3. In some aspects, the Pong ORF2 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 3. A nucleic acid sequence encoding the Pong ORF2 protein can comprise about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 4. A nucleic acid sequence encoding the Pong ORF2 protein can comprise at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 4.
[0216] The transposition sequences can be transposition sequences of a miniature inverted-repeat transposable element (MITE). In some aspects, the MITE is an mPing MITE or a derivative of mPing with sequences added or removed. In some aspects, transposition sequences of the mPing MITE comprise mPing inverted repeat 1 and inverted repeat 2. In some aspects, mPing inverted repeat 1 comprises a nucleotide sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%,or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 7, SEQ ID NO: 111 , or SEQ ID NO: 108 . In some aspects, mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 7, SEQ ID NO: 111 , or SEQ ID NO: 108 . In some aspects, mPing inverted repeat 2 comprises a nucleotide sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 8, SEQ ID NO: 112, or SEQ ID NO: 109. In some aspects, mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 8, SEQ ID NO: 112, or SEQ ID NO: 109.
[0217] The system comprises an expression construct for expressing the Pong ORF1 protein, wherein the expression construct for expressing the Pong ORF1 protein can comprise at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 100. In some aspects, the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100.
[0218] The programmable targeting system can be a CRISPR / Cas system comprising a Cas9 nuclease and a guide RNA (gRNA). In some aspects, the Cas9 nuclease comprises an amino acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 5. In some aspects, the Cas9 nuclease comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 5.
[0219] In some aspects, the Cas9 nuclease is encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%,84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 6. In some aspects, the Cas9 nuclease is encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 6. In some aspects, the gRNA comprises a nucleic acid sequence of SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 80, SEQ ID NO: 113, SEQ ID NO: 67 and SEQ ID NO: 113, or any combination thereof.
[0220] The transposase can be linked to the Cas9 nuclease. When the transposase is linked to the Cas9 nuclease, an engineered system of the instant disclosure comprises a Pong ORF2 protein is linked to the Cas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64. In some aspects, the Pong ORF2 protein linked to the Cas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64 comprises an amino acid sequence encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 106 or a nucleic acid sequence starting at base 8392 to base 14052 of SEQ ID NO: 74. In some aspects, the Pong ORF2 protein linked to the Cas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64 comprises an amino acid sequence encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 106 or a nucleic acid sequence starting at base 8392 to base 14052 of SEQ ID NO: 74.
[0221] In some aspects, the engineered system comprises an expression construct for expressing the Pong ORF2 protein linked to the Cas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64, wherein the expression construct comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence starting at base 7451 to base 15799 of SEQ ID NO: 74. In some aspects, the engineered system comprises an expression construct for expressing the Pong ORF2 protein linked to the Cas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64, wherein theexpression construct comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence starting at base 7451 to base 15799 of SEQ ID NO: 74. In some aspects, the cell is an Arabidopsis thaliana cell.
[0222] In some aspects, the programmable targeting system of the instant disclosure comprises a CRISPR nuclease system comprising dCas9 and a gRNA. In some aspects, the dCas9 nuclease is linked to Pong ORF2 by one copy of a G4S linker of SEQ ID NO: 64. In some aspects, the Pong ORF2 protein linked to the dCas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64 comprises an amino acid sequence encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 110. In some aspects, the Pong ORF2 protein linked to the dCas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64 comprises an amino acid sequence encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 110.
[0223] In some aspects, the Pong ORF2 protein linked to the dCas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64 is expressed using an expression construct for expressing the Pong ORF2 protein linked to the dCas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64, wherein the expression construct comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 115. In some aspects, the expression construct comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 115. In some aspects, the genetically modified cell is an Arabidopsis thaliana cell.
[0224] In some aspects, the dCas9 nuclease is linked to Pong ORF2 by three copies of a G4S linker of SEQ ID NO: 64. In some aspects, the Pong ORF2 protein linked to the dCas9 nuclease by three copies of a G4S linker of SEQ ID NO: 64 comprises an amino acid sequence encoded by a nucleic acid sequence comprising atleast about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 107. In some aspects, the Pong ORF2 protein linked to the dCas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64 comprises an amino acid sequence encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 107.
[0225] In some aspects, the Pong ORF2 protein linked to the Cas9 nuclease by three copies of a G4S linker of SEQ ID NO: 64 is expressed using an expression construct for expressing the Pong ORF2 protein linked to the Cas9 nuclease by three copies of a G4S linker of SEQ ID NO: 64, wherein the expression construct comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 104. In some aspects, the expression construct comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 104. In some aspects, the genetically modified cell is a soybean cell.
[0226] In some aspects, the Pong ORF2 protein is not linked to the targeting nuclease. When the Pong ORF2 protein is not linked to the targeting nuclease, the engineered system can comprise a nucleic acid expression construct for expressing a Cas9 nuclease, wherein the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 92 or a nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94. In some aspects, the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 92 or a nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94.
[0227] When the Pong ORF2 protein is not linked to the targeting nuclease, the engineered system can comprise a nucleic acid expression construct for expressing a Pong ORF2 protein, wherein the expression construct for expressing the Pong ORF2 protein comprises a nuclueic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO 101 or a nucleic acid sequence starting at base 5073 to base 8215 of SEQ ID NO: 89. In some aspects, the expression construct for expressing the Pong ORF2 protein comprises a nuclueic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO 101 or a nucleic acid sequence starting at base 5073 to base 8215 of SEQ ID NO: 89.
[0228] The first mPing transposition sequence and the second mPing transposition sequence can flank a cargo polynucleotide. In some aspects, the cargo polynucleotide comprises HSEs. When the cargo polynucleotide comprises HSEs, the first mPing transposition sequence can comprise at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 7 and the second mPing transposition sequence can comprise at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 8. In some aspects, the first mPing transposition sequence comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 7 and wherein the second mPing transposition sequence comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 8. In some aspects, the donor polynucleotide comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 81 . In some aspects, the donor polynucleotide comprises at least about 75% or more, at least about 85% ormore, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 81 .
[0229] In some aspects, the cargo polynucleotide comprises an expression construct for expressing a herbicide resistance function. The herbicide resistance function can be resistance to bialaphos herbicide. When the herbicide resistance function can be resistance to bialaphos herbicide, the first mPing transposition sequence can comprise a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 108 and the second mPing transposition sequence can comprise a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 109. In some aspects, the first mPing transposition sequence comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 108 and the second mPing transposition sequence comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 109.
[0230] In some aspects, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding a bialaphos resistance gene wherein the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97 or SEQ ID NO: 99. In some aspects, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding a bialaphos resistance gene wherein the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97 or SEQ ID NO: 99.
[0231] In some aspects, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding a bialaphos resistance gene wherein the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97. In some aspects, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding a bialaphos resistance gene wherein the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97.
[0232] In some aspects, the engineered system comprises an expression construct for expressing a gRNA for targeting the transposase and nuclease to a target nucleic acid locus in an Arabidopsis thaliana PDS3 gene, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 2632 to base 3343 of SEQ ID NO: 74. In some aspects, the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 2632 to base 3343 of SEQ ID NO: 74.
[0233] In some aspects, the engineered system comprises an expression construct for expressing a gRNA for targeting the transposase and nuclease to a target nucleic acid locus in an Arabidopsis thaliana ADH1 gene, wherein the expression construct for expressing a gRNA comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 254 to base 965 of SEQ ID NO: 89. In some aspects, the expression construct for expressing a gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% ormore, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 254 to base 965 of SEQ ID NO: 89.
[0234] In some aspects, the engineered system comprises an expression construct for expressing a gRNA for targeting the transposase and nuclease to a target nucleic acid locus in an Arabidopsis thaliana ACT8 gene, wherein the expression construct for expressing a gRNA comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 103 or the nucleic acid sequence starting at base 729 to base 1440 of SEQ ID NO: 92. In some aspects, the expression construct for expressing a gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 103 or the nucleic acid sequence starting at base 729 to base 1440 of SEQ ID NO: 92.
[0235] In some aspects, the engineered system comprises an expression construct for expressing a gRNA for targeting the transposase and nuclease to a target nucleic acid locus in a soybean DD20 intergenic region, wherein the expression construct for expressing a gRNA comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 105. In some aspects, the expression construct for expressing a gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 105.
[0236] Another aspect of the instant disclosure encompasses an engineered system for generating a genetically modified cell, wherein the engineered system comprises
[0237] In some aspects, the system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequenceidentity with the nucleic acid sequence of SEQ ID NO: 100; a nucleic acid expression construct for expressing a Pong ORF2 protein linked to Cas9 nuclease with one copy of a G4S linker, wherein the expression construct for expressing the Pong ORF2 protein linked to Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 7451 to base 14807 of SEQ ID NO: 74; a donor polynucleotide comprising first and second mPing transposition sequences; and an expression construct for expressing a gRNA, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 103. In some aspects, the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100. In some aspects, the expression construct for expressing the Pong ORF2 protein linked to Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 7451 to base 14807 of SEQ ID NO: 74. In some aspects, the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 103.
[0238] In some aspects, the donor polynucleotide comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 81. In some aspects, the donor polynucleotide comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 81 .
[0239] In some aspects, the system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct forexpressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100; a nucleic acid expression construct for expressing a Pong ORF2 protein, wherein the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 101 ; a nucleic acid expression construct for expressing a Cas9 nuclease, wherein the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102; a donor polynucleotide comprising first and second mPing transposition sequences; and an expression construct for expressing a gRNA, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 103. In some aspects, the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100. In some aspects, the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 101. In some aspects, the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102. In some aspects, the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at leastabout 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 103.
[0240] The donor polynucleotide comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 81 . In some aspects, the donor polynucleotide comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 81.
[0241] In some aspects, the engineered system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100; a nucleic acid expression construct for expressing a Pong ORF2 protein, wherein the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 101 ; a nucleic acid expression construct for expressing a Cas9 nuclease, wherein the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102; a donor polynucleotide comprising first and second mPing transposition sequences; and an expression construct for expressing a gRNA, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 103. In some aspects, the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95%or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100. In some aspects, the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 101. In some aspects, the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102. In some aspects, the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 103.
[0242] The donor polynucleotide comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 81 . In some aspects, the donor polynucleotide comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 81.
[0243] In some aspects, the system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100; a nucleic acid expression construct for expressing a Pong ORF2 protein, wherein the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 101 ; a nucleic acid expression construct for expressing a Cas9 nuclease, wherein the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%,88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102; a donor polynucleotide comprising first and second mPing transposition sequences; and an expression construct for expressing a gRNA, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 105. In some aspects, the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100. In some aspects, the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 101. In some aspects, the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102. In some aspects, the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 105.
[0244] In some aspects, the donor polynucleotide comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 81. In some aspects, the donor polynucleotide comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 81 .
[0245] In some aspects, the engineered system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%,85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100; a nucleic acid expression construct for expressing a Pong ORF2 protein, wherein the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 101 ; a nucleic acid expression construct for expressing a Cas9 nuclease, wherein the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102; a donor polynucleotide comprising first and second mPing transposition sequences; and an expression construct for expressing a gRNA of SEQ ID NO: 67 and a gRNA of SEQ ID NO: 113, wherein the expression construct comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 114. In some aspects, the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100. In some aspects, the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 101 . In some aspects, the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102. In some aspects, the expression construct comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 114.
[0246] In some aspects, the system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100; a nucleic acid expression construct for expressing a Pong ORF2 protein linked to dCas9 nuclease with one copy of a G4S linker, wherein the expression construct for expressing the Pong ORF2 protein linked to Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 115; a donor polynucleotide comprising first and second mPing transposition sequences; and an expression construct for expressing a gRNA of SEQ ID NO: 67 and a gRNA of SEQ ID NO: 113, wherein the expression construct comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 114. In some aspects, the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100. In some aspects, the expression construct for expressing the Pong ORF2 protein linked to Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 115. In some aspects, the expression construct comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 114.
[0247] As explained in Section II further below, a system of the instant disclosure can be encoded on one or more nucleic acid constructs encoding the components of the system. Depending on an intended use of the system of the instant disclosure, the nucleic acid constructs encoding the components of the system can be present orcloned into one or more vectors, such as plasmid vectors, based on intended use. For instance, the systems can be a single-component system comprising all the nucleic acid constructs encoding the components of the system cloned into one vector. Such a system can provide the convenience and simplicity of introducing a single nucleic acid construct into a cell.
[0248] In some aspects, an engineered system of the instant disclosure comprises a Pong transposase, wherein the nucleic acid transposition sequences are mPing inverted repeat 1 and inverted repeat 2, and the programmable targeting nuclease comprises a Cas9 nuclease and a gRNA. In some aspects, the Pong ORF2 protein is linked to the Cas9 nuclease. In some aspects, the Pong ORF2 protein is not linked to the Cas9 nuclease.
[0249] In some aspects, an engineered system of the instant disclosure comprises a donor polynucleotide comprising a first and second mPing miniature inverted-repeat transposable element (MITE) transposition sequences; one or more nucleic acid expression constructs for expressing a tranposase comprising a Pong ORF1 protein and a Pong ORF2 protein, wherein each of the one or more expression constructs comprises a promoter operably linked to a nucleic acid sequence encoding the Pong ORF1 protein and the Pong ORF2 protein; and a nucleic acid expression construct for expressing a programmable targeting system, wherein the expression construct comprises a promoter operably linked to a nucleic acid sequence encoding the programmable targeting system. The programmable targeting system is programmed to target the transposase and the donor polynucleotide to a target nucleic acid locus in the cell, to introduce a cut in the target nucleic acid locus, or both, thereby accomplishing insertion of the donor polynucleotide at the target nucleic acid locus to generate a genetically modified cell comprising the donor polynucleotide inserted at the target nucleic acid locus.
[0250] In some aspects, the system further comprises a reporter nucleic acid construct for expressing a reporter, wherein the reporter nucleic acid construct comprises a promoter operably linked to a polynucleotide sequence encoding the reporter, wherein the donor polynucleotide is inserted in the reporter nucleic acid construct thereby inactivating expression of the reporter, and wherein expression of the reporter is activated by excision of the inserted donor polynucleotide from the reporternucleic acid construct by the transposase. In some aspects, the reporter is GFP, and the nucleic acid expression construct comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at nucleotide 2414 to nucleotide 23460 and nucleotide 1 to nucleotide 42 of SEQ ID NO: 74. In some aspects, the reporter is GFP, and the nucleic acid expression construct comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at nucleotide 2414 to nucleotide 23460 and nucleotide 1 to nucleotide 42 of SEQ ID NO: 74.
[0251] A system of the instant disclosure can be encoded on more than one nucleic acid construct. In some aspects, a system of the instant disclosure comprises a two-component system comprising a donor nucleic acid construct comprising the nucleic acid construct comprising a donor polynucleotide of the instant disclosure, and a helper nucleic acid construct comprising a nucleic acid expression construct for expressing a transposase and the nucleic acid expression construct for expressing the programmable targeting nuclease of the instant disclosure. Two-component and one- components systems can be as described in Section ll(a) herein below.(e) Engineered system for generating a genetically modified cell comprising no off-target insertions
[0252] An engineered nucleic acid modification system for generation of a genetically modified cell comprising no or reduced off-target insertions. A biproduct of a transposable element system can be the generation of unintended transposition events, i.e. , off target transposition events catalyzed by the nuclease catalytic activity of the transposase. Further, the transposase activity could activate endogenous transposable elements that it may recognize. This is particularly likely if the transposase of the system is derived from a transposable element already present in the cell being genetically modified. For instance, if the transposase of a TATS I system being used to generate a genetically modified rice plant is derived from a rice Pong / mPing, there is arisk that the transposase will recognize the endogenous mPing elements in the rice genome and create off-target events and potentially genome instability.
[0253] To avoid potential off-target transposition events catalyzed by the nuclease activity of the transposase and prevent the activation of endogenous transposable elements, the inventors generated a TATSI system comprising a catalytically deficient (nuclease-deficient) transposase. In this system, the nuclease activity of the transposase of the system is replaced by the nuclease activity of the programmable targeting nuclease, referred to herein as the TAHITI system. Put differently, the transposase of the TAHITI system comprises a DNA-binding protein comprising the DNA-binding domain of the transposase responsible for binding the DNA transposition sequences of a donor of the TAHITI system. The DNA binding protein is nuclease deficient. Importantly, the inventors discovered that the TAHITI system generated higher rates of targeted insertion, and significantly fewer or even no off-target insertions. Further, the TAHITI system generates precise junction sequences at insertion sites.
[0254] Accordingly, in some aspects, an engineered system of the instant disclosure comprises a nuclease-deficient transposase and a programmable targeting nuclease or one or more nucleic acid constructs for expressing the nuclease-deficient transposase and / or the programmable targeting nuclease. An engineered system of the instant disclosure further comprises a donor polynucleotide comprising transposition sequences compatible with the nuclease-deficient transposase. In some aspects, an engineered system of the instant disclosure comprises a DNA-binding protein comprising the DNA-binding domain of the transposase, a programmable targeting nuclease or one or more nucleic acid constructs for expressing the DNA-binding protein and / or a programmable targeting nuclease. The engineered system further comprises a donor polynucleotide comprising transposition sequences compatible with the nuclease- deficient transposase. Transposable nucleic acid constructs can be as described in Sections 1(b) and 1(d) herein above.
[0255] The engineered system comprises a nuclease-deficient transposase or an expression construct for expressing the nuclease-deficient transposase. A nuclease- deficient transposase can be as described in Section 1(a) herein above. In some aspects, a nuclease-deficient transposase is a transposase comprising an inactivatednuclease domain. In some aspects, when the transposase is a Pong transposase, a nuclease-free transposase is a Pong transposase comprising an inactivating mutation in the nuclease domain of Pong ORF2. In some aspects, when the transposase is a Pong transposase, a nuclease-free transposase is a Pong transposase comprising an inactivating mutation in the DDE nuclease motif (Aspartate-Aspartate-Glutamate).
[0256] In some aspects, a nuclease-free transposase is a split transposase wherein the ORF comprising the nuclease domain is omitted. In some aspects, the split transposase is Pong transposase, and the nuclease-free transposase comprises ORF1 of the transposase but omits ORF2 of the transposase.
[0257] When the system comprises an expression construct for expressing a nuclease-deficient transposase, the expression construct comprises a promoter operably linked to a nucleic acid sequence encoding the nuclease-deficient transposase. Transposases and expression constructs for expressing a nuclease- deficient transposase can be as described in Section l(a) herein above and Section II herein below, respectively.
[0258] The engineered system also comprises a programmable targeting nuclease of one or more nucleic acid constructs for expressing the programmable targeting nuclease. The expression construct for expressing the programmable targeting nuclease comprises a promoter operably linked to a nucleic acid sequence encoding the programmable targeting nuclease. Programmable targeting nucleases and expression constructs for expressing targeting nucleases can be as described in Section l(c) herein above and Section II herein below.
[0259] The targeting nuclease is engineered to excise the donor transposable nucleic acid construct. The targeting nuclease is also engineered to introduce a cut in a target nucleic acid locus thereby guiding insertion of the excised donor polypeptide at the target nucleic acid locus to generate a genetically modified cell comprising the donor polynucleotide inserted at the target nucleic acid locus. The gRNAs are designed so that the targeting nuclease cannot cut the target site position once the donor polynucleotide has been inserted.
[0260] In some aspects, the targeting nuclease can be linked to the nuclease- deficient transposase. In other aspects, the nuclease-deficient transposase is not linked to the targeting nuclease.
[0261] In some aspects, the nuclease-deficient transposase is a split transposase. In some aspects, the nuclease-deficient transposase is a Pong or Ponglike split transposase comprising a Pong ORF1 protein and a nuclease-deficient Pong ORF2 protein, wherein the Pong ORF2 protein comprises an inactivating mutation in the nuclease domain of Pong ORF2. In some aspects, the nuclease-deficient transposase is a Pong or Pong-like split transposase comprising a Pong ORF1 protein and a nuclease-deficient Pong ORF2 protein, wherein the nuclease-deficient Pong ORF2 protein comprises a modification in a DDE catalytic site that inactivates the nuclease function of the protein.
[0262] In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 1. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 1.
[0263] In some aspects, the nuclease deficient Pong ORF2 protein comprises an amino acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 121 . In some aspects, the nuclease deficient Pong ORF2 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 121.
[0264] In some aspects, the expression construct for expressing the nuclease- deficient transposase comprises an expression construct for expressing the Pong ORF1 protein wherein the expression construct for expressing the Pong ORF1 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 100; and an expression construct for expressing the nuclease-deficient Pong ORF2 protein, wherein the expression construct for expressing the nuclease deficient Pong ORF2 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%,91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 120. In some aspects, the expression construct for expressing the nuclease-deficient transposase comprises: an expression construct for expressing the Pong ORF1 protein wherein the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100; and an expression construct for expressing the nuclease deficient Pong ORF2 protein wherein the expression construct for expressing the nuclease deficient Pong ORF2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 120.
[0265] In some aspects, the nuclease-deficient transposase is a Pong or Ponglike split transposase comprising a Pong ORF1 protein, wherein the modified nuclease deficient Pong ORF2 protein comprises an absence of a Pong ORF2 protein. In some aspects, the nuclease-deficient transposase comprises a Pong ORF1 protein. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 1. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 1 . In some aspects, an expression construct for expressing the Pong ORF1 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 100 and wherein the engineered system comprises an absence of an expression construct for expressing Pong ORF2. In some aspects, the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100 and wherein the engineered system comprises an absence of an expression construct for expressing Pong ORF2.
[0266] The transposition sequences of a donor polynucleotide of an engineered system of the instant disclosure can be transposition sequences of a miniature inverted-repeat transposable element (MITE). In some aspects, the MITE is an mPing MITE. In some aspects, transposition sequences of the mPing MITE comprise mPing inverted repeat 1 and inverted repeat 2, wherein mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 7, SEQ ID NO: 111 , or SEQ ID NO: 108, and mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 8, SEQ ID NO: 112, or SEQ ID NO: 109. In some aspects, transposition sequences of the mPing MITE comprise mPing inverted repeat 1 and inverted repeat 2, wherein mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 7, SEQ ID NO: 111 , or SEQ ID NO: 108, and mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 8, SEQ ID NO: 112, or SEQ ID NO: 109. In some aspects, transposition sequences of the mPing MITE comprise mPing inverted repeat 1 and inverted repeat 2, wherein mPing inverted repeat 1 comprises a nucleotide sequence comprising the nucleic acid sequence of SEQ ID NO: 7, SEQ ID NO: 111 , or SEQ ID NO: 108, and mPing inverted repeat 2 comprises a nucleotide sequence comprising the nucleic acid sequence of SEQ ID NO: 8, SEQ ID NO: 112, or SEQ ID NO: 109.
[0267] A programmable targeting nuclease of a system of the instant disclosure can comprise a programmable, sequence-specific nucleic acid-binding domain and a nuclease domain. For instance, the programmable targeting nuclease can be an RNA- guided clustered regularly interspersed short palindromic repeats (CRISPR) nuclease system, a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, a ssDNA-guided Argonaute endonuclease, a meganuclease, a rare-cutting endonuclease, or any combination thereof.
[0268] In some aspects, the programmable targeting nuclease is a CRISPR- associated (Cas) (CRISPR / Cas) nuclease system comprising a nuclease and a guide RNA (gRNA). In some aspects, the programmable targeting nuclease comprises a Cas9 nuclease and a gRNA.
[0269] In some aspects, the Cas9 nuclease is linked to the nuclease deficient Pong ORF2 protein. In some aspects, the nuclease deficient Pong ORF2 protein is linked to the Cas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64. In other aspects, the nuclease deficient Pong ORF2 protein is linked to the Cas9 nuclease by three copies of a G4S linker of SEQ ID NO: 64.
[0270] In some aspects, the Cas9 nuclease is not linked to the nuclease deficient Pong ORF2 protein. When the Cas9 nuclease is not linked to the nuclease deficient Pong ORF2 protein, the engineered system can comprise a Cas9 nuclease or a nucleic acid expression construct for expressing the Cas9 nuclease. The cas9 nuclease can comprise an amino acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 5. In some aspects, the cas9 nuclease can comprise an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 5. In some aspects, a nucleic acid sequence encoding the Cas9 protein comprises about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 6. In some aspects, a nucleic acid sequence encoding the Cas9 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 6. In some aspects, the system comprises a nucleic acid construct for expressing a Cas9 nuclease, wherein the expression construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 92 or a nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94. In some aspects, the expression construct for expressing the Cas9 nuclease comprises a nucleic acidsequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 92 or a nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94.
[0271] In some aspects, the engineered system comprises (a) one or more excision gRNAs or one or more nucleic acid expression constructs for expressing one or more excision gRNAs that guide the Cas9 nuclease to nucleic acid sequences in or flanking the transposition sequences of the transposable nucleic acid construct to guide excision of the transposable nucleic acid construct and (b) one or more targeting gRNA or one or more expression constructs for expressing the one or more targeting gRNAs that guide the Cas9 nuclease to the target nucleic acid locus to introduce a cut at the target nucleic acid locus, thereby guiding insertion of the excised transposable nucleic acid construct at the target nucleic acid locus by the nuclease-deficient transposase to generate a genetically modified cell. In some aspects, the engineered system comprises one excision gRNA that guides the Cas9 nuclease to a common sequence flanking the transposition sequences of a transposable nucleic acid construct of the system or an expression construct for expressing one excision gRNA. In some aspects, the excision gRNA comprises a nucleic acid sequence of SEQ ID NO: 118. In some aspects, the targeting gRNA that guides the Cas9 nuclease to the target nucleic acid locus to introduce a cut at the target nucleic acid locus comprises a nucleic acid sequence of SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 80, SEQ ID NO: 113, SEQ ID NO: 67 and SEQ ID NO: 113, or any combination thereof.
[0272] In some aspects, the one or more nucleic acid expression constructs for expressing the excision and targeting gRNAs comprise a nucleic acid sequence comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 122 and wherein a nucleic acid sequence starting at base 425 to base 444 is replaced with a nucleic acid sequence of a gRNA for targeting the transposase and nuclease to a target nucleic acid locus. In some aspects, the one or more nucleic acid expression constructs for expressing the excision and targeting gRNAs comprise a nucleic acid sequence comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 122 and wherein a nucleic acid sequence starting at base 425to base 444 is replaced with a nucleic acid sequence of a gRNA for targeting the transposase and nuclease to a target nucleic acid locus.
[0273] In some aspects, the engineered system of the instant disclosure comprises a Pong ORF1 protein; a nuclease deficient Pong ORF2 protein; a Cas9 nuclease; one or more excision gRNAs and one or more targeting gRNAs; and a donor transposable nucleic acid construct. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 1. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 1 . In some aspects, the nuclease deficient Pong ORF2 protein comprises an amino acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 121. In some aspects, the nuclease deficient Pong ORF2 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 121 . The cas9 nuclease can comprise an amino acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 5. In some aspects, the cas9 nuclease can comprise an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 5. In some aspects, the excision gRNA comprises a nucleic acid sequence of SEQ ID NO: 118. In some aspects, the targeting gRNA that guides the Cas9 nuclease to the target nucleic acid locus to introduce a cut at the target nucleic acid locus comprises a nucleic acid sequence of SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 80, SEQ ID NO: 113, SEQ ID NO: 67 and SEQ ID NO: 113, or any combination thereof.
[0274] In some aspects, the engineered system of the instant disclosure comprises a Pong ORF1 protein; a Cas9 nuclease; one or more excision gRNAs andone or more targeting gRNAs; and a donor transposable nucleic acid construct, wherein the system does not comprise a Pong ORF2 protein. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 1. In some aspects, the Pong ORF1 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 1 . The cas9 nuclease can comprise an amino acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 5. In some aspects, the cas9 nuclease can comprise an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 5. In some aspects, the excision gRNA comprises a nucleic acid sequence of SEQ ID NO: 118. In some aspects, the targeting gRNA that guides the Cas9 nuclease to the target nucleic acid locus to introduce a cut at the target nucleic acid locus comprises a nucleic acid sequence of SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 80, SEQ ID NO: 113, SEQ ID NO: 67 and SEQ ID NO: 113, or any combination thereof.
[0275] In some aspects, the engineered system of the instant disclosure comprises an expression construct for expressing the Pong ORF1 protein; an expression construct for expressing the nuclease deficient Pong ORF2 protein; an expression construct for expressing the Cas9 nuclease; one or more expression constructs for expressing one or more excision gRNAs and a targeting gRNA; and a donor transposable nucleic acid construct. In some aspects, the expression construct for expressing the Pong ORF1 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 100; the expression construct for expressing the nuclease-deficient Pong ORF2 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 120; the expression construct for expressinga Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94; and the one or more excision gRNAs comprises a nucleic acid sequence of SEQ ID NO: 118. In some aspects, the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100; the expression construct for expressing the modified Pong ORF2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 120; the expression construct for expressing a Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94; and the one or more excision gRNAs comprises a nucleic acid sequence of SEQ ID NO: 118.
[0276] In some aspects, the engineered system of the instant disclosure comprises an expression construct for expressing the Pong ORF1 protein; an expression construct for expressing the Cas9 nuclease; one or more expression constructs for expressing one or more excision gRNAs and a targeting gRNA; and a donor transposable nucleic acid construct, wherein the system does not comprise an expression construct for expressing the ORF2 protein. In some aspects, the expression construct for expressing the Pong ORF1 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %,92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ IDNO: 100; the expression construct for expressing a Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%,83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%,98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94; and the one or more excision gRNAs comprises a nucleic acid sequence of SEQ ID NO: 118. In some aspects, the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100%sequence identity with SEQ ID NO: 100; the expression construct for expressing a Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94; and the one or more excision gRNAs comprises a nucleic acid sequence of SEQ ID NO: 118.
[0277] In some aspects, the donor polynucleotide comprises a cargo polynucleotide flanked by the transposition sequences of the system of the instant disclosure. Cargo polynucleotides can be as described in Section IV herein below. In one aspects, the cargo polynucleotide comprises HSEs. In another aspects, the cargo polynucleotide comprises an expression construct for expressing a herbicide resistance function. The herbicide resistance function can be resistance to bialaphos herbicide. In some aspects, when the herbicide resistance function can be resistance to bialaphos herbicide, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding resistance to bialaphos and wherein the nucleic acid construct comprising the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97 or SEQ ID NO: 99. In some aspects, the cargo polynucleotide comprises an expression construct comprising a promoter operably linked to a polynucleotide encoding resistance to bialaphos and wherein the nucleic acid construct comprising the donor polynucleotide comprises a nucleic acid sequencing comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97.
[0278] The target nucleic acid locus can be in a nuclear, organellar, or extrachromosomal nucleic acid sequence. In some aspects, the target nucleic acid locus is in a protein-coding gene, an RNA coding gene, or an intergenic region.
[0279] In some aspects, the cell is a eukaryotic cell. In some aspects, the cell is a plant cell, a plant or part thereof, or seed. In some aspects, the plant is an Arabidopsis sp. or a soybean plant cell, a plant or part thereof, or seed.II. Nucleic Acid Constructs
[0280] A further aspect of the present disclosure provides one or more nucleic acid constructs encoding the components of the engineered system described above in Section I. In some aspects, the engineered system of nucleic acid constructs encodes the engineered system described in Section 1(d).
[0281] Any of the multi-component engineered systems described herein are to be considered modular, in that the different components can optionally be distributed among two or more nucleic acid constructs as described herein. The nucleic acid constructs can be DNA or RNA, linear or circular, single-stranded or double-stranded, or any combination thereof. The nucleic acid constructs can be codon optimized for efficient translation into protein, and possibly for transcription into an RNA donor polynucleotide transcript in the cell of interest. Codon optimization programs are available as freeware or from commercial sources.
[0282] The nucleic acid constructs can be used to express one or more components of the engineered system for later introduction into a cell to be genetically modified. Alternatively, the nucleic acid constructs can be introduced into the cell to be genetically modified for expression of the components of the engineered system in the cell.
[0283] Expression constructs generally comprise DNA coding sequences operably linked to at least one promoter control sequence for expression in a cell of interest. Promoter control sequences can control expression of the transposase, the programmable targeting nuclease, the donor polynucleotide, or combinations thereof in bacterial (e.g., E. coli) cells or eukaryotic (e.g., yeast, insect, mammalian, or plant) cells. Suitable bacterial promoters include, without limit, T7 promoters, lac operon promoters, trp promoters, tac promoters (which are hybrids of trp and lac promoters), variations of any of the foregoing, and combinations of any of the foregoing. Non-limiting examples of suitable eukaryotic promoters include constitutive, regulated, or cell- or tissue-specific promoters. Suitable eukaryotic constitutive promoter control sequences include, but are not limited to, cytomegalovirus immediate early promoter (CMV), simian virus (SV40) promoter, adenovirus major late promoter, Rous sarcoma virus (RSV) promoter, mouse mammary tumor virus (MMTV) promoter, phosphoglycerate kinase (PGK) promoter, elongation factor (EDI)-alpha promoter, ubiquitin promoters, actin promoters, tubulinpromoters, immunoglobulin promoters, fragments thereof, or combinations of any of the foregoing. Examples of suitable eukaryotic regulated promoter control sequences include, without limit, those regulated by heat shock, metals, steroids, antibiotics, or alcohol. Non-limiting examples of tissue-specific promoters include B29 promoter, CD14 promoter, CD43 promoter, CD45 promoter, CD68 promoter, desmin promoter, elastase-1 promoter, endoglin promoter, fibronectin promoter, Flt-1 promoter, GFAP promoter, GPIIb promoter, ICAM-2 promoter, INF-|3 promoter, Mb promoter, Nphsl promoter, OG-2 promoter, SP-B promoter, SYN1 promoter, and WASP promoter.
[0284] Promoters can also be plant-specific promoters, or promoters that can be used in plants. A wide variety of plant promoters are known to those of ordinary skill in the art, as are other regulatory elements that can be used alone or in combination with promoters. Preferably, promoter control sequences control expression in cassava such as promoters disclosed in Wilson et al., 2017, The New Phytologist, 213(4): 1632-1641 , the disclosure of which is incorporated herein in its entirety.
[0285] Promoters can be divided into two types, namely, constitutive promoters and non-constitutive promoters. Constitutive promoters are classified as providing for a range of constitutive expression. Thus, some are weak constitutive promoters, and others are strong constitutive promoters. Non-constitutive promoters include tissuepreferred promoters, tissue-specific promoters, cell-type specific promoters, and inducible-promoters. Suitable plant-specific constitutive promoter control sequences include, but are not limited to, a CaMV35S promoter, CaMV 19S, GOS2, Arabidopsis At6669 promoter, Rice cyclophilin, Maize H3 histone, Synthetic Super MAS, an opine promoter, a plant ubiquitin (Ubi) promoter, an actin 1 (Act-1 ) promoter, pEMU, Cestrum yellow leaf curling virus promoter (CYMLV promoter), and an alcohol dehydrogenase 1 (Adh-1 ) promoter. Other constitutive promoters include those in U.S. Pat. Nos.5,659,026; 5,608,149; 5,608,144; 5,604,121 ; 5,569,597; 5,466,785; 5,399,680; 5,268,463; and 5,608,142.
[0286] Regulated plant promoters respond to various forms of environmental stresses, or other stimuli, including, for example, mechanical shock, heat, cold, flooding, drought, salt, anoxia, pathogens such as bacteria, fungi, and viruses, and nutritional deprivation, including deprivation during times of flowering and / or fruiting, and other forms of plant stress. For example, the promoter can be a promoter which is induced byone or more, but not limited to one of the following: abiotic stresses such as wounding, cold, desiccation, ultraviolet-B, heat shock or other heat stress, drought stress or water stress. The promoter can further be one induced by biotic stresses including pathogen stress, such as stress induced by a virus or fungi, stresses induced as part of the plant defense pathway or by other environmental signals, such as light, carbon dioxide, hormones or other signaling molecules such as auxin, hydrogen peroxide and salicylic acid, sugars and gibberellin or abscisic acid and ethylene. Suitable regulated plant promoter control sequences include, but are not limited to, salt-inducible promoters such as RD29A; drought-inducible promoters such as maize rab17 gene promoter, maize rab28 gene promoter, and maize Ivr2 gene promoter; heat-inducible promoters such as heat tomato hsp80-promoter from tomato.
[0287] Tissue-specific promoters can include, but are not limited to, fiber-specific, green tissue-specific, root-specific, stem-specific, flower-specific, callus-specific, pollenspecific, egg-specific, and seed coat-specific. Suitable tissue-specific plant promoter control sequences include, but are not limited to, leaf-specific promoters [such as described, for example, by Yamamoto et al., Plant J. 12:255-265, 1997; Kwon et al., Plant Physiol. 105:357-67, 1994; Yamamoto et al., Plant Cell Physiol. 35:773-778, 1994; Gotor et al., Plant J. 3:509-18, 1993; Orozco et al., Plant Mol. Biol. 23:1129-1138, 1993; and Matsuoka et al., Proc. Natl. Acad. Sci. USA 90:9586-9590, 1993], seedpreferred promoters [e.g., from seed-specific genes (Simon et al., Plant Mol. Biol. 5. 191 , 1985; Scofield et al., J. Biol. Chem. 262: 12202, 1987; Baszczynski et al., Plant Mol. Biol. 14: 633, 1990), Brazil Nut albumin (Pearson et al., Plant Mol. Biol. 18: 235- 245, 1992), legumin (Ellis et al., Plant Mol. Biol. 10: 203-214, 1988), Glutelin (rice) (Takaiwa et al., Mol. Gen. Genet. 208: 15-22, 1986; Takaiwa et al., FEBS Letts. 221 : 43-47, 1987), Zein (Matzke et al., Plant Mol Biol, 143: 323-32, 1990), napA (Stalberg et al., Planta 199: 515-519, 1996), Wheat SPA (Albanietal, Plant Cell, 9: 171-184, 1997), sunflower oleosin (Cummins et al., Plant Mol. Biol. 19: 873-876, 1992)], endosperm specific promoters [e.g., wheat LMW and HMW, glutenin-1 (Mol Gen Genet 216:81-90, 1989; NAR 17:461-2), wheat a, b and g gliadins (EMBO3: 1409-15, 1984), Barley Itrl promoter, barley B1 , C, D hordein (Theor Appl Gen 98:1253-62, 1999; Plant J 4:343-55, 1993; Mol Gen Genet 250:750-60, 1996), Barley DOF (Mena et al., The Plant Journal, 116(1 ): 53-62, 1998), Biz2 (EP99106056.7), Synthetic promoter (Vicente-Carbajosa etal., Plant J. 13: 629-640, 1998), rice prolamin NRP33, rice-globulin Glb-1 (Wu et al., Plant Cell Physiology 39(8) 885-889, 1998), rice alpha-globulin REB / OHP-1 (Nakase et al., Plant Mol. Biol. 33: 513-S22, 1997), rice ADP-glucose PP (Trans Res 6:157-68, 1997), maize ESR gene family (Plant J 12:235-46, 1997), sorgum gamma-kafirin (PMB 32:1029-35, 1996)], embryo-specific promoters [e.g., rice OSH1 (Sato et al., Proc. Natl. Acad. Sci. USA, 93: 8117-8122), KNOX (Postma-Haarsma et al., Plant Mol. Biol.39:257-71 , 1999), rice oleosin (Wu et al., J. Biochem., 123:386, 1998)], and flowerspecific promoters [e.g., AtPRP4, chalene synthase (chsA) (Van der Meer et al., Plant Mol. Biol. 15, 95-109, 1990), LAT52 (Twell et al., Mol. Gen Genet. 217:240-245; 1989), apetala-3],
[0288] Any of the promoter sequences can be wild type or can be modified for more efficient or efficacious expression. The DNA coding sequence also can be linked to a polyadenylation signal (e.g., SV40 polyA signal, bovine growth hormone (BGH) polyA signal, etc.) and / or at least one transcriptional termination sequence. In some situations, the complex or fusion protein can be purified from the bacterial or eukaryotic cells.
[0289] Nucleic acid constructs encoding one or more components of an engineered system of the instant disclosure can be present in, or cloned into, one or more vectors. Suitable vectors include plasmid vectors, viral vectors, and selfreplicating RNA (Yoshioka et al., Cell Stem Cell, 2013, 13:246-254). For instance, the nucleic acid encoding one or more components of an engineered system of the instant disclosure can be present in a plasmid vector.
[0290] Non-limiting examples of suitable plasmid constructs include pUC, pBR322, pET, pBluescript, and variants thereof. Alternatively, the nucleic acid encoding one or more components of an engineered system of the instant disclosure can be part of a viral vector (e.g., lentiviral vectors, adeno-associated viral vectors, adenoviral vectors, and so forth).
[0291] The plasmid or viral vector can comprise additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcriptional termination sequences, etc.), selectable reporter sequences (e.g., antibiotic resistance genes), origins of replication, T-DNA border sequences, and the like. The plasmid or viral vector can further comprise RNA processing elements such asglycine tRNAs, or Csy4 recognition sites. Such RNA processing elements can, for instance, intersperse polynucleotide sequences encoding multiple gRNAs under the control of a single promoter to produce the multiple gRNAs from a transcript encoding the multiple gRNAs. When a cys4 recognition cite is used, a vector can further comprise sequences for expression of Csy4 RNAse to process the gRNA transcript. Additional information about vectors and use thereof can be found in “Current Protocols in Molecular Biology”, Ausubel et al., John Wiley & Sons, New York, 2003, or “Molecular Cloning: A Laboratory Manual”, Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.(a) Aspects of constructs expressing an engineered system
[0292] In some aspects, a nucleic acid construct of the instant disclosure comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing the Pong ORF1 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO:100. In some aspects, the nucleic acid expression construct for expressing a Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100.
[0293] In some aspects, a nucleic acid construct of the instant disclosure comprises a nucleic acid expression construct for expressing a Pong ORF2 protein, wherein the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO:101. In some aspects, the nucleic acid expression construct for expressing a Pong ORF2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 101.
[0294] In some aspects, a nucleic acid construct of the instant disclosure comprises a nucleic acid expression construct for expressing a nuclease-deficient Pong ORF2 protein. In some aspects, an expression construct for expressing the nuclease deficient Pong ORF2 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 120. In some aspects, an expression construct for expressing the nuclease deficient Pong ORF2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 120.
[0295] In some aspects, a nucleic acid construct of the instant disclosure comprises a nucleic acid expression construct for expressing a Cas9 protein, wherein the expression construct for expressing the Cas9 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102. In some aspects, the nucleic acid expression construct for expressing a Cas9 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102.
[0296] In some aspects, a nucleic acid construct of the instant disclosure comprises a nucleic acid expression construct for expressing a gRNA for targeting a transposase and nuclease to the DD20 intergenic region of soybean, wherein the expression construct for expressing the gRNA for targeting a transposase and nuclease of the instant disclosure to the DD20 intergenic region of soybean comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 105. In some aspects, the nucleic acid expression construct for expressing a gRNA directed to the DD20 intergenic region of soybean comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 105.
[0297] In some aspects, a system of the instant disclosure is a one-component system, wherein the Pong ORF2 protein is linked to the Cas9 nuclease and the donorpolynucleotide is inserted in a nucleic acid expression construct encoding a GFP reporter, thereby inactivating the reporter. In these aspects, the target nucleic acid locus is in an Arabidopsis PDS3 gene. The system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing the Pong ORF1 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100 or the nucleic acid sequence starting at base 5073 to base 8215 of SEQ ID NO: 89. In some aspects, the nucleic acid expression construct for expressing a Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 100 or the nucleic acid sequence starting at base 5073 to base 8215 of SEQ ID NO: 89. The system also comprises a nucleic acid expression construct for expressing a Pong ORF2 protein linked to Cas9 nuclease by a single copy of the G4S linker (SEQ ID NO: 64), wherein the construct comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 115 or a nucleic acid sequence starting at base 7451 to base 15799 of SEQ ID NO: 74. In some aspects, the construct for expressing a Pong ORF2 protein linked to Cas9 nuclease by a single copy of the G4S linker (SEQ ID NO: 64) comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 115 or a nucleic acid sequence starting at base 7451 to base 15799 of SEQ ID NO: 74. The system further comprises a nucleic acid expression construct comprising a promoter operably linked to a polynucleotide sequence encoding GFP, wherein the donor polynucleotide inserted in the nucleic acid expression construct. In some aspects, the GFP expression construct comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at nucleotide 2414 to nucleotide 23460 and nucleotide 1 to nucleotide 42 of SEQ ID NO: 74. In some aspects, the GFP expression construct comprises a nucleicacid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at nucleotide 2414 to nucleotide 23460 and nucleotide 1 to nucleotide 42 of SEQ ID NO: 74. The system further comprises an expression construct for expressing a gRNA, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 2632 to base 3343 of SEQ ID NO: 74. In some aspects, the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 2632 to base 3343 of SEQ ID NO: 74. In some aspects, the system is encoded on a plasmid comprising a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 74. In some aspects, the system is encoded on a plasmid comprising a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 74.
[0298] In some aspects, a system of the instant disclosure is a one-component system, wherein the Pong ORF2 protein is linked to the Cas9 nuclease and the donor polynucleotide is inserted in a nucleic acid expression construct encoding a GFP reporter, thereby inactivating the reporter. In these aspects, the target nucleic acid locus is in an actin 8 (ACT8) gene. The system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing the Pong ORF1 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 1456 to base 5362 of SEQ ID NO: 92. In some aspects, the nucleic acid expression construct for expressing a Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base1456 to base 5362 of SEQ ID NO: 92. The system also comprises a nucleic acid expression construct for expressing a Pong ORF2 protein linked to Cas9 nuclease, wherein the construct comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 108 or the nucleic acid sequence starting at base 5548 to base 12904 of SEQ ID NO: 92. In some aspects, the construct for expressing a Pong ORF2 protein linked to Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 108 or the nucleic acid sequence starting at base 5548 to base 12904 of SEQ ID NO: 92. The system further comprises a nucleic acid construct comprising the donor polynucleotide, wherein the nucleic acid construct comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 69 to base 498 of SEQ ID NO: 92. In some aspects, the nucleic acid construct comprising the donor polynucleotide comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 69 to base 498 of SEQ ID NO: 92. The system comprises an expression construct for expressing a gRNA, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 729 to base 1440 of SEQ ID NO: 92. In some aspects, the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 729 to base 1440 of SEQ ID NO: 92. In some aspects, the system is encoded on a plasmid comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 92. Insome aspects, the system is encoded on a plasmid comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 92.
[0299] In other aspects, a system of the instant disclosure is a one-component system, wherein the Pong ORF2 protein linked to a Cas9 nuclease and the target nucleic acid locus is in an Arabidopsis actin 8 (ACT8) gene. In these aspects, the donor polynucleotide comprises a nucleotide sequence comprising heat shock element (HSE) sequences flanked by mPing inverted repeat 1 and inverted repeat 2. The system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing the Pong ORF1 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 1481 to base 5390 of SEQ ID NO: 93. In some aspects, the nucleic acid expression construct for expressing a Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 1481 to base 5390 of SEQ ID NO: 93. The system also comprises a nucleic acid expression construct for expressing a Pong ORF2 protein linked to Cas9 nuclease, wherein the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 1481 to base 5390 of SEQ ID NO: 93. In some aspects, the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 1481 to base 5390 of SEQ ID NO: 93. The system further comprises a nucleic acid construct comprising the donor polynucleotide, wherein the donor polynucleotide comprises a nucleotide sequence comprising HSE sequences flanked by mPing inverted repeat 1 and inverted repeat 2, and wherein the donor polynucleotide comprises about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%,99%, or 100% sequence identity with the nucleic acid sequence starting at base 69 to base 512 of SEQ ID NO: 93. In some aspects, the donor polynucleotide comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 69 to base 512 of SEQ ID NO: 93. The system comprises an expression construct for expressing a gRNA, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 754 to base 1465 of SEQ ID NO: 93. In some aspects, the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 754 to base 1465 of SEQ ID NO: 93. In some aspects, the system is encoded on a plasmid comprising a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 93. In some aspects, the system is encoded on a plasmid comprising a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 93.
[0300] In some aspects, a system of the instant disclosure is a one-component system, wherein the Cas9 protein is not linked to the Pong ORF2 protein, and the target nucleic acid locus is in a soybean DD20 intergenic region. The system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing the Pong ORF1 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with nucleic acid sequence starting at base 3593 to base 7502 of SEQ ID NO: 94. In some aspects, the nucleic acid expression construct for expressing a Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 3593 to base 7502 of SEQ ID NO: 94. The system alsocomprises a nucleic acid expression construct for expressing a Pong ORF2 protein, wherein the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 7685 to base 10827 of SEQ ID NO: 94. In some aspects, the expression construct for expressing the Pong ORF2 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 7685 to base 10827 of SEQ ID NO: 94. The system also comprises a nucleic acid expression construct for expressing a Cas9 nuclease, wherein the construct comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94. In some aspects, the construct for expressing the Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94. The system comprises a nucleic acid construct comprising the donor polynucleotide, wherein the nucleic acid construct comprising the donor polynucleotide comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 2201 to base 2630 of SEQ ID NO: 94. The system also comprises an expression construct for expressing a gRNA targeting the soybean DD20 intergenic region, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 103 or thenucleic acid sequence starting at base 2861 to base 3572 of SEQ ID NO: 94. In some aspects, the system is encoded on a plasmid comprising a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 103 or the nucleic acid sequence starting at base 2861 to base 3572 of SEQ ID NO: 94. In some aspects, the system is encoded on a plasmid comprising a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 94.
[0301] In some aspects, a system of the instant disclosure is a one-component system, wherein the Cas9 protein is linked to the Pong ORF2 protein, the donor construct is inserted in an expression construct expressing a GFP reporter, and the target nucleic acid locus is in a soybean DD20 intergenic region. The system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing the Pong ORF1 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 5490 to base 9399 of SEQ ID NO: 95. In some aspects, the nucleic acid expression construct for expressing a Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 5490 to base 9399 of SEQ ID NO: 95. The system also comprises a nucleic acid expression construct for expressing a Pong ORF2 protein linked to a Cas9 nuclease, wherein the expression construct for expressing the Pong ORF2 protein linked to a Cas9 nuclease comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 9582 to base 16938 of SEQ ID NO: 95. In some aspects, the expression construct for expressing the Pong ORF2 protein linked to a Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 9582 to base16938 of SEQ ID NO: 95. The system comprises a nucleic acid construct comprising the donor polynucleotide, wherein the nucleic acid construct comprising the donor polynucleotide comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 4545 to base 2173 of SEQ ID NO: 95. The system also comprises an expression construct for expressing a gRNA, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 4763 to base 5474 of SEQ ID NO: 95. In some aspects, the system is encoded on a plasmid comprising a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 95. In some aspects, the system is encoded on a plasmid comprising a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 95.
[0302] In some aspects, the system of the instant disclosure comprises a helper construct and a donor construct, wherein the helper construct comprises a nucleic acid expression construct for expressing Pong ORF1 and a nucleic acid expression construct for expressing Pong ORF2 protein linked to a Cas9 nuclease. The system comprises a nucleic acid expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing the Pong ORF1 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 981 to base 4890 of SEQ ID NO: 75. In some aspects, the nucleic acid expression construct for expressing a Pong ORF1 protein comprises at least about 75% or more, atleast about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 981 to base 4890 of SEQ ID NO: 75. The system also comprises a nucleic acid expression construct for expressing a Pong ORF2 protein linked to Cas9 nuclease, wherein the construct comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 5073 to base 12429 of SEQ ID NO: 75. In some aspects, the construct for expressing a Pong ORF2 protein linked to Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 5073 to base 12429 of SEQ ID NO: 75. The system further comprises an expression construct for expressing a gRNA, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 254 to base 965 of SEQ ID NO: 75. In some aspects, the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 254 to base 965 of SEQ ID NO: 75. In some aspects, the system is encoded on a plasmid comprising a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 75. In some aspects, the system is encoded on a plasmid comprising a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 75.
[0303] In some aspects, the donor polynucleotide is inserted in a nucleic acid expression construct encoding a GFP reporter, thereby inactivating the reporter. In some aspects, the expression construct is inserted in nucleic acid sequence in the genome of the cell. In some aspects, the target nucleic acid locus is in an Arabidopsis PDS3 gene.
[0304] In some aspects, the system of the instant disclosure comprises a helper construct and a donor construct. In some aspects, the donor construct comprises a nucleic acid expression construct encoding a GFP reporter. The donor nucleic acid construct is inserted into the expression construct thereby inactivating the reporter. In these aspects, the target nucleic acid locus is an Arabidopsis ADH1 gene. The helper construct comprises a nucleic acid expression construct for expressing Pong ORF1 , a nucleic acid expression construct for expressing Pong ORF2 protein, and a nucleic acid construct for expressing a deCas9 nickase. The expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 981 to base 4890 of SEQ ID NO: 89. In some aspects, the nucleic acid expression construct for expressing a Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 981 to base 4890 of SEQ ID NO: 89. The system also comprises a nucleic acid expression construct for expressing a Pong ORF2 protein, wherein the construct comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 5073 to base 8215 of SEQ ID NO: 89. In some aspects, the construct for expressing a Pong ORF2 protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 5073 to base 8215 of SEQ ID NO: 89. The system also comprises a nucleic acid expression construct for expressing a deCas9 nickase, wherein the construct comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at nucleotide 8218 to nucleotide 13856 of SEQ ID NO: 89. In some aspects, the construct for expressing a deCas9 nickase protein comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity withthe nucleic acid sequence starting at nucleotide 8218 to nucleotide 13856 of SEQ ID NO: 89. The system further comprises an expression construct for expressing a gRNA, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 254 to base 965 of SEQ ID NO: 89. In some aspects, the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 254 to base 965 of SEQ ID NO: 89. In some aspects, the helper construct is encoded on a plasmid comprising a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 89. In some aspects, the helper construct is encoded on a plasmid comprising a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 89.
[0305] In some aspects, the system of the instant disclosure comprises a helper construct and a donor construct. In some aspects, the donor construct comprises a nucleic acid expression construct encoding a GFP reporter, wherein the donor nucleic acid construct is inserted into the expression construct thereby inactivating the reporter. In these aspects, the target nucleic acid locus is an Arabidopsis ACT8 gene. The helper construct comprises a nucleic acid expression construct for expressing Pong ORF1 and a nucleic acid expression construct for expressing Pong ORF2 protein linked to a Cas9 nuclease. In some aspects, the expression construct for expressing a Pong ORF1 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 981 to base 4890 of SEQ ID NO: 91. In some aspects, the nucleic acid expression construct for expressing a Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or100% sequence identity with the nucleic acid sequence starting at base 981 to base 4890 of SEQ ID NO: 91. The system also comprises a nucleic acid expression construct for expressing a Pong ORF2 protein linked to Cas9 nuclease, wherein the construct comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 5073 to base 12429 of SEQ ID NO: 91 . In some aspects, the construct for expressing a Pong ORF2 protein linked to Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 5073 to base 12429 of SEQ ID NO: 91 . The system further comprises an expression construct for expressing a gRNA, wherein the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 254 to base 965 of SEQ ID NO: 91 . In some aspects, the expression construct for expressing the gRNA comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 254 to base 965 of SEQ ID NO: 91 . In some aspects, the helper construct is encoded on a plasmid comprising a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 91 . In some aspects, the helper construct is encoded on a plasmid comprising a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 91 .
[0306] The donor construct comprises a nucleic acid expression construct comprising a promoter operably linked to a polynucleotide sequence encoding GFP, wherein the donor polynucleotide inserted in the nucleic acid expression construct. In some aspects, the GFP expression construct comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%,87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence starting at base 3037 clockwise to base 665 of SEQ ID NO: 90. In some aspects, the GFP expression construct comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 3037 clockwise to base 665 of SEQ ID NO: 90. In some aspects, the donor construct is encoded on a plasmid comprising a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 90. In some aspects, the donor construct is encoded on a plasmid comprising a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 90.
[0307] In some aspects, the programmable targeting system of the instant disclosure comprises a CRISPR nuclease system comprising dCas9 and a gRNA. In some aspects, the dCas9 nuclease is linked to Pong ORF2 by one copy of a G4S linker of SEQ ID NO: 64. In some aspects, the Pong ORF2 protein linked to the dCas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64 comprises an amino acid sequence encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 110. In some aspects, the Pong ORF2 protein linked to the dCas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64 comprises an amino acid sequence encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 110.
[0308] In some aspects, the Pong ORF2 protein linked to the dCas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64 is expressed using an expression construct for expressing the Pong ORF2 protein linked to the dCas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64, wherein the expression construct comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%,83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 115. In some aspects, the expression construct comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 115. In some aspects, the genetically modified cell is an Arabidopsis thaliana cell.
[0309] In some aspects, the Pong ORF2 protein linked to the Cas9 nuclease by three copies of a G4S linker of SEQ ID NO: 64 is expressed using an expression construct for expressing the Pong ORF2 protein linked to the Cas9 nuclease by three copies of a G4S linker of SEQ ID NO: 64, wherein the expression construct comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 104. In some aspects, the expression construct comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 104. In some aspects, the genetically modified cell is a soybean cell.(b) Constructs for generating genetically modified cells with no off-target insertions
[0310] Another aspect of the instant disclosure encompasses one or more nucleic acid constructs expressing an engineered nucleic acid modification system for generating a non-transgenic genetically modified cell comprising no or reduced off- target insertions. The nucleic acid modification system for generating a non-transgenic genetically modified cell comprising no or reduced off-target insertions can be described in Section l(e) herein above. In some aspects, the cell is a eukaryotic cell. The eukaryotic cell can be a plant cell, a plant or part thereof, or seed.
[0311] In some aspects, the one or more constructs comprise one or more expression constructs for expressing a Pong nuclease-free transposase, one or more constructs for expressing a programmable nuclease, and a transposable nucleic acid construct.
[0312] In some aspects, the one or more constructs comprise an expression construct for expressing a Pong ORF1 , one or more expression constructs for expressing a programmable nuclease, and a donor polynucleotide. In some aspects, an expression construct for expressing a Pong ORF1 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 100. In some aspects, an expression construct for expressing a Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100.
[0313] In some aspects, a system of the instant disclosure does not comprise an expression construct for expressing Pong ORF2. In other aspects, a system of the instant disclosure further comprises an expression construct for expressing the nuclease deficient Pong ORF2 protein wherein the expression construct for expressing the nuclease deficient Pong ORF2 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 120. In some aspects, the expression construct for expressing the nuclease-deficient transposase comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 120.
[0314] In some aspects, one or more expression constructs for expressing a programmable nuclease comprise an expression construct for expressing Cas9 and one or more expression constructs for expressing a gRNA. In some aspects, the nucleic acid expression construct for expressing a Cas9 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102. In some aspects, the nucleic acid expression construct for expressing a Cas9 protein comprises a nucleic acid sequence comprising about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102. In some aspects, the one or more nucleic acid expression constructs comprise one or more constructs for expressing one or more excision gRNAs that guide the Cas9 nuclease to nucleic acid sequences in or flanking the transposition sequencesof the transposable nucleic acid construct to guide excision of the transposable nucleic acid construct and for expressing a targeting gRNA that guides the Cas9 nuclease to the target nucleic acid locus to introduce a cut at the target nucleic acid locus, thereby guiding insertion of the excised transposable nucleic acid construct at the target nucleic acid locus by the nuclease-deficient transposase to generate a genetically modified cell. In some aspects, the engineered system comprises an expression construct for expressing one excision gRNA that guides the Cas9 nuclease to a common sequence flanking the transposition sequences of a transposable nucleic acid construct of the system. In some aspects, the excision gRNA comprises a nucleic acid sequence of SEQ ID NO: 118. In some aspects, the targeting gRNA that guides the Cas9 nuclease to the target nucleic acid locus to introduce a cut at the target nucleic acid locus comprises a nucleic acid sequence of SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 80, SEQ ID NO: 113, SEQ ID NO: 67 and SEQ ID NO: 113, or any combination thereof.
[0315] In some aspects, the one or more nucleic acid expression constructs for expressing the excision and targeting gRNAs comprise a nucleic acid sequence comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 122 and wherein a nucleic acid sequence starting at base 425 to base 444 is replaced with a nucleic acid sequence of a gRNA for targeting the transposase and nuclease to a target nucleic acid locus. In some aspects, the one or more nucleic acid expression constructs for expressing the excision and targeting gRNAs comprise a nucleic acid sequence comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 122 and wherein a nucleic acid sequence starting at base 425 to base 444 is replaced with a nucleic acid sequence of a gRNA for targeting the transposase and nuclease to a target nucleic acid locus.
[0316] In some aspects, the one or more nucleic acid constructs expressing an engineered nucleic acid modification system for generating a non-transgenic genetically modified cell comprising no or reduced off-target insertions comprises (a) an expression construct an expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing a Pong ORF1 protein comprises at least about75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 100; (b) an expression construct for expressing the nuclease deficient Pong ORF2 protein wherein the expression construct for expressing the nuclease deficient Pong ORF2 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 120; (c) a nucleic acid expression construct for expressing a Cas9 protein, wherein the expression construct for expressing the Cas9 protein comprises a nucleic acid sequence comprising about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102; (d) a nucleic acid expression construct for expressing a targeting gRNA, an excision gRNA, or both, wherein the expression construct comprises about 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 122. In some aspects, the one or more nucleic acid constructs expressing an engineered nucleic acid modification system for generating a non-transgenic genetically modified cell comprising no or reduced off-target insertions comprises (a) an expression construct an expression construct for expressing a Pong ORF1 protein, wherein the expression construct for expressing a Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100%sequence identity with SEQ ID NO: 100; (b) an expression construct for expressing the nuclease deficient Pong ORF2 protein wherein the expression construct for expressing the nuclease deficient Pong ORF2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 120; (c) a nucleic acid expression construct for expressing a Cas9 protein, wherein the expression construct for expressing the Cas9 protein comprises a nucleic acid sequence comprising about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 102; (d) a nucleic acid expression construct for expressing a targeting gRNA, an excision gRNA, or both, wherein the expression constructcomprises about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 122. In some aspects, the excision gRNA comprises a nucleic acid sequence of SEQ ID NO: 118.III. Cells
[0317] Another aspect of the instant disclosure encompasses a cell, a tissue, or an organism comprising an engineered system described in Section I above. One or more components of the engineered system in the cell can be encoded by one or more nucleic acid constructs of a system of nucleic acid constructs as described in Section II above.
[0318] A variety of cells are suitable for use in the methods disclosed herein. The cell can be a prokaryotic cell. Alternatively, the cell is a eukaryotic cell. For example, the cell can be a prokaryotic cell, a human mammalian cell, a non-human mammalian cell, a non-mammalian vertebrate cell, an invertebrate cell, an insect cell, a plant cell, a yeast cell, or a single cell eukaryotic organism. The cell can also be a one-cell embryo. For example, a non-human mammalian embryo including rat, hamster, rodent, rabbit, feline, canine, ovine, porcine, bovine, equine, plant, and primate embryos. The cell can also be a stem cell such as embryonic stem cells, ES-like stem cells, fetal stem cells, adult stem cells, and the like. The cell can be in vitro, ex vivo, or in vivo (i.e. , within an organism or within a tissue of an organism).
[0319] Non-limiting examples of suitable mammalian cells or cell lines include human embryonic kidney cells (HEK293, HEK293T); human cervical carcinoma cells (HELA); human lung cells (W138); human liver cells (Hep G2); human U2-OS osteosarcoma cells, human A549 cells, human A-431 cells, and human K562 cells; Chinese hamster ovary (CHO) cells; baby hamster kidney (BHK) cells; mouse myeloma NS0 cells; mouse embryonic fibroblast 3T3 cells (NIH3T3); mouse B lymphoma A20 cells; mouse melanoma B16 cells; mouse myoblast C2C12 cells; mouse myeloma SP2 / 0 cells; mouse embryonic mesenchymal C3H-10T1 / 2 cells; mouse carcinoma CT26 cells; mouse prostate DuCuP cells; mouse breast EMT6 cells; mouse hepatoma Hepa1c1c7 cells; mouse myeloma J5582 cells; mouse epithelial MTD-1A cells; mouse myocardial MyEnd cells; mouse renal RenCa cells; mouse pancreatic RIN-5F cells; mouse melanoma X64 cells; mouse lymphoma YAC-1 cells; rat glioblastoma 9L cells;rat B lymphoma RBL cells; rat neuroblastoma B35 cells; rat hepatoma cells (HTC); buffalo rat liver BRL 3A cells; canine kidney cells (MDCK); canine mammary (CMT) cells; rat osteosarcoma D17 cells; rat monocyte / macrophage DH82 cells; monkey kidney SV-40 transformed fibroblast (COS7) cells; monkey kidney CVI-76 cells; Afrimay green monkey kidney (VERO-76) cells. An extensive list of mammalian cell lines can be found in the Amerimay Type Culture Collection catalog (ATCC, Manassas, VA).
[0320] The cell can be a plant cell, a plant part, or a plant. Plant cells include germ cells and somatic cells. Non-limiting examples of plant cells include parenchyma cells, sclerenchyma cells, collenchyma cells, xylem cells, and phloem cells. Plant parts include, but are not limited to, stems, roots, ovules, stamens, leaves, embryos, meristematic regions, callus tissue, gametophytes, sporophytes, pollen, microspores, and the like. The plant can be a monocot plant or a dicot plant. For instance, the plant can be soybean; maize; sugar cane; beet; tobacco; wheat; barley; poppy; rape; sunflower; alfalfa; sorghum; rose; carnation; gerbera; carrot; tomato; lettuce; chicory; pepper; melon; cabbage; oat; rye; cotton; millet; flax; potato; pine; walnut; citrus (including oranges, grapefruit etc.); hemp; oak; rice; petunia; orchids; Arabidopsis; broccoli; cauliflower; brussels sprouts; onion; garlic; leek; squash; pumpkin; celery; pea; bean (including various legumes); strawberries; grapes; apples; cherries; pears; peaches; banana; palm; cocoa; cucumber; pineapple; apricot; plum; sugar beet; lawn grasses; maple; teosinte; Tripsacum; Coix; triticale; safflower; peanut; cassava, and olive.
[0321] The invention also provides an agricultural product produced by any of the described transgenic plants, plant parts, and plant seeds. Agricultural products include, but are not limited to, plant extracts, proteins, amino acids, carbohydrates, fats, oils, polymers, vitamins, and the like.IV. Methods
[0322] A further aspect of the present disclosure encompasses a method of genetic modification of genomic nucleic acid sequences of a cell, including targeted insertion of nucleic acid sequence into a target nucleic acid locus in a cell. In a method of the instant disclosure, the cell can be ex vivo or in vivo. The locus can be in a chromosomal DNA, organellar DNA, or extrachromosomal DNA. The method can beused to insert a single donor polynucleotide or more than one donor polynucleotide at one or more target loci.
[0323] The method comprises providing or having provided an engineered system for generating a genetically modified cell or one or more nucleic acid constructs for expressing the engineered system and introducing the system or the expression constructs for expressing the engineered system into the cell. The method further comprises maintaining the cell under appropriate conditions such that the donor polynucleotide is inserted in the target locus. Optionally, the method further comprises identifying an accurate insertion of the donor polynucleotide in the nucleic acid locus. The engineered system can be as described in Section I; nucleic acid constructs encoding one or more components of the engineered system can be as described in Section II; and the cells can be as described in Section III.
[0324] It will be recognized that the components of the engineered system described herein are modular and do not need to be introduced into a cell simultaneously or as a single unit. In some aspects, one or more components of the system can already be present in the cell prior to the introduction of additional components necessary for the system to function. For example, a cell can already contain a donor polynucleotide flanked by transposition sequences, and the remaining components, such as the nuclease-deficient transposase and programmable targeting nuclease, can be introduced later to facilitate the desired genetic modification. Alternatively, the programmable targeting nuclease can be introduced first to prepare the target nucleic acid locus, followed by the introduction of the donor polynucleotide and transposase to complete the insertion process.
[0325] In other aspects, the components of the system can be introduced sequentially to optimize the timing and efficiency of the genetic modification process. For instance, the transposase can be introduced initially to bind the transposition sequences of the donor polynucleotide, followed by the introduction of the programmable targeting nuclease to excise the donor polynucleotide and guide its insertion into the target locus. Similarly, the donor polynucleotide can be introduced into the cell first, allowing it to integrate into the cellular environment before the transposase and targeting nuclease are introduced to initiate the modification process.
[0326] The modularity of the system also allows for flexibility in delivery methods. For example, one or more components can be delivered via plasmid vectors, viral vectors, or synthetic oligonucleotides, while other components can be delivered as preassembled protein complexes or RNA molecules. The components can be distributed across multiple vectors, with each vector encoding a specific component of the system. This approach can be particularly advantageous for applications requiring precise control over the expression levels or timing of each component.
[0327] Furthermore, the system can be adapted to scenarios where certain components are stably integrated into the genome of the cell or organism, while other components are transiently introduced. For instance, a cell can be genetically engineered to stably express the nuclease-deficient transposase, and the programmable targeting nuclease and donor polynucleotide can be introduced transiently to perform specific genetic modifications. Alternatively, the donor polynucleotide can be stably integrated into the genome, and the transposase and targeting nuclease can be introduced transiently to excise and reinsert the donor polynucleotide at a different locus.
[0328] In some embodiments, the components of the system can be introduced in a staged manner to accommodate specific experimental or therapeutic requirements. For example, in a therapeutic context, the donor polynucleotide can be introduced into a patient’s cells first, followed by the introduction of the transposase and targeting nuclease at a later time to complete the genetic modification. Similarly, in agricultural applications, a plant cell can be transformed with the donor polynucleotide during one stage of development, and the transposase and targeting nuclease can be introduced during a subsequent stage to achieve the desired genetic modification.
[0329] The flexibility of the system also extends to the use of pre-existing cellular components. For instance, if a cell already comprises endogenous transposase or targeting nuclease activity, the system can be designed to leverage these pre-existing components by introducing only the donor polynucleotide and any additional components required to achieve the desired genetic modification. Alternatively, the system can be tailored to suppress or modify the activity of endogenous components to ensure precise and controlled genetic modifications.
[0330] Overall, the modular and adaptable nature of the engineered system allows for a wide range of configurations and delivery strategies, enabling its use in diverse applications across research, therapeutic, and agricultural contexts. The ability to introduce components sequentially, transiently, or stably, and to leverage pre-existing cellular components, provides significant flexibility in designing and implementing genetic modification strategies tailored to specific needs.
[0331] Insertion of the donor polynucleotide into a target nucleic acid locus in a cell can have a number of uses known to individuals of skill in the art. For instance, insertion of the donor polynucleotide can introduce cargo nucleic acid sequences of interest into nucleic acid sequences in a cell, including genes of interest or regulatory nucleic acid sequences of interest. Alternatively, insertion of a donor polynucleotide can be used to introduce nucleic acid modifications in nucleic acid sequences in the cell. Non-limiting examples of nucleic acid modifications that can be introduced in nucleic acid sequences in a cell include point mutations, partial sequence deletions, replacements, or additions, ribosomal skipping sequences, antibody epitopes and tags such as AcV5, AU1 , AU5, E, ECS, E2, FLAG, Glu-Glu, HSV, KT3, myc, S, S1 , T7, V5, VSV-G, and 6xHis and variants thereof, TAP tag, recombinase recognition sites, gene expression regulatory sequences, spacers, capture sequences, small RNA target sites, miRNA trigger sites, tasiRNA sequences.
[0332] The system can also be used to modulate transcriptional or post- transcriptional expression of an endogenous nucleic acid sequence in the cell, to investigate RNA-protein interactions, or to determine the function of a protein or RNA, or investigate RNA-protein interactions, or to alter the stability, accumulation, and protein production from the RNA.
[0333] In general, cargo nucleic acid sequences can be introduced into a nucleic acid sequence of a cell by flanking the nucleic acid sequence to be introduced with the transposition sequences compatible with the transposase. Introduced cargo nucleic acid sequences can include, without limitation, nucleic acid sequences encoding herbicide resistance, disease resistance such as viral coat proteins and R gene families, insect resistance such as Bt toxin genes, antibiotic resistance, short RNAs, reporters, programmable nucleic acid-modification systems, epigenetic modification systems, regulatory elements, viral vectors, agronomic traits of interest such drought and salinityresistance, and any combination thereof. Non-limiting examples of cargo nucleic acid sequences include Bt toxin genes (Cry Genes), RNAi (RNA Interference) constructs, pathogen-derived resistance genes, R gene families, herbicide resistance genes, nitrogen fixation genes (Modulation Genes), drought tolerance genes, salinity tolerance genes, cold tolerance genes, vitamin and nutrient enrichment genes, fruit ripening control genes, photosynthetic efficiency genes, flower color modification genes, plant growth regulator genes, phytoremediation genes, altered oil or protein content genes, biofortification genes, and aroma and flavor enhancement genes.
[0334] In some aspects, a method of the instant disclosure comprises altering expression of a gene of interest. The method comprises introducing expression regulatory elements to a location on the genome where expression of a gene of interest is controlled. In some aspects, the regulatory elements are heat shock enhancer elements. In some aspects, the method comprises introducing an array of six heatshock enhancer elements flanked by the mPing transposition sequences for insertion into the promoter of the Arabidopsis ACT8 gene. These enhancers have a short size and regulate expression of the gene irrespective of the orientation of the introduced sequences. Donor constructs comprising heat-shock enhancer elements flanked by the mPing transposition sequences can be as described in Sections l(b) and Section II.
[0335] In some aspects, a method of the instant disclosure is used to introduce a herbicide resistance gene. Non-limiting examples of genes that can be used in cargo nucleic acids of the instant disclosure to introduce herbicide resistance include EPSPS (5-Enolpyruvylshikimate-3-Phosphate Synthase) that can provide resistance to glyphosate herbicides such as Roundup, PAT (Phosphinothricin Acetyltransferase) that can confer resistance to glufosinate herbicides, including Liberty and Basta, modified ALS (Acetolactate Synthase) genes that can confer resistance to sulfonylurea and imidazolinone herbicides, BAR (Bialaphos Resistance) that can provide resistance to herbicides like Bialaphos and phosphinothricin (the active ingredient in glufosinate herbicides), modified ACCase (Acetyl-CoA Carboxylase) genes that can provide resistance to ACCase-inhibiting herbicides, such as clethodim and sethoxydim, modified PPG (Protoporphyrinogen Oxidase) genes that can provide resistance to saflufenacil, GST (Glutathione S-Transferase) genes that can be used to enhance the plant's ability to detoxify a range of herbicides by conjugating them with glutathione, rendering themless toxic, Vip3A (Vegetative Insecticidal Protein) gene that can confer resistance to some herbivorous insects that damage crops alongside herbicide resistance, modified HPPD (4-Hydroxyphenylpyruvate Dioxygenase) genes that can confer resistance to certain herbicides, like mesotrione, inhibit the HPPD enzyme, AAD-12 (Aryloxyalkanoate Dioxygenase-12) gene that can provide resistance to 2,4-D herbicides, and DSF (Dinitroaniline Herbicide Resistance). In some aspects, a method of the instant disclosure comprises introducing resistance to bialophos herbicide. In some aspects, a method of the instant disclosure comprises introducing a donor construct comprising an expression construct expressing the BAR gene flanked by the mPing transposition sequences into a cell. Donor constructs comprising heat-shock enhancer elements flanked by the mPing transposition sequences can be as described in Sections l(b) and Section II.(a) Introduction into the Cell
[0336] The method comprises introducing the engineered system into a cell of interest. The engineered system can be introduced into the cell as a purified isolated composition, purified isolated components of a composition, as one or more nucleic acid constructs encoding the engineered system, or combinations thereof. Further, components of the engineered system can be separately introduced into a cell. For example, a transposase, a donor polynucleotide, and a programmable targeting nuclease can be introduced into a cell sequentially or simultaneously.
[0337] The engineered system described above can be introduced into the cell by a variety of means. Suitable delivery means include microinjection, electroporation, sonoporation, biolistics, calcium phosphate-mediated transfection, cationic transfection, liposomes and other lipids, dendrimer transfection, heat shock transfection, nucleofection transfection, gene gun delivery, dip transformation, supercharged proteins, cell-penetrating peptides, implantable devices, magnetofection, lipofection, impalefection, optical transfection, proprietary agent-enhanced uptake of nucleic acids, Agrobacterium tumefaciens mediated foreign gene transformation, proprietary agent- enhanced uptake of nucleic acids, and delivery via liposomes, immunoliposomes, virosomes, or artificial virions. The choice of means of introducing the system into a cellcan and will vary depending on the cell, or the system or nucleic acid constructs encoding the system, among other variables.(b) Culturing a Cell
[0338] The method further comprises maintaining the cell under appropriate conditions such that the donor polynucleotide is inserted in the target locus. When the cell is in tissue ex vivo, or in vivo within an organism or within a tissue of an organism, the tissue and / or organism can also be maintained under appropriate conditions for insertion of the donor polynucleotide. In general, the cell is maintained under conditions appropriate for cell growth and / or maintenance. Those of skill in the art appreciate that methods for culturing cells are known in the art and can and will vary depending on the cell type. Routine optimization can be used, in all cases, to determine the best techniques for a particular cell type. See for example, in Santiago et al. (2008) PNAS 105:5809-5814; Moehle et al. (2007) PNAS 104:3055-3060; Umov et al. (2005) Nature 435:646-651 ; and Lombardo et al. (2007) Nat. Biotechnology 25:1298-1306; Taylor et al., (2012) Tropical Plant Biology 5: 127-139.
[0339] In some aspects, the method further comprises identifying an accurate insertion of the donor polynucleotide using methods known in the art. Upon confirmation that an accurate insertion has occurred, single cell clones can be isolated. Additionally, cells comprising one accurate insertion can undergo one or more additional rounds of targeted insertions of additional polynucleotides.(c) Method of generating a genetically modified cell comprising no off-target insertion
[0340] Another aspect of the instant disclosure encompasses a method of generating a genetically modified cell comprising no or reduced off-target insertion. The method comprises the steps of: (a) introducing an engineered nucleic acid modification system for generation of a genetically modified cell comprising no off-target insertion or one or more nucleic acid constructs for expressing the engineered nucleic acid modification system; (b) maintaining the cell under conditions and for a time sufficient for the one or more nucleic acid constructs for expressing the transposase and programmable targeting nuclease to express the transposase and programmabletargeting nuclease, for the donor polynucleotide to be inserted in the target locus to generate a genetically modified cell comprising the donor polynucleotide inserted in the target locus; (c) concurrently with or after step (b), maintaining the genetically modified cell under conditions and for a time sufficient for generating a genetically modified cell wherein the one or more nucleic acid constructs for expressing the transposase and programmable targeting nuclease are absent. In some aspects, the method further optionally comprises identifying an insertion of the donor polynucleotide in the nucleic acid locus in the cell. In some aspects, an inserted donor polynucleotide replaces a target locus.
[0341] In some aspects, the cell is a eukaryotic cell. In some aspects, the eukaryotic cell is a plant cell, a plant or part thereof, or seed. In some aspects, the cell is ex vivo.V. Kits
[0342] A further aspect of the present disclosure encompasses kits for generation of a non-transgenic genetically modified cell. The kit comprises one or more engineered systems detailed above in Section l(e). The engineered systems can be encoded by a system of one or more nucleic acid constructs encoding the components of the system as described herein above in Section ll(b). Alternatively, the kit can comprise one or more cells comprising one or more engineered systems of the instant disclosure, one or more components of an engineered system, one or more nucleic acid constructs, or combinations thereof.
[0343] A further aspect of the present disclosure provides a system of one or more nucleic acid constructs encoding the components of the system described above.
[0344] The kits can further comprise transfection reagents, cell growth media, selection media, in-vitro transcription reagents, nucleic acid purification reagents, protein purification reagents, buffers, and the like. The kits provided herein generally include instructions for carrying out the methods detailed below. Instructions included in the kits can be affixed to packaging material or can be included as a package insert. While the instructions are typically written or printed materials, they are not limited to such. Any medium capable of storing such instructions and communicating them to an end user is contemplated by this disclosure. Such media include, but are not limited to,electronic storage media (e.g., magnetic discs, tapes, cartridges, chips), optical media (e.g., CD ROM), an internet address that provides the instructions, and the like. As used herein, the term “instructions” can include the address of an internet site that provides the instructions.DEFINITIONS
[0345] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991 ); and Hale & Marham, The Harper Collins Dictionary of Biology (1991 ). As used herein, the following terms have the meanings ascribed to them unless specified otherwise.
[0346] When introducing elements of the present disclosure or the aspects(s) thereof, the articles "a", "an", "the" and "said" are intended to mean that there are one or more of the elements. The terms "comprising", "including" and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.
[0347] As used herein, the term "gene" refers to a DNA region (including exons and introns) encoding a gene product, as well as all DNA regions which regulate the production of the gene product, whether or not such regulatory sequences are adjacent to coding and / or transcribed sequences. Accordingly, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment sites, and locus control regions.
[0348] A “genetically modified” cell refers to a cell in which the nuclear, organellar or extrachromosomal nucleic acid sequences of a cell has been modified, i.e. , the cell contains at least one nucleic acid sequence that has been engineered to contain aninsertion of at least one nucleotide, a deletion of at least one nucleotide, and / or a substitution of at least one nucleotide.
[0349] The terms “genome modification” and “genome editing” refer to processes by which a specific nucleic acid sequence in a genome is changed such that the nucleic acid sequence is modified. The nucleic acid sequence may be modified to comprise an insertion of at least one nucleotide, a deletion of at least one nucleotide, and / or a substitution of at least one nucleotide. The modified nucleic acid sequence is inactivated such that no product is made. Alternatively, the nucleic acid sequence may be modified such that an altered product is made.
[0350] As used herein, the term “compatible transposition sequences” refers to any transposition sequences recognized by the transposase for transposition. For instance, the transposition sequences can be transposition sequences of the TE from which the transposase is derived, or from another autonomous or non-autonomous TE recognized by the transposase for transposition.
[0351] As used herein, the term “engineered” when applied to a targeting protein refers to targeting proteins modified to specifically recognize and bind to a nucleic acid sequence at or near a target nucleic acid locus. A “genetically modified” plant refers to a cell in which the nuclear, organellar or extrachromosomal nucleic acid sequences of a cell have been modified, i.e. , the cell contains at least one nucleic acid sequence that has been engineered to contain an insertion of at least one nucleotide, a deletion of at least one nucleotide, and / or a substitution of at least one nucleotide.
[0352] The term “nucleic acid modification” refers to processes by which a specific nucleic acid sequence in a polynucleotide is changed such that the nucleic acid sequence is modified. The nucleic acid sequence may be modified to comprise an insertion of at least one nucleotide, a deletion of at least one nucleotide, and / or a substitution of at least one nucleotide. The modified nucleic acid sequence is inactivated such that no product is made. Alternatively, the nucleic acid sequence may be modified such that an altered product is made.
[0353] As used herein, “protein expression” includes but is not limited to one or more of the following: transcription of a gene into precursor mRNA; splicing and other processing of the precursor mRNA to produce mature mRNA; mRNA stability; translation of the mature mRNA into protein (including codon usage and tRNAavailability); production of a mutant protein comprising a mutation that modifies the activity of the protein, including the calcium channel activity; and glycosylation and / or other modifications of the translation product, if required for proper expression and function. The term "heterologous" refers to an entity that is not native to the cell or species of interest.
[0354] The terms “nucleic acid” and “polynucleotide” refer to a deoxyribonucleotide or ribonucleotide polymer, in linear or circular conformation. For the purposes of the present disclosure, these terms are not to be construed as limiting with respect to the length of a polymer. The terms may encompass known analogs of natural nucleotides, as well as nucleotides that are modified in the base, sugar and / or phosphate moieties. In general, an analog of a particular nucleotide has the same basepairing specificity, i.e., an analog of A will base-pair with T. The nucleotides of a nucleic acid or polynucleotide may be linked by phosphodiester, phosphothioate, phosphoram idite, phosphorodiamidate bonds, or combinations thereof.
[0355] The term "nucleotide" refers to deoxyribonucleotides or ribonucleotides. The nucleotides may be standard nucleotides (i.e., adenosine, guanosine, cytidine, thymidine, and undine) or nucleotide analogs. A nucleotide analog refers to a nucleotide having a modified purine or pyrimidine base or a modified ribose moiety. A nucleotide analog may be a naturally occurring nucleotide (e.g., inosine) or a non-naturally occurring nucleotide. Non-limiting examples of modifications on the sugar or base moieties of a nucleotide include the addition (or removal) of acetyl groups, amino groups, carboxyl groups, carboxymethyl groups, hydroxyl groups, methyl groups, phosphoryl groups, and thiol groups, as well as the substitution of the carbon and nitrogen atoms of the bases with other atoms (e.g., 7-deaza purines). Nucleotide analogs also include dideoxy nucleotides, 2’-O-methyl nucleotides, locked nucleic acids (LNA), peptide nucleic acids (PNA), and morpholinos.
[0356] The terms “polypeptide” and “protein” are used interchangeably to refer to a polymer of amino acid residues.
[0357] As used herein, the terms "target site", "target sequence", or “nucleic acid locus” refer to a nucleic acid sequence that defines a portion of a nucleic acid sequence to be modified or edited and to which a homologous recombination composition is engineered to target.
[0358] The terms "upstream" and "downstream" refer to locations in a nucleic acid sequence relative to a fixed position. Upstream refers to the region that is 5' (i.e. , near the 5' end of the strand) to the position, and downstream refers to the region that is 3' (i.e., near the 3' end of the strand) to the position.
[0359] As used herein, the term “encode” is understood to have its plain and ordinary meaning as used in the biological fields, i.e., specifying a biological sequence. For instance, when a construct is encoding a protein of the system, the term is understood to mean that the construct further comprises nucleic acid sequences required for expressing the components of the system.
[0360] As various changes could be made in the above-described cells and methods without departing from the scope of the invention, it is intended that all matter contained in the above description and in the examples given below, shall be interpreted as illustrative and not in a limiting sense.EXAMPLES
[0361] All patents and publications mentioned in the specification are indicative of the levels of those skilled in the art to which the present disclosure pertains. All patents and publications are herein incorporated by reference to the same extent as if each individual publication was specifically and individually indicated to be incorporated by reference.
[0362] The publications discussed throughout are provided solely for their disclosure before the filing date of the present application. Nothing herein is to be construed as an admission that the invention is not entitled to antedate such disclosure by virtue of prior invention.
[0363] The following examples are included to demonstrate the disclosure. It should be appreciated by those of skill in the art that the techniques disclosed in the following examples represent techniques discovered by the inventors to function well in the practice of the disclosure. Those of skill in the art should, however, in light of the present disclosure, appreciate that many changes could be made in the disclosure and still obtain a like or similar result without departing from the spirit and scope of the disclosure, therefore all matter set forth is to be interpreted as illustrative and not in a limiting sense.Example 1. Targeted integration of a transposable element
[0364] Transgenesis in plants is accomplished via bombardment or agrobacterium-mediated transformation and results in the integration of foreign DNA into a plant’s genome. During this process, the transgene integration site within the plant DNA is not controlled, and follow-up experiments must be performed to determine where in the genome the transgene integrated. En mass transformation experiments have demonstrated that the integration typically occurs at sites of open chromatin configuration, such as actively transcribing genes, however integration into heterochromatic closed chromatin can also occur. Transgene integration into or near genes can generate new mutations or alter the regulation of nearby genes, while insertions into heterochromatic regions are often not permissive to the desired high levels of transgene expression or do not provide stable expression over multiple generations. Insertion of transgenes is also associated with mutations (deletions and rearrangements) of the target region and transferred DNA. In addition, to study or create a product from a gene of interest, it needs to be taken out of its native context and added back to the plant as a transgene, and key distal regulatory enhancers or repressor elements can be missed or rearranged during this process. The lack of user- defined control of transgene integration site generates variability and inconsistency in experiments and products.
[0365] The control of transgene integration site is desired to direct transgenes to the same expression-permissive regions of the genome (to reduce variability), to add sequences to genes at their native locations, and / or to maintain gene order on the chromosome. Multiple attempts have been made to overcome these issues and perform target site-directed integration. The FLP-FRT recombination system has been used to reproducibly target transgene insertion into one location in plant genomes. However, this insertion site must also be transgenic to carry the correct targeting sequences. Current methods to insert DNA into any user-defined targeted region of a plant genome involve homology-directed repair (HDR) off a provided DNA template after a doublestrand DNA break induced by a Meganuclease, Zinc Finger Nuclease, TALEN or CRISPR / Cas9 (or related) system. In plants, currently available tools using targeted insertion of a transgene via HDR are inefficient for two reasons. First, thecomplementary repair template and nuclease system must be added to the cell via traditional transgenesis, which particularly in crop plants is laborious. Second, plant cells favor the resolution of double-strand DNA breaks by the non-homology end joining (NHEJ) pathway, which bypasses the integration of new DNA.
[0366] Recently, research has uncovered naturally-occurring fusions between transposase proteins and the CRISPR / Cas system in prokaryotes. The CRISPR / Cas system provides sequence specificity to the transposase for selection of the integration site, and was proven to be programmable by altering the sequence of the CRISPR guide RNA (gRNA). However, none of the systems currently available that use CRISPR-targeting of a transposase protein were successful in targeting to a specific gene location in eukaryotic cells. To date, the programmability of transposase-mediated integration of DNA has not been accomplished in a eukaryote.
[0367] In an attempt to overcome the difficulties in accomplishing insertion of a transgene into a target locus, the inventors linked a TE-encoded transposase protein to the CRISPR / Cas9 system to achieve targeted integration of DNA in plants. The inventors reasoned that the transposase protein would need to have two features to broadly function in this system. First, a wide host-range of functionality in plants was desired to create a universal tool for plant biology. Second, using split-transposase proteins (where the single transposase was encoded by two proteins that function together to achieve excision and insertion) would have a lower probability of disturbing protein function. It was reasoned that the rice mPing / Pong system would provide the highest probably of functioning when linked to Cas9, as the Pong transposase is split into two proteins (ORF1 and ORF2) and can mobilize the mPing non-autonomous (nonprotein coding) TE in a range of plant species. An mPing / Pong engineered system was used that had the Pong transposase ORF1 and ORF2 immobilized by the removal of the Pong TIRs. In this system, mPing excision can be visualized by its removal from a constitutively expressed GFP gene (FIG. 1). The Pong ORF1 / ORF2 system was engineered with the G4S (GSSSS) flexible protein linker to allow efficient fusions to Cas9 proteins on either the N- or C-terminus of ORF1 or ORF2, and an SV40 nuclear localization signal (NLS) was added to these protein fusions. Three versions of the Cas9 protein were used, the catalytically active Cas9, the single-stranded nickase deCas9, and the catalytically inactive dCas9. A total of 12 constructs were generated (3Cas9 proteins x 4 ORF1 / ORF2 positions; FIG. 2) with a gRNA known to target the Arabidopsis PDS3 gene.
[0368] To determine if the Pong transposase was functional when linked to Cas9 derivatives, GFP fluorescence was visualized in seedlings. GFP fluorescence is a marker of mPing excision from the GFP donor site, and this fluorescence was detected for all 12 fusion proteins, but not the negative control without ORF1 / ORF2 (FIG. 3A), verifying that ORF1 and ORF2 are co-creating a functional transposase protein even while linked to Cas9. A functional CRISPR / Cas9 system was verified through the observation of white seedlings and sectors in plants with the Cas9 and deCas9 proteins (in this experiment, dCas9 plants did not display white plants or sectors) (FIG. 3B). Overall, the results demonstrate that fusion of the Cas9 and transposase proteins does not stop their function.
[0369] A PCR amplification strategy was used to detect targeted mPing insertions into the Arabidopsis PDS3 gene (FIG. 4A). T2 seedling pools were screened using negative control lines that either lack ORF1 / ORF2, or that lack the Cas9 fusion (FIG. 4B). It was found that clone #2 displayed the correct size PCR band in all PCR assays (FIG. 4B). The PCR can identify mPing insertions in the forward or reverse orientation (FIG. 4A), and the fact that clone #2 amplified for both suggests that there is more than one mPing insertion in this pool of plants. Clone #2 encodes for ORF1 + ORF2-Cas9, where ORF2 has a C-terminal fusion to the Cas9 protein. This data demonstrates targeted insertion of mPing into the PDS3 gene using a targeting nuclease having full double stranded cleavage activity of Cas9.Example 2. Characterization of target site insertions
[0370] The target-site PCR assay was replicated (FIG. 4C), and PCR products cloned and sequenced. In all, 36 clones were sequenced. The sequenced clones represent at least nine (9) unique targeted transposition events (FIG. 5). Both mPing forward and reverse orientation insertions were identified, demonstrating the random directionality of the targeted insertion event.
[0371] The targeted insertion occurred between the third and fourth base of the gRNA target sequence, as expected based on the known cleavage activity of Cas9 (FIG. 5). The results show that mPing is intact in each sequenced clone except one. Ineach case there is one target site duplication, on either the 5’ or 3’ of mPing. Additional single-base insertions are found in some clones. The sequencing represents at least nine distinct events, meaning that mPing inserted into the PDS3 gene in the line with clone #2 at least nine different times. Most insertions have either intact or partial TTA I TAA sequence on only one end of the insertion. This sequence originates from the donor site and is part of the known target site duplication (TSD) of the Pong / mPing TE system. The presence of only one TSD, rather than one on either side of the TE insertion, signifies that Cas9 created a blunt cut at the insertion site, but the transposase protein made a staggered cut at the donor site before the integration event. This demonstrates that both the Cas9 and transposase proteins are functional for generating this set of insertions.
[0372] For each insertion, the gRNA target sequence was preserved and mPing had inserted at the expected Cas9 cleavage point between the third and fourth nucleotide. In all but one sequence read the mPing element is complete, with only single base insertions. The lack of deletions or other insertions at these insertion sites demonstrates the seamless repair of the insertion events by the transposase protein compared to typical sites of blunt-end DNA breaks.Example 3. Integration into any DNA break
[0373] Several previous reports have demonstrated that transgenes will insert at a low frequency into any site of double-strand break. To determine if the mPing targeted insertion detected in Examples 1 and 2 requires the transposase protein, a PCR assay was performed for the integration of the transgene backbone encoding the ORF2-Cas9 protein into the DNA break generated at PDS3. It was reasoned that if the mPing insertion into PDS3 was a product of transgene insertion, rather than transposition, it would be equally likely to detect other parts of the transgene at this insertion site location. However, transgene was detected at PDS3 (FIG. 6A), demonstrating that mPing insertion requires the transposase to excise the mPing element from the donor position.
[0374] Next, it was assayed whether it was essential that the transposase protein and Cas9 were directly linked, or if both proteins unlinked in the same cell could perform targeted insertion. It was discovered that in some cases, the two proteins could beunlinked and targeted insertion would take place (FIG. 6B). At the same time, it was demonstrated that both proteins are functional and that in this instance, the catalytic activity of Cas9 is used (FIG. 6B). Together, this data demonstrates that to obtain targeted insertion, it is essential that the transposase excise the element out of the donor position, and that Cas9 cleave the insertion site, but the two proteins do not necessarily need to be linked together (see FIGs. 8A and 8B and Example 5).Example 4. Programmability of target sites
[0375] Multiple sites in the Arabidopsis genome were targeted using the system of the instant disclosure. Two additional gRNAs were designed for integration into two additional target loci; the ADH1 gene and a non-coding region upstream of the ACT8 gene of Arabidopsis. The gRNAs were used in a system described herein to integrate mPing into the two target loci (FIG. 7A). FIG. 7B shows the Sanger sequencing results of junctions of each identified target insertion into the PDS3 gene, the ADH1 gene, and the promoter of ACT8 gene. The chromatograms above the sequence show the sequences at the insertion sites. The sequences below mPing are the expected sequence if a perfect “seamless” insertion is obtained. These results clearly confirm that the insertion of a donor polynucleotide is surprisingly and unexpectedly inserted on target and unexpectedly accurate and seamless.Example 5. Direct Fusion of the transposase proteins ORF1 and ORF2 to the nuclease is not required for targeted insertions
[0376] Using methods described in Example 3, whether a system wherein the transposase proteins ORF1 and ORF2 are not directly linked to the Cas9 nuclease was tested. FIG. 8A shows that mPing can be targeted to the Arabidopsis PDS3 gene by the CRISPR gRNA and can insert in either the forward direction (above the PDS3 region) or reverse direction (below the PDS3 region). A combination of 2 out of 4 PCR primers corresponding to the PDS3 exon (U,D) and the mPing gene (R, L) were used. FIG. 8A shows the location of these 4 PCR primers (R,L,U,D) for orientation.
[0377] The mPing targeted insertion was detected with PCR using the primer sets from part A. FIG. 8B shows a representative agarose gel with PCR products observed. Arrowheads denote the correct size of the PCR products for each set ofprimers. “mPing only”, “+ORF1 / 2” and “+Cas9” are negative controls. Any bands from these lanes near the correct size were sequenced and shown not to be specific targeted insertions of mPing. The bands shown in the “+unlinked ORF1 / 2 and Cas9” lane show that using unlinked constructs can generate real targeted insertions, as does the biological replicate of ORF2 linked to Cas9 in the “ORF1 / ORF2-Cas9” lane. All PCR products from this assay were also verified by Sanger sequencing. These data confirm the results from FIG. 6B and demonstrate that direct fusion of the transposase proteins to the nuclease is not required for targeted insertions.Example 6: Targeted insertion driven by single transgene vector
[0378] In the previously described experiments, the system comprised a donor construct and a helper construct. Here, a single transgene vector was developed containing all the elements required for targeted insertion in a plant cell. The vector is diagrammed in FIG. 9A and contains the CRISPR / Cas9 system (including gRNA), the mPing donor element, and ORF1 and ORF2 transposase proteins.
[0379] Using methods described in the examples above, mPing was targeted to the Arabidopsis PDS3 gene by the CRISPR gRNA. As shown in FIG. 9B, mPing can insert in either the forward direction (above the PDS3 region) or reverse direction (below the PSD3 region). The location of 4 PCR primers (R, L, U, D) are shown for orientation. FIG. 9C shows a representative agarose gel with PCR detection of mPing targeted insertion in the Arabidopsis genome using the primer sets from part B. The largest PCR fragment for each primer set is the correct size and was Sanger sequenced to ensure that it is a bonafide targeted insertion of mPing into the PDS3 gene.Example 7: Targeted and seamless integration in plant genomes using CRISPR- transposasesIntroduction
[0380] Transgenesis in plants is accomplished via bombardment or agrobacterium-mediated transformation and results in the integration of foreign DNA into a plant’s genome. During this process, the transgene integration site within the plant DNA is not controlled, and follow-up experiments must be performed to determine where in the genome the transgene integrated. En mass transformation experimentshave demonstrated that the integration typically occurs at sites of open chromatin configuration, such as actively transcribing genes, however integration into heterochromatic closed chromatin can also occur. Transgene integration into or near genes can generate new mutations or alter the regulation of nearby genes, while insertions into heterochromatic regions are often not permissive to the desired high levels of transgene expression or do not provide stable expression over multiple generations. Insertion of transgenes is also associated with mutations (deletions and rearrangements) of the target region and transferred DNA. In addition, to study or create a product from a gene of interest, it needs to be taken out of its native context and added back to the plant as a transgene, and key distal regulatory enhancers or repressor elements can be missed or rearranged during this process. The lack of user- defined control of transgene integration site generates variability and inconsistency in experiments and products.
[0381] The control of transgene integration site is desired to direct transgenes to the same expression-permissive regions of the genome (to reduce variability), to add sequences to genes at their native locations, and / or to maintain gene order on the chromosome. Multiple attempts have been made to overcome these issues and perform targeted site-directed integration. Recombination systems have been used to reproducibly target transgene insertion into one location in plant genomes, however, this insertion site must also be transgenic to carry the correct targeting sequences. Current methods to insert DNA into any user-defined targeted region of a plant genome involve homology-directed repair (HDR) off a provided DNA template after a double-strand DNA break induced by a Meganuclease, Zinc Finger Nuclease, TALEN or CRISPR / Cas9 (or related) system. In plants, targeting insertion of a transgene via HDR is inefficient for two reasons. First, the complementary repair template and nuclease system must be added to the cell via traditional transgenesis, which particularly in crop plants is laborious. Second, plant cells favor the resolution of double-strand DNA breaks by the non-homology end joining (NHEJ) pathway, which bypasses the integration of new DNA. Therefore, addition of custom sequences to a targeted location in a plant genome is laborious, requiring screening for a low-frequency event. In addition, because free ends of DNA are exposed during this process, the ends of the inserted fragment of DNAor the native DNA at the insertion site is often subject to degradation, creating deletions and unintended base changes at the HDR site.
[0382] Transposases are transposable element (TE)-derived proteins that naturally mobilize pieces of DNA from one location in the genome to another. Transposases function by binding the repeated ends of a TE called the terminal inverted repeats (TIRs) within the same TE family. The transposase cleaves the DNA, removing the TE from the excision / donor site, then cleaves and integrates the TE at the insertion site. Plant transposases select their insertion site by chromatin context and DNA accessibility but are not targeted to individual regions or specific sequences of plant genomes. Recently, research has uncovered naturally-occurring fusions between transposase proteins and the CRISPR / Cas system in prokaryotes. The CRISPR / Cas system provides sequence specificity to the transposase for selection of the integration site, and was proven to be programmable by altering the sequence of the CRISPR guide RNA (gRNA). Several laboratories have taken the approach to identify natural Cas protein fusions to transposable elements in prokaryotic genomes, with the intent of moving these fusion proteins into eukaryotes. In human cell culture, CRISPR-targeting of a transposase protein has been attempted but failed to target to a specific gene location, although the integration into targeted repetitive retrotransposon sites were enriched. The inventors took the approach of starting with a transposase protein known to work in a wide variety of plants, and Cas9 and CFP1 , which have also been shown to work in plants. Rather than identifying a natural fusion in a prokaryotic genome, both of these proteins were artificially used at the same time, including fusing these proteins together, to accomplish targeted insertion in a plant genome. An overview of this process is shown in FIG. 10.ResultsTargeted integration of a transposable element
[0383] The goal was to fuse a TE-encoded transposase protein to the CRISPR / Cas9 system to achieve targeted integration of DNA in plants. The reason lies in that the transposase protein would need to have two features to broadly function in this system. First, a wide host-range of functionality in plants was desired to create a universal tool for plant biology. Second, using split-transposase proteins (where thesingle transposase was encoded by two proteins that function together to achieve excision and insertion) would have a lower probability of disturbing protein function. It was reasoned that the rice mPing / Pong system would provide the highest probably of functioning when linked to Cas9, as the Pong transposase is split into two proteins (ORF1 and ORF2) and can mobilize the mPing non-autonomous (non-protein coding) TE in a range of plant species. mPing / Pong engineered system was obtained where the Pong transposase ORF1 and ORF2 were immobilized by the removal of the Pong TIRs, and mPing excision can be visualized by its removal from a constitutively expressed GFP gene (cartoons in FIG. 11). The Pong ORF1 / ORF2 system was engineered with the G4S (GSSSS; SEQ ID NO: 64) flexible protein linker to allow efficient fusions to Cas9 proteins on either the N- or C-terminus of ORF1 or ORF2 and added an SV40 nuclear localization signal (NLS) to these protein fusions. Three versions of the Cas9 protein where used, the catalytically active Cas9, the single-stranded nickase deCas9, and the catalytically inactive dCas9. A total of 12 constructs were generated (3 Cas9 proteins x 4 ORF1 / ORF2 positions) (FIG. 11 ) with a gRNA known to target the Arabidopsis PDS3 gene (https: / / doi.org / 10.1038 / nbt.2655).
[0384] To determine if the Pong transposase was functional when linked to Cas9 derivatives, mPing excision from the donor site within GFP was assayed by visualizing the GFP fluorescence of seedlings (FIG. 12A and FIG. 13A). GFP fluorescence is a marker of mPing excision from the GFP donor site, and this fluorescence was detected for all 12 fusion proteins, but not the negative control without ORF1 / ORF2 (summarized in FIG. 12A, full data in FIG. 13A), verifying that ORF1 and ORF2 are co-creating a functional transposase protein even while linked to Cas9. The function of the transposase was additionally verified using a PCR assay to detect mPing excision from the donor site. mPing excises out of its donor position when the transposase is linked to Cas9 (FIG. 12B), although the frequency may be decreased compared to transposase proteins with no fusion (FIG. 12B). A functional CRISPR / Cas9 system was verified through the observation of white seedlings and sectors in plants with the Cas9 proteins (dCas9 plants did not display white plants or sectors) (FIG. 13B). These whi...
Claims
CLAIMSWhat is claimed is:1 . An engineered nucleic acid modification system for generation of a genetically modified cell, the system comprising: a. one or more nucleic acid constructs for expressing a nuclease-deficient transposase and a programmable targeting nuclease, wherein the one or more nucleic acid constructs comprise: i. an expression construct for expressing the nuclease-deficient transposase, the expression construct comprising a promoter operably linked to a nucleic acid sequence encoding the nuclease- deficient transposase; and ii. an expression construct for expressing the programmable targeting nuclease, wherein the expression construct comprises a promoter operably linked to a nucleic acid sequence encoding the programmable targeting nuclease; and b. a donor polynucleotide comprising transposition sequences compatible with the nuclease-deficient transposase, wherein the donor polynucleotide is optionally comprised in a source polynucleotide; wherein the programmable targeting nuclease is engineered to excise the donor polynucleotide from the source polynucleotide, wherein the nuclease-deficient transposase recognizes and binds the transposition sequences of the donor polynucleotide, and wherein the programmable targeting nuclease is engineered to introduce a cut in a target nucleic acid locus in the cell thereby guiding insertion of the excised donor polynucleotide at the target nucleic acid locus to generate a genetically modified cell comprising the transposable nucleic acid construct inserted at the target nucleic acid locus.
2. The engineered system of claim 1 , wherein the programmable nuclease-deficient transposase is linked to the programmable targeting nuclease.
3. The engineered system of claim 1 or claim 2, wherein the nuclease-deficient transposase is not linked to the programmable targeting nuclease.
4. The engineered system of any one of the preceding claims, wherein the nuclease-deficient transposase is a nuclease-deficient split transposase.
5. The engineered system of claim 4, wherein the nuclease-deficient split transposase is a Pong or Pong-like split transposase comprising a Pong ORF1 protein and a Pong ORF2 protein and wherein the nuclease deficient Pong ORF2 protein comprises a modification that inactivates the nuclease function of the Pong ORF2 protein.
6. The engineered system of any one of claims 4-5, wherein the nuclease-deficient transposase is a Pong or Pong-like split transposase comprising a Pong ORF1 protein and a nuclease deficient Pong ORF2 protein and wherein the nuclease-deficient Pong ORF2 protein comprises a modification in a nuclease catalytic site of the Pong ORF2 protein that inactivates the nuclease function of the Pong ORF2 protein.
7. The engineered system of any one of claims 4-6, wherein the nuclease-deficient transposase is a Pong or Pong-like split transposase comprising a Pong ORF1 protein and a nuclease-deficient Pong ORF2 protein and wherein the nuclease deficient Pong ORF2 protein comprises a modification in a DDE catalytic site of the nuclease-deficient Pong ORF2 protein that inactivates the nuclease function of the Pong ORF2 protein.
8. The engineered system of any one of claims 4-7, wherein the Pong ORF1 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 1.
9. The engineered system of any one of claims 4-8, wherein the nuclease deficient Pong ORF2 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 121.
10. The engineered system of any one of claims 4-9, wherein the expression construct for expressing the nuclease-deficient transposase comprises: a. an expression construct for expressing the Pong ORF1 protein wherein the expression construct for expressing the Pong ORF1 protein comprisesat least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100; and b. an expression construct for expressing the nuclease deficient Pong ORF2 protein wherein the expression construct for expressing the nuclease deficient Pong ORF2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 120.11 . The engineered system of any one of claims 4-8, wherein the nuclease-deficient split transposase is a Pong or Pong-like split transposase comprising a Pong ORF1 protein and wherein the Pong ORF2 protein is excluded from the system.
12. The engineered system of claim 11 , wherein the engineered system comprises an expression construct for expressing the Pong ORF1 protein wherein the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100 and wherein the Pong ORF2 protein is excluded from the system.
13. The engineered system of any one of the preceding claims, wherein the donor polynucleotide comprises a cargo polynucleotide flanked by the transposition sequences compatible with the nuclease-deficient transposase of the system.
14. The engineered system of any one of the preceding claims, wherein the transposition sequences are transposition sequences of a miniature inverted-repeat transposable element (MITE).
15. The engineered system of claim 14, wherein the MITE is an mPing MITE.
16. The engineered system of claim 15, wherein transposition sequences of the mPing MITE comprise mPing inverted repeat 1 and mPing inverted repeat 2, wherein mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 7, SEQ ID NO: 111 , or SEQ ID NO: 108, and mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95%or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 8, SEQ ID NO: 112, or SEQ ID NO: 109.
17. The engineered system of claim 15 or claim 16, wherein transposition sequences of the mPing MITE comprise mPing inverted repeat 1 and mPing inverted repeat 2, wherein mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 108, and mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 109.
18. The engineered system of any one of the preceding claims, wherein the programmable targeting nuclease comprises a programmable, sequence-specific nucleic acid-binding domain and a nuclease domain.
19. The engineered system of any one of the preceding claims, wherein the programmable targeting nuclease is an RNA-guided clustered regularly interspersed short palindromic repeats (CRISPR) nuclease system, a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, a ssDNA- guided Argonaute endonuclease, a meganuclease, a rare-cutting endonuclease, or any combination thereof.
20. The engineered system of any one of the preceding claims, wherein the programmable targeting nuclease is a CRISPR-associated (Cas) (CRISPR / Cas) nuclease system comprising a nuclease and a guide RNA (gRNA).21 . The engineered system of claim 20, wherein the programmable targeting nuclease comprises a Cas9 nuclease and a gRNA.
22. The engineered system of claim 21 , wherein cas9 nuclease comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 5.
23. The engineered system of claim 20 or claim 21 , wherein the Cas9 nuclease is linked to the nuclease deficient Pong ORF2 protein.
24. The engineered system of claim 22, wherein the nuclease deficient Pong ORF2 protein is linked to the Cas9 nuclease by one copy of a G4S linker of SEQ ID NO: 64.
25. The engineered system of claim 22, wherein the nuclease deficient Pong ORF2 protein is linked to the Cas9 nuclease by three copies of a G4S linker of SEQ ID NO: 64.
26. The engineered system of claim 20 or claim 21 , wherein the Cas9 nuclease is not linked to the nuclease deficient Pong ORF2 protein.
27. The engineered system of claim 26, wherein the engineered system comprises a nucleic acid expression construct for expressing a Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94.
28. The engineered system of any one of claims 21-27, wherein the engineered system comprises one or more nucleic acid expression constructs for expressing one or more excision gRNAs that guide the Cas9 nuclease to nucleic acid sequences in or flanking the transposition sequences of the donor polynucleotide to guide excision of the donor polynucleotide from the source polynucleotide and for expressing a targeting gRNA that guides the Cas9 nuclease to the target nucleic acid locus to introduce a cut at the target nucleic acid locus, thereby guiding insertion of the excised donor polynucleotide at the target nucleic acid locus to generate a genetically modified cell.
29. The engineered system of claim 28, wherein the one or more excision gRNAs comprise a nucleic acid sequence of SEQ ID NO: 118.
30. The engineered system of any one of claims 28-29, wherein the targeting gRNA that guides the Cas9 nuclease to the target nucleic acid locus to introduce a cut at the target nucleic acid locus comprises a nucleic acid sequence of SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 80, SEQ ID NO: 113, SEQ ID NO: 67 and SEQ ID NO: 113, or any combination thereof.31 . The engineered system of any one of claims 28-30, wherein the one or more nucleic acid expression constructs for expressing the excision and targeting gRNAs comprise a nucleic acid sequence comprises at least about 75% or more, at least about85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 122 and wherein a nucleic acid sequence starting at base 425 to base 444 is replaced with a nucleic acid sequence of a gRNA for targeting the transposase and nuclease to a target nucleic acid locus.
32. The engineered system of any one of claims 27 or 31 , wherein the engineered system comprises: a. an expression construct for expressing the Pong ORF1 protein; b. an expression construct for expressing the Cas9 nuclease; c. one or more expression constructs for expressing an excision gRNAs and a targeting gRNA; and d. a source polynucleotide comprising the donor polynucleotide.
33. The engineered system of claim 32, wherein: a. the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100; b. the expression construct for expressing a Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94; c. the excision gRNA comprises a nucleic acid sequence of SEQ ID NO: 118; and d. the donor polynucleotide comprises mPing inverted repeat 1 and mPing inverted repeat 2, wherein mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 108, and mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75% or more,at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 109.
34. The engineered system of any one of claims 27-31 , wherein the engineered system comprises: a. an expression construct for expressing the Pong ORF1 protein; b. an expression construct for expressing the modified nuclease deficient Pong ORF2 protein; c. an expression construct for expressing the Cas9 nuclease; d. one or more expression constructs for expressing an excision gRNAs and a targeting gRNA; and e. a source polynucleotide comprising the donor polynucleotide.
35. The engineered system of claim 34, wherein: a. the expression construct for expressing the Pong ORF1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 100; b. the expression construct for expressing the nuclease-deficient Pong ORF2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with SEQ ID NO: 120; c. the expression construct for expressing a Cas9 nuclease comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence starting at base 10857 to base 16495 of SEQ ID NO: 94; d. the excision gRNA comprises a nucleic acid sequence of SEQ ID NO: 118; and e. the donor polynucleotide comprises mPing inverted repeat 1 and mPing inverted repeat 2, wherein mPing inverted repeat 1 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% ormore, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 108, and mPing inverted repeat 2 comprises a nucleotide sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 109.
36. The engineered system of any one of the preceding claims, wherein the donor polynucleotide comprises a cargo polynucleotide.
37. The engineered system of claim 36, wherein the cargo polynucleotide comprises HSEs.
38. The engineered system of claim 36, wherein the cargo polynucleotide comprises an expression construct for expressing a herbicide resistance function.
39. The engineered system of claim 38, wherein the herbicide resistance function is resistance to bialaphos herbicide, resistance to glyphosate herbicides, or both.
40. The engineered system of claim 39, wherein the cargo polynucleotide comprises an expression construct for expressing resistance to bialophos herbicide, wherein the expression construct comprises a promoter operably linked to a polynucleotide encoding resistance to bialaphos, and wherein the donor polynucleotide comprising the expression construct for expressing resistance to bialophos herbicide comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97 or SEQ ID NO: 99.41 . The engineered system of claim 39, wherein the cargo polynucleotide comprises an expression construct for expressing resistance to bialophos herbicide, wherein the expression construct comprises a promoter operably linked to a polynucleotide encoding resistance to bialaphos.
42. The engineered system of claim 41 , wherein the donor polynucleotide comprising the expression construct for expressing resistance to bialophos herbicide comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 97.
43. The engineered system of claim 38, wherein the cargo polynucleotide comprises an expression construct for expressing resistance to glyphosate herbicide, wherein the expression construct comprises a promoter operably linked to a polynucleotide encoding resistance to glyphosate.
44. The engineered system of claim 43, wherein the donor polynucleotide comprising the expression construct for expressing resistance to glyphosate herbicide comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with the nucleic acid sequence of SEQ ID NO: 124The engineered system of any one of the preceding claims, wherein the target nucleic acid locus is in a nuclear, organellar, or extrachromosomal nucleic acid sequence.
45. The engineered system of any one of the preceding claims, wherein the target nucleic acid locus is in a protein-coding gene, an RNA coding gene, or an intergenic region.
46. The engineered system of any one of the preceding claims, wherein the cell is a eukaryotic cell.
47. The system of any one of the preceding claims, wherein the cell is a plant cell, a plant or part thereof, or seed.
48. The system of claim 47, wherein the plant is an Arabidopsis sp. or a soybean plant cell, a plant or part thereof, or seed.
49. An engineered nucleic acid modification system for generation of a genetically modified cell, the system comprising: a. one or more nucleic acid expression constructs for expressing a nuclease- deficient transposase and a programmable targeting nuclease, wherein the nuclease-deficient transposase is a Pong or Pong-like split transposase comprising a Pong ORF1 protein and a Pong ORF2 protein and wherein the modified nuclease-deficient Pong ORF2 protein comprises an absence of a Pong ORF2 protein, and wherein the one or more nucleic acid constructs comprise:i. an expression construct for expressing the Pong ORF1 protein, the expression construct comprising a promoter operably linked to a nucleic acid sequence encoding the ORF1 protein; and ii. an expression construct for expressing the programmable targeting nuclease, wherein the expression construct comprises a promoter operably linked to a nucleic acid sequence encoding the programmable targeting nuclease; and b. a donor polynucleotide comprising transposition sequences compatible with the nuclease-deficient transposase, wherein the donor polynucleotide is optionally comprised in a source polynucleotide; wherein the programmable targeting nuclease is engineered to excise the donor polynucleotide from the source polynucleotide, wherein the nuclease-deficient transposase recognizes and binds the transposition sequences of the donor polynucleotide, and wherein the programmable targeting nuclease is engineered to introduce a cut in a target nucleic acid locus in the cell thereby guiding insertion of the excised donor polynucleotide at the target nucleic acid locus to generate a genetically modified cell comprising the transposable nucleic acid construct inserted at the target nucleic acid locus.
50. The engineered system of claim 49, wherein the transposition sequences are transposition sequences of a miniature inverted-repeat transposable element (MITE).51 . The engineered system of claim 50, wherein the MITE is an mPing MITE.
52. The engineered system of any one of claims 49-51 , wherein the programmable targeting nuclease is a CRISPR-associated (Cas) (CRISPR / Cas) nuclease system comprising a nuclease and a guide RNA (gRNA).
53. The engineered system of claim 52, wherein the programmable targeting nuclease comprises a Cas9 nuclease and a gRNA.
54. One or more nucleic acid constructs encoding an engineered nucleic acid modification system of one of claims 1 to 53.
55. A cell comprising the engineered system of one of claims 1 to 48 or one or more nucleic acid constructs of claims 49-53.
56. The cell of claim 55, wherein the cell is a eukaryotic cell.
57. The cell of claim 55, wherein the cell is a plant cell, a plant or part thereof, or seed.
58. A method of generating a genetically modified cell comprising no off-target insertion, the method comprising: a. introducing an engineered nucleic acid modification system for generation of a genetically modified cell of claims 1-49 or one or more nucleic acid constructs of claims 54 to57 into the cell; b. maintaining the cell under conditions and for a time sufficient for the one or more nucleic acid constructs for expressing the nuclease-deficient transposase, expressing the programmable targeting nuclease, for the programmable targeting nuclease to excise the donor polynucleotide from a source polynucleotide, and for the donor polynucleotide to be inserted in the target locus to generate a genetically modified cell comprising the donor polynucleotide inserted in the target locus; c. optionally identifying an insertion of the donor polynucleotide in the nucleic acid locus in the cell; and d. optionally confirming an absence of off-target insertions of the transposable nucleic acid construct.
59. The method of claim 58, wherein an inserted donor polynucleotide replaces a target locus.
60. The method of one of claims 58-59, wherein the cell is a eukaryotic cell.61 . The method of one of claims 58-60, wherein the eukaryotic cell is a plant cell, a plant or part thereof, or seed.
62. The method of one of claims 58-61 , wherein the cell is ex vivo.
63. A kit for generating a genetically modified cell comprising no off-target insertion, the kit comprising one or more engineered nucleic acid modification systems of claims 1-49 or one or more nucleic acid constructs of claims 58-62, wherein each of the engineered systems generates a genetically modified cell comprising an accurate insertion of the donor polynucleotide into the target nucleic acid locus.
64. The kit of claim 63, wherein the kit comprises one or more cells comprising one or more engineered systems, one or more nucleic acid constructs, or any combination thereof.
65. The kit of claim 63 or claim 64, wherein the one or more cells are eukaryotic.
66. The kit of claim 65, wherein the eukaryotic cell is a plant cell, a plant or part thereof, or seed.
Citation Information
Patent Citations
Crispr-associated transposase systems and methods of use thereof
US20230056577A1
Targeted insertion via transportation
US20240150795A1
Targeted insertion via transposition
WO2024098063A2