Nucleic acid-guided TIGR recombinase systems and methods of use
Patent Information
- Application Number
- PCT/US2026/016903
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-26
- Filing Date
- 2026-02-26
- Publication Date
- 2026-09-03
Smart Images

Figure US2026016903_03092026_PF_FP_ABST
Abstract
Description
NUCLEIC ACID-GUIDED TIGR RECOMBINASE SYSTEMS AND METHODS OF USE CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U. S. Provisional Patent Application No. 63 / 763,843 filed February 26, 2025, U. S. Provisional Patent Application No. 63 / 763,865, filed February 26, 2025, and U. S. Provisional Patent Application No. 63 / 763,906, filed February 26, 2025, each of which is incorporated herein by reference in its entirety.REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0002] Reference is made to the electronic sequence listing (“BROD-6055WP_ST26.xml"; Size is 27,221,558 bytes, created on February 26, 2026) is herein incorporated by reference in its entirety.TECHNICAL FIELD
[0003] The subject matter disclosed herein is generally directed to novel RNA guided endonucleases for gene editing.BACKGROUND
[0004] RNA-guided functional systems are versatile in which a protein binds to a guide RNA that directs its activity to a complementary nucleic acid sequence. In nature, RNA-guided systems enable a single enzyme to target multiple sequences depending on the RNA guide loaded. For instance, in bacteria and archaea, CRISPR-Cas adaptive immune systems can recognize and cleave DNA or RNA from mobile genetic elements (MGE) based on RNA guide sequences stored in a CRISPR array, and these arrays can be updated in response to new invaders. In eukaryotes, small RNAs such as microRNAs and PlWI-interacting RNAs (piRNAs) can direct regulatory responses upon interaction with complementary target RNAs. In eukaryotes and archaea, snoRNAs guide methylation and pseudouridylation of complementary RNAs and direct ribosome biogenesis. RNA-guided systems hold immense potential for developing molecular tools in the laboratory due to their easy programmability. Over the last decade, applications of bacterial CRISPR-Cas systems for genome and epigenome editing, as well as RNA editing, regulation, and detection, haverevolutionized basic bioscience, molecular medicine, and biotechnology. Additionally, eukaryotic RNA interference has been employed as a molecular tool for gene knockdown. The recent discovery of prokaryotic OMEGA systems, which are apparent ancestors of the effector nucleases of type II and type V CRISPR systems (Cas9 and Casl2, respectively), along with their eukaryotic homologs, Fanzors, has significantly expanded the known diversity of RNA-guided systems. Concurrently, CRISPR-associated transposons and proteases have contributed distinct RNA-guided functionalities beyond RNA-guided endonuclease activity to the repertoire of natural and potentially applicable RNA-guided molecular machinery. These RNA-guided systems were identified based on their evolutionary relationship with CRISPR-Cas, but to further expand the diversity of RNA-guided systems, more general database mining methodologies are necessary.
[0005] As a result, there is an increasing demand for innovative RNA-guided endonucleases that possess advantageous properties suitable for genome and epigenome editing.
[0006] Citation or identification of any document in this application is not an admission that such a document is available as prior art to the present invention.SUMMARY
[0007] Described in certain example embodiments herein are engineered Tandem Interspaced Guide RNA (TIGR) recombinase system including a Tandem Interspaced Guide RNA (TIGR)-associated (Tas) polypeptide; a recombinase capable of associating with the Tas polypeptide and directing site-specific insertion of a donor insertion sequence from a donor construct at a target insertion site; and a tandem-interspersed guide molecule (TIGR guide) capable of forming a complex with the Tas polypeptide and directing the Tas polypeptide and recombinase to the target insertion site.
[0008] In one embodiment, the tandem-interspersed guide molecule comprises a target binding domain that recognizes the target insertion site and a donor binding domain that recognizes the donor construct.
[0009] In one embodiment, the target binding domain comprises a first stem loop comprising a first spacer and second spacer that bind both stands of a double-stranded target oligonucleotide at the target insertion site.
[0010] In one embodiment, the donor binding domain comprises a second stem loop comprising a first spacer and second spacer that bind the donor construct.
[0011] In one embodiment, the Tas polypeptide comprises a Nop domain. In one embodiment, the Nop domain comprises a coiled-coil domain.
[0012] In one embodiment, the Tas polypeptide further comprises a nuclease domain. In one embodiment, the nuclease domain is catalytically inactive. In one embodiment, the nuclease domain is a RuvC domain or a HNH domain.
[0013] In one embodiment, the Tas polypeptide comprises or consists of a polypeptide that is about 90% to 100% identical to any one of SEQ ID NO. 22199-23909 or a fragment thereof.
[0014] In one embodiment, the Tas polypeptide comprises or consists of a polypeptide that is about 90% to 100% identical to any one of SEQ ID NO: 23910-23938.
[0015] In one embodiment, the recombinase polypeptide is a tyrosine recombinase polypeptide.
[0016] In one embodiment, the tyrosine recombinase comprises a tyrosine recombinase catalytic domain but does not comprise a tyrosine recombinase DNA binding domain.
[0017] In one embodiment, the recombinase polypeptide comprises or consists of a sequence that is 90-100% identical to any one of SEQ ID NO: 23939-23967.
[0018] In one embodiment, the recombinase non-covalently complexes with the Tas polypeptide.
[0019] In one embodiment, the recombinase is fused to the Tas polypeptide or covalently attached via a linker.
[0020] In one embodiment, the donor construct comprises recombinase recognition site(s) and a donor insertion sequence.
[0021] In one embodiment, the size of the donor insertion sequence is 1 to 20 kb.
[0022] In one embodiment, binding to the target sequence is in a PAM- or TAM-independent manner.
[0023] Described in certain example embodiments herein are polynucleotides that encode a (TIGR) recombinase system of the present disclosure.
[0024] Described in certain example embodiments herein are vectors or vector systems comprising a polynucleotide encode a TIGR recombinase system of the present disclosure. In one embodiment, the polynucleotide is operatively coupled to one or more regulatory elements.
[0025] Described in certain example embodiments herein are delivery vehicles configured to deliver one or more components of the TIGR recombinase system of the present disclosure or apolynucleotide encoding one or more components of the TIGR recombinase system of the present disclosure, or the vector or vector system of the present disclosure.
[0026] In one embodiment, the delivery vehicle is a viral or viral-like particle. In one embodiment, the delivery vehicle is a lentiviral, retroviral, or an adeno associated virus (AAV) particle.
[0027] In one embodiment, the delivery vehicle is a non-viral particle.
[0028] In one embodiment, the delivery vehicle is a lipid nanoparticle comprising one or more mRNAs encoding one or more components of the TIGR recombinase system.
[0029] In one embodiment, the delivery vehicle is a lipid nanoparticle comprising a ribonucleoprotein complex.
[0030] Described in certain example embodiments is a cell comprising (a) TIGR recombinase system of the present disclosure; (b) a polynucleotide encoding (a) or a component thereof; (c) a vector or vector system comprising (d); (e) a delivery vehicle a configured to deliver one or more components of the TIGR recombinase system of the present disclosure or a polynucleotide encoding one or more components of the TIGR recombinase system of the present disclosure, or the vector or vector system of the present disclosure; or (f) any combination thereof.
[0031] Described in certain example embodiments is a pharmaceutical formulation comprising: (a) TIGR recombinase system of the present disclosure; (b) a polynucleotide encoding (a) or a component thereof; (c) a vector or vector system comprising (b); (d) a delivery vehicle a configured to deliver one or more components of the TIGR recombinase system of the present disclosure or a polynucleotide encoding one or more components of the TIGR recombinase system of the present disclosure, or the vector or vector system of the present disclosure; (e) a cell comprising (a)-(e) or any combination thereof; or (f) any combination thereof; and a pharmaceutically acceptable carrier.
[0032] Described herein are methods of modifying a target polynucleotide comprising contacting a target polynucleotide with (a) the TIGR recombinase system of the present disclosure; (b) a polynucleotide encoding (a) or a component thereof; (c) a vector or vector system comprising (b); (d) a delivery vehicle a configured to deliver one or more components of the TIGR recombinase system of the present disclosure or a polynucleotide encoding one or more components of the TIGR recombinase system of the present disclosure, or the vector or vector system of the present disclosure; (e) a pharmaceutical formulation comprising (a)-(d) or anycombination thereof; or (f) any combination thereof, whereby the, system mediates site-directed insertion of a donor insert polynucleotide at a target insertion site.
[0033] In one embodiment, delivering occurs in vitro, ex vivo, or in vivo.
[0034] In one embodiment, the target polynucleotide is in a cell. In one embodiment, the cell is a eukaryotic or prokaryotic cell. In one embodiment, the cell is a bacterial cell, a plant cell, a fungal cell, a mammalian cell, a reptile cell, an insect cell, an avian cell, a fish cell, a parasite cell, an arthropod cell, a cell of an invertebrate, a cell of a vertebrate, a rodent cell, a muse cell, a rat cell, a primate cell, or a non-human primate cell. In one embodiment, the cell is a human cell. In one embodiment, the cell is a diseased cell.
[0035] In one embodiment, the target polynucleotide is genomic DNA or extrachromosomal DNA.
[0036] In one embodiment, the target polynucleotide is viral DNA. In one embodiment, the viral DNA is papovavirus, human papillomavirus (HPV), hepadnavirus, Hepatitis B Virus (HBV), herpesvirus, varicella zoster virus (VZV), Epstein-Barr virus (EBV), adenovirus, poxvirus, or parvovirus DNA.
[0037] In one embodiment, the target polynucleotide is a tumor suppressor gene.
[0038] In one embodiment, the target polynucleotide is associated with a blood disorder. In one embodiment, the blood disorder is sickle cell disease or beta-thalassemia. In one embodiment, target polynucleotide is BCL11 A.
[0039] In one embodiment, the target polynucleotide is associate with an inflammatory disease. In one embodiment, the inflammatory disease is transthyretin amyloidosis. In one embodiment, the target nucleotide sequence encodes TTR protein.
[0040] In one embodiment, the target nucleotide sequence is associated with cardiovascular disease. In one embodiment, the target nucleotide sequence is PCSK9, ANGPTL3, or Lp(a).
[0041] In one embodiment, the modification is to a cell ex vivo. In one embodiment, the cell is a T cell, a CAR T cell, a natural killer (NK) cell, a stem cell. In one embodiment, the stem cell is a hematopoietic stem cell or a pancreatic derived stem cell.
[0042] These and other aspects, objects, features, and advantages of the example embodiments will become apparent to those having ordinary skill in the art upon consideration of the following detailed description of example embodiments.BRIEF DESCRIPTION OF THE DRAWINGS
[0043] An understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention may be utilized, and the accompanying drawings of which:
[0044] FIG. 1: Discovery and Genomic Organization of TIGR Systems. FIG. 1A: Structural mining and identification of TIGR systems. The RNA-binding domain (RBD) of SpCas9 (PDB: 8G11) was used as a seed to identify IS110 as a structural homolog. Further structural mining revealed similarity to the Nop domain-containing family. Genomic mining of Nop domaincontaining proteins followed by community detection from embedding of the candidates identified a distinct family of Nop domain-containing proteins associated with Tandem Interspaced Guide RNA (TIGR) systems. Protein structures or models are all colored similarly: RBD, red, blue and green; rest of the Nop domain, wheat. FIG. IB: Genomic organization of TIGR systems and architectures of Tas proteins. TIGR-associated (Tas) protein structural models (left) and genomic locus architecture (right) of three representative Tas proteins (TasA, no nuclease; TasR, RuvC; TasH, HNH). Protein amino acid (top) and nucleotide (bottom) coordinates are shown, as are the number of repeat units in the TIGR arrays. FIG. 1C (SEQ ID NO: 84): Sequence composition of TIGR arrays. Alignment of TaTIGR array individual repeat units. Conserved regions corresponding to edge and loop repeats are shown in black, and variable regions corresponding to spacers A and B are shown in gray.
[0045] FIG. 2. TIGR arrays are processed into 36-nt tigRNAs. FIG. 2A (SEQ ID NO: 85): Small RNA-seq of RNA pulled down with Thermoproteota archaeon TasR (TaTasR) mapped to native TaTIGR array expressed in E. coli. Top: All reads mapped to the TIGR array. Bottom: Filtered 36-nt long reads. The pre-tigRNA transcript is expressed and processed into distinct 36-nt tigRNA units (see also FIG. 12). FIG. 2B (SEQ ID NO: 86-87): Top: Schematic representation of pre-tigRNA processing into tigRNA units. Bottom: Schematic representation of pre-crRNA processing (SpCas9 CRISPR system) into crRNA units. FIG. 2C: Experimental design to test wildtype (WT) or catalytically inactive TasR (dTasR, DI 1A) association with native tigRNA arrays or a minimal tigRNA array consisting of three repeats of tigRNAl (3X tRl) via ribonucleic acid protein (RNP) pulldown. FIG. 2D: Composition of purified TaTasR RNPs using the experimental design shown in FIG. 2C. Left: SDS-PAGE protein gel stained with Coomassie blue. Right: 10% denaturing PAGE gel stained with SYBR Gold to show nucleic acids. FIG. 2E: Small RNA-seq ofRNAs present in TaTasR RNPs showing comparison between the processing of the native array (top) and the 3X tRl array (bottom). FIG. 2F: Small RNA-seq of total RNA from E. coll. The TIGR-TasR expression plasmid contained an arabinose-inducible promoter (PBAD) before the TasR gene, and cells were grown in the presence of arabinose. Coverage and read starts / ends are shown for reads that mapped to the expression plasmid. From top to bottom, the plasmid expresses the 1) WT system, 2) WT system with a nonsense mutation at Asp200 of TasR, 3) Residues 1-322 of TasR (331 residues total) replaced with GFP, and 4) Residues 1-322 to TasR deleted.
[0046] FIG. 3: Identification of the TasR nuclease target. FIG. 3 A: Experimental scheme to identify TasR cleavage sites in the E. coli genome. FIG. 3B (SEO ID NO: 88-89): Mapping of read starts to the E. coli genome. The annotated peak is at the wcaD gene. Example reads start with a ‘T’ due to the non-templated A added during NGS library preparation with KI enow DNA polymerase and identify the 5' end of the input DNA fragment. The pArray variant was either native, minimized to three identical units (3X tRl), or absent (apoprotein). FIG. 3C (SEQ ID NO: 90-91): Top: The wcaD gene sequence with potential matches to spacer A and spacer B indicated, and the sequence of a synthetic target where these matches are improved to full complementarity. Bottom: Denaturing polyacrylamide gel of an in vitro cleavage reaction using purified TasR RNP (coexpressed with a 3X tRl TIGR array) and strand-specifically labelled substrates. FL, fluorescein label on the strand targeted by spacer A.
[0047] FIG.4: DNA targeting rules for TIGR-TasR. FIG. 4A (SEQ ID NO: 92-94): Top: basepairing scheme of tigRNAl from the TaTasR TIGR array with the synthetic optimized target sequence. Triangles indicate cleavage sites. Bottom: comparison with the targeting rules for CRISPR Cas9. PAM, protospacer-adjacent motif. FIG. 4B: In vitro cleavage reactions with TasR. Left: TaTasR (WT) or catalytically inactivated RuvC domain (d, D11A mutation) RNPs were purified and incubated with synthetic target DNA matching the first tigRNA from the TaTasR TIGR array (RNP indicates apoprotein was used). Right: Purified WT TaTasR apoprotein was incubated with the synthesized cognate tigRNAl, no RNA, or a non-targeting (NT) tigRNA, with the target (T) DNA substrate or non-target (NT) substrate. tigRNA2 from the TaTasR TIGR array was used as the non-targeting tigRNA, and PCR amplicon of the SpTasH target was used for the non-target DNA substrate. FIG. 4C (SEQ ID NO: 95-96): Sanger traces for sequencing of the TasR in vitro-cleaved optimized DNA target. The polymerase used in Sanger sequencing adds a non-templated A after running off the template, indicated with an asterisk, and this delineates theprecise cleavage site. FIG. 4D: In vitro activity of TasR containing tigRNAl on labelled singlestranded or double-stranded DNA or RNA substrates containing matches to spacer A and / or spacer B. FIG. 4E (SEQ ID NO: 92, 94): Effects of transversions on the in vitro cleavage of the optimized DNA target by TasR RNPs containing tigRNAl. All transversions were strictly A— > T, T— > A, G^C and C^G. Both strands of the target were mutated. The seed and nickase mutations are indicated. For (B, D, and E) reactions were resolved by denaturing PAGE and visualized by fluorescent labels on the DNA 5' ends.
[0048] FIG. 5: Human genome editing with TIGR-TasR. FIG. 5 A (SEQ ID NO: 97-99): Experimental scheme to test programmable gene editing with TasR in HEK293FT cells, hs, Homo sapiens; CMV, cytomegalovirus; NLS, nuclear localization signal. FIG. 5B: Average indel rates (%) generated by TaTasR and a second TasR ortholog from Parcubacteria of the candidate phyla radiation (ParTasR) at six genomic loci in HEK293FT; data are presented as mean ± s.d. (n = 3). FIG. 5C (SEQ ID NO: 100-104): Sequence of the CXCR4 target site (highlighted in green) in the human genome and corresponding TasR gRNA with spacer matches shown in green. FIG. 5D (SEQ ID NO: 103, 105-125): Indels generated by TaTasR (left) and ParTasR (right) at the CXCR4 target site. FIG. 5E: Distribution of indel size generated by TaTasR (top) and ParTasR (bottom) at the CXCR4 target site.
[0049] FIG. 6: Structural basis for RNA-guided DNA cleavage by TIGR-TasR. FIG. 6A: Cryo-EM structure of TasR containing tigRNAl and the optimized DNA substrate in a postreaction product state. The protein is a dimer, and each protomer is labelled ‘A’ or ‘B’ according to whether its RuvC domain has cleaved the DNA strand targeted by spacer A or B. FIG. 6B: Detail of the RuvC domain of monomer A interacting with the spacer A RNA / DNA heteroduplex. The 5' phosphate of the nick is still in proximity to the RuvC active site, key residues of which are shown and labelled. FIG. 6C: Interaction of the coiled-coil dimerization domain of TasR monomer A with the seed region of the spacer B RNA / DNA heteroduplex (see FIG. 4E). CTD, C-terminal domain. FIG. 6D: Recognition of the pseudosymmetric tigRNA by the symmetric TasR protein dimer. FIG. 6E: Interactions of the box C and box D motifs of the edge (left) and loop (right) repeats with TasR residues. FIG. 6F (SEO ID NO: 126-134): Consensus sequences for edge and loop repeats from nine different TasR TIGR arrays. Excerpts from the full alignments used to derive these consensuses can be found in FIG. 11. Arcs show the Watson / Crick base pairs observed in the structure (which are between different edge repeats on a processed tigRNA but are shownhere on the same edge repeat for simplicity) which are either absolutely conserved (box C / D pairing) or show covariance.
[0050] FIG. 7: Evolutionary and functional diversity of TIGR systems. FIG. 7A: Phylogenetic tree of the Tas Nop domain indicating array association, viral origin, and presence of RuvC (regardless of catalytic activity) or HNH nuclease domain. FIG. 7B (SEQ ID NO: 135): Example of a TasA TIGR system with a stem-loop array (from the gut microbiome of Peromyscus leucopus). Top: locus architecture with TasA protein domain coordinates (in aa) above (coiled-coil region in light green; RNA binding C-terminal domain in blue) and nucleotide coordinates below. Bottom: Alignment of individual stem units, with hairpin sequences (palindromic sequence) in blue flanking each unit. Conserved C and D motifs are shown in black, spacers are in gray, and linkers between units are indicated. FIG. 7C: Small RNA-seq mapping of the TIGR array following RNP pulldown. All three stem arrays are actively expressed and processed at linker regions. FIG. 7D (SEQ ID NO: 136): Schematic of a representative stem-loop repeat unit (tigRNA2 from the TIGR system found in the gut microbiome of Peromyscus leucopus. FIG. 7E (SEQ ID NO: 137-140): Comparative schematic of TIGR systems, IS110, and box C / D. tigRNAs are the simplest and shortest ncRNAs, while IS110 bridge RNAs resemble a fusion of two tigRNAs. box C / D snoRNAs are depicted as tigRNA-like molecules stabilizing the stem structure via the L7 k-turn motif. Bottom: comparison of the TasR-product complex with the crystal structure of an archaeal snoRNP (PDB 3PLA) and a dimer excerpted from a cryo-EM structure of the IS110 bridge RNA complex (PDB 8WT6). The RNP structures of all systems harbor the same domain architecture. TasR is structurally closer to IS110, while the box C / D protein (Nop5) harbors an inactive RuvC recruiting the fibrillarin methylase (green) and stabilizing the C and D box via an interaction with L7 (white). A proposed model positions TIGR as an ancestral system to IS 110 and box C / D snoRNAs, highlighting its evolutionary relevance.
[0051] FIG. 8: Pipeline for TIGR system discovery. FIG. 8A: Structural superimposition of the RNA-binding domain (RBD) from an AlphaFold model of SpCas9 (dark purple) with a region of Tropicimonas sp. IS110 (TroISllO, pink). The best-aligned regions are indicated with coordinates for SpCas9 and TroISllO. FIG. 8B: Structural mining pipeline starting from the RBD of E. coli IS110 (EcISl 10), including the region structurally similar (purple) to the SpCas9 RBD extended to the C-terminal domain of IS 110 (wheat). Structural mining was performed using Foldseek and DALI across databases: AlphaFold clustered at 50% sequence identity (AF50), theEMBL-EBI metagenomic database folded by ESMFold (clustered at 30% sequence identity, Mgnify30), and a protein-virus database folded by AlphaFold (BFVD). This process revealed structural similarities to the Nop domain family (see panels 8C and 8D). Twelve domain profiles related to the Nop domain-containing family were used for genome profile mining using hmmsearch across JGI, NCBI, WGS, EMBL, and MG-RAST databases. FIG. 8C: Phylogenetic tree of Nop domain candidates obtained through structural mining. The outer ring indicates the presence of the RuvC domain detected by structural comparison (regardless of catalytic activity) to EcISllO RuvC using DALI. Box C / D Nop domains are highlighted in purple. Two clades lack RuvC domains, corresponding to Prp31 (green) and an unknown clade (red). Proteins from the unknown clade were converted into a profile for further genomic profile mining. FIG. 8D: Structural comparison of the RBD of SpCas9 (interacting with its guide RNA), EcISllO (seed region used for mining and interacting with bridge RNA), and members of the Nop domain family, including box C / D snoRNAs and Prp31. Shared structural features are highlighted with matching colors. The right panel provides an in-depth comparison of the C-terminal region of the Nop domain, shared by EcISllO, box C / D, and Prp31. Distinct helices are colored to emphasize structural similarities. FIG. 8E: Genomic profile mining pipeline. Hits from genomic mining were clustered at varying sequence identity thresholds down to 50%. The representative sequences were folded using AlphaFold2. Candidates were filtered to retain only those with structural similarity to the E. coli Nop domain seed using DALI. Protein embeddings were generated using the ESM2 1 B model, and embeddings were extracted only for Nop-specific positions. Nop embeddings for each candidate were averaged and compared via cosine similarity. Leiden community detection was applied to identify families of Nop domains. Leiden communities were mapped onto a t-SNE projection of the Nop embeddings. FIG. 8F: Leiden communities were mapped onto a t-SNE projection of the Nop embeddings. Each community was manually inspected, examining structural models and full protein features. RuvC nucleases structural domains (regardless of catalytic activity) are shown in yellow, Nop domain regions in magenta (N-terminal coiled coil helices), green (coiled coil linker), and wheat (C-terminal). Extensions are depicted in grey. The box C / D and Prp31 community are shown. A large group near Prp31 and box C / D, largely lacking RuvC domains, was selected for further investigation.
[0052] FIG. 9: IS110 is an RNA-dependent insertion element. FIG. 9A (SEP ID NO: 142-143): RNA-seq analysis of the Bifidobacterium breve IS 110 locus reveals reads mapping to theupstream region of the IS 110 gene. The schematic of the locus highlights the IS110 gene (pink) and the associated ncRNA upstream. The ncRNA contains two distinct spacers, termed D-spacer and U-spacer, embedded within the loop of a predicted hairpin structure. If these spacers match complementary sequences located upstream (U-target) and downstream (D-target) of the IS 110 gene, suggesting RNA-mediated targeting. FIG. 9B: Disruption of either the ncRNA or the protein integrity significantly reduces transposition efficiency. The graph illustrates the quantitative impact on transposition rates under various conditions, underscoring the requirement of both the ncRNA and the IS110 protein for optimal transposition activity. Probe used for the efficiency quantification by ddPCR is annotated as star.
[0053] FIG. 10: Tas nuclease diversity. FIG. 10A: Phylogenetic tree of Tas Nop domains. The outer ring indicates the presence of nuclease domains: light blue for group 1 (active RuvC), purple for group 2 (inactive RuvC), dark blue for additional RuvC domains (sparse RuvC across the tree, and a small clade with inactive RuvC), and light / dark green for two distinct groups of HNH nucleases. FIG. 10B (SEP ID NO: 11-25): HNH nuclease catalytic site. Structural models highlight the conserved catalytic residues for both HNH groups. Below, a sequence alignment shows conserved positions across HNH nuclease sequences, with red stars marking the catalytic residues. The red arrow indicates which sequences were used for the structural models shown above. FIG. 10C (SEP ID NO: 36-64): Active RuvC catalytic sites. Structural models display the positions of conserved catalytic residues within active RuvC domains. A sequence alignment below highlights the catalytic residues, marked with red stars. FIG. 10D (SEP ID NO: 144-175): Inactive RuvC domains. Structural models show the positions of non-conserved residues within catalytically inactive RuvC domains. A zoomed-in view reveals a conserved structural insertion (in blue) within the RuvC domain, including a residue forming a salt bridge with RuvC and the insertion. Sequence alignment indicates the conserved residue positions with stars, and the insertion sequence is shown in blue.
[0054] FIG. 11: Diversity of Dual Repeat TIGR Arrays. FIG. 11A: Phylogenetic tree of the Nop domain in the Tas family. The inner ring indicates the presence and type of nuclease: green for HNH and blue for RuvC. Nodes labelled around the tree correspond to loci manually inspected for dual repeat arrays; alignments are shown in panel FIG. 1 IB. Blue branches denote association with dual repeat arrays. FpTasA, TaTasR, and SpTasH are indicated in the tree. FIG. 11B (SEP ID NO: 176-328): Alignment of dual repeat arrays. Representative alignments of repeat arraysfrom different loci (labelled in the tree panel A) are shown, each consisting of edge repeat, spacer A, loop repeat, and spacer B. Typically, five repeats are displayed per array. WebLogos above each alignment highlight the conserved positions of repeats and spacers, showcasing sequence conservation and variability across arrays. Putative box C and box D motifs are annotated along with spacer lengths. For RuvC-clade arrays, spacers are annotated as ‘A’ if the preceding repeat is predicted to be an edge repeat by similarity to the TaTasR T1GR array. The repeat arrays of TaTIGR are shown in FIG. 1C in the main text. The first and last repeat arrays of FpTIGR and SpTIGR are shown in FIG. 12 A.
[0055] FIG. 12: Small RNA-Seq Analysis of TIGR Arrays. FIG. 12A (SEO ID NO: 329-330): Alignment of repeats from the FpTIGR and SpTIGR arrays shown in panels FIG. 12B and FIG.12C. Conserved regions corresponding to edge and loop repeats are shown in bold, while variable regions represent spacers A and B. Spacers are defined by their positions relative to the conserved box C and box D motifs. FpTIGR repeats end with a non-conserved nucleotide (CCN box C) that is strictly part of the repeat, whereas SpTIGR spacers end with a conserved A (N8A spacer) outside of the AG box D and are defined as part of the spacer. A WebLogo illustrates sequence conservation across the array below the alignment. FIGs. 12B and 12C (SEQ ID NO: 331): Small RNA-seq of RNA pulled down with (FIG. 12B) Flavonifractor pla tii TasA (FpTasA) or (FIG.12C) Salicolct phage TasH (SpTasH) mapped to the FpTIGR or SpTIGR loci, respectively. Top panels: All reads mapped to (FIG. 12B) FpTIGR or (FIG. 12C) SpTIGR arrays expressed in E. coli. Bottom panels: Mapping of 36-nt reads reveals the pre-tigRNA transcript is processed into distinct 36-nt tigRNA units. FIG. 12D: Distribution of small RNA-seq read lengths from RNA pulled down with Tas proteins. Histograms display the distribution of read lengths (after adapter trimming) from small RNA-seq data of TIGR systems expressed in E. coli. Data are shown for FpTIGR, Thermoproteota archaeon TIGR (TaTIGR), and SpTIGR. The fraction of reads (%) is normalized to the total number of reads, highlighting a prominent peak at 36 nt, indicative of processed tigRNA units. A peak at 72 nt, a multiple of 36, indicates partially processed tigRNA.
[0056] FIG. 13: Repeat Unit Number Requirement and Minimization of TIGR Array. FIG.13 A: Schematic of TaTIGR expression constructs. Diagram depicting five TIGR array plasmids used for expressing TaTasR and associated tigRNAs. Variations include different numbers and identities of repeat units. In pEffector, TasR has a Twin-Strep-SUMO tag. FIG. 13B: Validation of TaTIGR RNP assembly. Representative SDS-PAGE gel (left) shows the expression of TaTasRprotein, and a denaturing PAGE gel (right) verifies the presence of tigRNA in RNP complexes purified from E. coli. FIG. 13C: Schematic of SpTIGR expression constructs. Diagram depicting three TIGR array plasmids used for expressing SpTasH and associated tigRNAs, including minimized constructs. FIG. 13D: Validation of SpTIGR RNP assembly. Representative SDS-PAGE gel (left) shows the expression of SpTasH protein, and a denaturing PAGE gel (right) verifies the presence of tigRNA in RNP complexes purified from E. coli. FIG. 13E: Small RNA-seq mapping of SpTIGR array constructs. Top: Mapping of small RNA-seq reads to the half-edge repeat construct, showing that 5' of tigRNA was not processed at the half-edge repeat. Bottom: Mapping of small RNA-seq reads to the single repeat unit construct, demonstrating tigRNA production from a minimal TIGR array configuration.
[0057] FIG. 14: Biochemical Properties of TaTasR. FIG. 14A: Target dsDNA cleavage by TaTasR in the presence of various divalent metal ions. All subsequent experiments are performed using Mg2+. FIG. 14B: Target dsDNA cleavage by TaTasR at various temperatures. FIG. 14C: Effects of increasing numbers of mismatches to spacer A and / or spacer B on dsDNA cleavage by TaTasR. Each mismatch is a transversion of type A— T, T— > A, G— > C, or C^G. Both strands of the DNA target are mutated in this experiment. FIG. 14D: Effect of gaps or overlaps in the DNA spacer A / B-matching sequences. These experiments used TaTasR loaded with tigRNAl, which only permits testing of a 1 nt overlap, as further overlaps also introduce mismatches to the guides, which independently compromise activity (not shown). All gap sizes are compatible with the TaTasR structure, suggesting reduced activity with larger gaps is due to the energetic penalty of breaking more base-pairs than are formed. (For example, tigRNA spacer A / B pairing introduces 18 bp, a +3 gap breaks 21 bp.) FIG. 14E: Apo TaTasR protein was loaded with in vitro transcribed and dephosphorylated tigRNA variants as indicated. NT, non-targeting; RC, reverse-complemented. The schematic shows the equivalency of tigRNAs with A-B spacers and tigRNAs with B-A spacers. FIG. 14F: Scheme of in vitro TAM identification screen. FIG. 14G (SEO ID NO: 332-333): Upstream and downstream TAMs. Top: TaTasR. Bottom: SpTasH.
[0058] FIG. 15: Biochemical Properties of SpTasH. FIG. 15 A: In vitro cleavage reactions with TasH. Left: SpTasH (WT) or catalytically inactivated HNH domain (d, XxxX mutation) RNPs were purified and incubated with synthetic target DNA matching tigRNAxxx. (RNP indicates apoprotein was used). Right: Purified WT SpTasH apoprotein was incubated with synthesized tigRNAxxx, no RNA, or a non-targeting tigRNA, with the target (T) DNA substrate or non-targeted (NT) substrate. tigRNA2 from the TaTasR TIGR array was used as the non-targeting tigRNA, and PCR amplicon of the SpTasH target was used for the non-targeted DNA substrate. FIG. 15B (SEQ ID NO: 334-335): Sanger traces for sequencing of the TasH in vitro-c\Q NQ optimized DNA target. The polymerase used in Sanger sequencing adds a non-templated A after running off the template, indicated with an asterisk, and this delineates the precise cleavage site. FIG. 15C: Target dsDNA cleavage by SpTasH in the presence of various divalent metal ions. All subsequent experiments are performed using Mg2+. FIG. 15D: Target dsDNA cleavage by SpTasH at various temperatures.
[0059] FIG. 16: In vivo activities of TaTasR on plasmid maintenance, conjugation, and phage infection. FIG. 16A: Conjugation efficiency assay. An RP4 oriT-containing plasmid with or without the biochemically optimized target sequence for tigRNAl was conjugated into a recipient strain containing TaTasR and 3x tigRNAl (3X tRl). Transconjugants were selected after 3.5 hrs of growth, and conjugation efficiency was determined as indicated. Note that the donor cells cannot grow on LB in the absence of diaminopimelic acid (AdapA mutant) and are excluded from the CFU calculations, cat, chloramphenicol acetyltransferase (resistance marker). FIG. 16B (SEQ ID NO: 336-338): Phage defense assay. E. coli Stbl3 contained a TIGR-TasR plasmid either encoding 3 copies of tigRNAl (3X tRl) or 3 tigRNAs designed to target different regions of the phage ZL19 genome as indicated. Plaques are shown for 10-fold dilution series of phage ZL19 or T5 on these strains. FIG. 16C (SEQ ID NO: 339): Schematic of a plasmid maintenance assay. E. coli was cotransformed with arabinose-inducible TIGR-TasR system variants (on an ampicillin resistance marked plasmid) and a target plasmid with a chloramphenicol-resistance gene marker and the biochemically optimized target sequence for tigRNAl. The system was induced or repressed with arabinose or glucose, respectively, and cells were grown overnight without chloramphenicol selection. Plasmid maintenance was determined as the number of chloramphenicol -resistant colonies isolated after overnight induction of the TIGR system relative to total ampicillin-resistant colonies, cat, chloramphenicol acetyltransferase (resistance marker); Amp, ampicillin; Cm, chloramphenicol. FIG. 16D: Plasmid loss after one night of growth without antibiotic selection. Uninduced cells were grown with 0.2% glucose to repress the PBAD promoter; induced cells were grown with 0.2% arabinose to induce the promoter. Only the induced combination with matching TIGR spacers and target plasmid leads to detectable plasmid loss. FIG.16E (SEQ ID NO: 339-340): Selection of escape mutants. Cells were passaged for a further threenights in LB + ampicillin + arabinose without chloramphenicol selection. At this point, plasmid maintenance was estimated as 1 in 10,000 CFUs. Dilutions of cells were plated on media containing ampicillin and chloramphenicol to select for escape mutants. Plasmids were prepared and sequenced by NGS. All 38 unique escape mutants are shown aligned to the wild-type targeted plasmid, along with the number of colonies the mutation was found in. Periods indicate matches to the WT target plasmid, lowercase letters indicate point mutations, and dashes indicate gaps. Mutants are grouped by simple point mutations and larger insertions / deletions (indels). The two escape mutants with wild-type target plasmids had IS1 transposon insertions either in the PBAD promoter or TIGR array of the TIGR system plasmid.
[0060] FIG. 17: Cryo-EM data processing. FIG. 17A: Example cryo-EM micrograph for the TasR RNP + DNA product complex. The left panel shows the raw micrograph, and the right panel shows the micrograph after denoising in cryoSPARC. Denoised micrographs were used for particle picking (grey circles). The particles in the final reconstruction are circled in red. FIG. 17B: Example 2D classes for the particles in the final reconstruction. The right panel shows some example classes from the same particles but with a larger box and circular mask, showing no evidence for higher-order structure beyond dimerization (e.g., tetramers as in IS 110). FIG. 17C: Cryo-EM processing workflow. All steps except particle picking were performed in RELION-5.0. Classes from 3D classification are shown as central slices. The final sharpened map is shown on the right.
[0061] FIG. 18: Cryo-EM density assessment. FIG. 18A: Gold-standard Fourier Shell Correlation (FSC) curves for the final reconstruction as calculated in RELION. FIG. 18B: Orientation distribution for the final reconstruction. FIG. 18C: Map-model FSC as calculated with Phenix. FIG. 18D: Local resolution of the final map as calculated in RELION. The local resolution-filtered map is shown with an additional low-pass filter applied to visualize the peripheral DNA, which is flexible. FIG. 18E: Cryo-EM density masked around each base pair within the two RNA-DNA heteroduplexes. Due to the pseudosymmetry of the complex, if refined with C2 symmetry, the densities for the spacer A and spacer B duplexes would look identical. If correctly refined with Cl symmetry, these densities should appear distinct where spacers A and B differ in sequence. The left panel shows the heteroduplexes in the maps as modeled. The right panel shows the fit of the spacer A heteroduplex into the spacer B density, and vice versa. In most positions, the distinctlylarger density for purines vs. pyrimidines allows distinction of spacer A vs. spacer B and shows that the pseudosymmetry of the complex was accounted for during cryo-EM reconstruction.
[0062] FIG. 19. Structural comparison of TasR with IS110 and box C / D snoRNP. FIG. 19A: Comparison of the Nop and coiled-coil domains of TasR (this study), the IS110 transposase (32), and the box C / D snoRNP Nop5 protein71. One RNA / DNA heteroduplex is shown from each structure, and the equivalent box C / D motifs are highlighted. The box C / D snoRNP structure additionally shows the L7 protein (L7Ae archaeal equivalent) recruited to the k-turn formed by box C / D. Equivalent alpha helices are colored identically. Nop5 in the box C / D snoRNP contains the equivalent to alpha helix 4 in TasR and IS110 after the equivalent to alpha helix 5, so the numbering of these two elements is reversed. FIG. 19B: Comparison of the dimerization interfaces, which are all identical. FIG. 19C: Comparison of the box C / D motif-binding sites on the Nop domain.
[0063] FIG. 20: Diversity of Stem TIGR Arrays. FIG. 20A: Phylogenetic tree of the Nop domain in the Tas family. The inner ring indicates the presence and type of nuclease: green for HNH and blue for RuvC. Nodes around the tree correspond to loci manually inspected for stem loop arrays, with associated stem secondary structures shown in panel B. Red branches denote association with stem arrays. Blue branches represent dual repeat arrays depicted in FIG. 11. FIG.20B (SEQ ID NO: 341-359): Structural prediction of stem units. Predicted secondary structures of two representative stem units are depicted for loci associated with stem arrays. C and D motifs are highlighted in orange and yellow, respectively. Below, a sequence alignment is provided for a stem-less example for comparison, illustrating the diversity and conservation within the TIGR-associated stem arrays.
[0064] FIG. 21: TIGR Association with ParB-like Gene. FIG. 21 A: Locus conservation and variability. Sequence alignment of six loci highlights the conserved operonization of tasA (teal color) and tasParB (blue color) genes, while flanking regions exhibit variability, suggesting a functional association between the two genes. TIGR arrays are shown in light green and found upstream, downstream, or between tasA and tasParB. FIG. 2 IB (SEQ ID NO: 360): Locus organization of TIGR-associated tasA and tasParB genes. Genomic schematic showing an operon where tasA is operonized with a predicted ParB-like domain gene (tasParB)'. The tasA gene is colored according to its domain architecture: light blue for the coiled-coil region, blue for the RNA-binding domain, and grey for the C-terminal region. A conserved motif DGRHR (SEQ IDNO: 360) is highlighted within tasParB. Locus coordinates are provided. The operon is surrounded by two TIGR arrays. FIG. 21C (SEQ ID NO: 361-381): Repeat array alignment. Alignment of upstream and downstream array repeats from the locus in panel B. Both arrays feature truncated edge repeats at their terminal positions, appearing at both the beginning and the end of each array. FIG. 21D: AlphaFold2 multimer prediction of TasA and TasParB tetrameric (2xTasA, 2xTasParB) structure. TasA proteins are predicted to form a head-to-head dimer (light and dark green), a feature distinct from other Tas protein family members. TasParB proteins (beige and black) are predicted to dimerize via cofolding of their C-terminal regions, a characteristic also observed in other ParB-like domain systems. TasParB weakly interacts with TasA. Such prediction suggests more components, e.g., tigRNA and DNA target, might be required to predict the assembly. FIG.2 IE: Structural comparison of TasParB and ParB-like proteins. Structural comparisons with a known ParB-like CTPase (PDB: 7BNR) revealed similar ParB-like folds (highlighted in green), a conserved dimerization domain in the C-terminal region, and a distinct recruiting domain. In ParB-like CTPases, this domain recruits ParC, but in TasParB, this region is structurally divergent, suggesting a different functional interaction. The catalytic site of ParB-like CTPases is indicated in pink, and mapping these positions onto TasParB revealed a distinct and conserved motif in TasParB, suggesting an alternative enzymatic function for TasParB. FIG. 21F (SEO ID NO: 382-391): HHpred analysis of TasParB. Sequence-based homology searches using HHpred identify significant similarity between TasParB and various ParB-like proteins. Notably, ParB-like dndB DNA sulfur modification proteins show similar motifs (TasParB: DGFHR (SEQ ID NO: 360) in orange, vs dndB: DGQHR (SEQ ID NO: 22149) in red) to TasParB aligned with catalytic residue positions (see panel E), suggesting potentially similar enzymatic functionality for TasParB, potentially recognizing DNA phosphorothioate modification or playing a role in related modifications.
[0065] FIG. 22A-22B: An alpha fold2 predicted ribbon diagram of a Tas polypeptide multimer (FIG. 22A) in comparison with PDB:6EMY, nl549 conjugative (FIG. 22B). The multimer contains a dimer of Tas polypeptides that interact with a Yrec domain.
[0066] FIG. 23: Two example transposon contigs with the top being isolated from Nostoc sp. isolate P943 and the bottom contig being isolated from Nostoc sp. 'Peltigera membranacea cyanobionf. Both contain a transposase with an upstream tigRNA and a tyrosine recombinase.
[0067] FIG 24: Determination of the transposon ends by aligning inserted and uninserted loci. The system facilitated insertion of almost 17,000 kb in a target polynucleotide.
[0068] FIG 25: structure of a predicted bridge-like RNA upstream of the transposase. This predicted bridge-like RNA contains two tigRNAs.
[0069] FIG. 26A-26B: Target sequence and insertion site sequence matching with spacers of a first tigRNA molecule of a guide RNA that interaction with a target polypeptide (FIG. 26A) or a donor construct (FIG.26B).
[0070] FIG. 27: Predicted transposon boundaries as determined by using predicted tigRNA molecules. The accessory genes were observed to be variable, but all were aboutlO - 20 kb, all contain a tyrosine recombinase gene in the middle, and a Tas / ISl 10-like transposase at the right end.
[0071] FIG. 28A-28C: Candidate loci identified by TasA structural mining. FIGs. 22A and 22B show the candidate loci having Tas genes proximal to putative methyltransferases and methylases. FIG. 22C shows coverage of RNA-seq data across the locus after selecting for RNAs of length 36 and remapping to the locus.
[0072] FIG. 29A-29C: Electrophoretic mobility shift assays. FIG. 23A shows binding of the TasA TIGR RNA complex to a target DNA sequence. FIG. 23B shows off target binding. FIG.23C shows copurification of the locus-specific methyltransferase with TasA.
[0073] FIG. 30A-30C: Phylogenetic Tree and the Stem Array Clade of TIGR System. FIG.24A shows Stem Array Clade of TIGR system. FIG. 24B shows Representative Stem Array TIGR Loci. FIG. 24C shows co-immunoprecipitation of TasA Protein with candidate TIGR RNAs.
[0074] FIG. 31A-31D: Small RNA Sequencing Reveals Processing of Stem-Loop into Stem-Loop Units. FIGs. 25A and 25B show the structure, sequence and processing of stem and loop (SL) units. FIG. 25C shows stem-loop units as putative guide RNAs for the TasA protein and alignment of 10 stem and loop RNAs. FIG. 25 D shows proposed models of dna targeting by the TasA-D Stem-Loop RNP. FIG 27 A
[0075] FIG. 32A-32D: Electrophoretic mobility shift assays (EMSA) reveals target DNA binding by the TasA RNP (FIG. 26A). FIG. 26B shows AF models of the upstream small protein reveals structure conservation. FIG. 26C shows AF model of complex formation between TasA and the upstream protein. FIG. 26D shows the use of an ESM-2 Language Model to identify conserved proteins in the locus.
[0076] FIG. 33A-33B shows upstream protein and Cas4-Like protein likely associate with TasA and facilitate RNA Processing. FIG. 27 A shows phylogenetic tree with structure of TasH candidates. FIG. 27B shows TasH Candidates in Stem Array TIGR System.
[0077] FIG. 34A and FIG. 34B: Phylogenetic tree of the Tas-YR family constructed using Tas protein sequences. Members of this family are predominantly found in organisms of the phylum Cyanobacteriota, as indicated by the representative photomicrograph inset. FIG. [X]B: Illustration of the genomic architecture, bridge RNA structure, and transposition mechanism of the Tas-YR family. Top: Schematic of a representative Tas-YR transposon element spanning approximately 10-20 kb, with a consensus sequence identity plot shown above. The element encodes a tyrosine recombinase (YR) gene toward the left end and a Tas / ISl 10-like transposase (Tas) gene toward the right end, flanked by terminal repeat sequences recognized by the bridge-RNA. A bridge RNA-like motif is encoded upstream of the Tas gene. Middle: Sequence alignment showing complementarity between the bridge RNA spacers and their respective DNA recognition sequences. The Target Site A (TSA) spacer (5'-UCAGATTAA-3') directs recognition of the genomic target strand, while the Donor Site A (DSA) spacer (5'-AAACACAAA-3') directs recognition of the donor strand. Corresponding target site B (TSB) and donor site B (DSB) recognition sequences are also indicated. Two transposition configurations are depicted: donor circle insertion (top) and linear donor integration (bottom), illustrating how the bridge RNA coordinates simultaneous engagement of target and donor DNA. Right: Secondary structure of the bridge RNA-like motif, comprising BoxC and BoxD structural elements flanking four functional spacer domains — TSA, TSB, DSA, and DSB — organized around a central scaffold, demonstrating the structural basis for dual-site recognition during transposition.
[0078] FIG. 35A, 35B, 35C, and 35D: FIG. 35A: Schematic of the ortholog screening strategy. The native transposon locus encoding a tyrosine recombinase (YR) and a TasR transposase, is co-transformed into A. coli with a plasmid expressing either a TwinStrep-tagged TasR protein or a TwinStrep-tagged YR protein. Following transformation, protein purification via the TwinStrep affinity tag and high-throughput small RNA sequencing are performed to detect RNA species co-purified with each tagged protein. FIG. 35B: Small RNA sequencing read coverage profiles obtained from pull-down experiments using ortholog #6 Tas (SEQ ID NO: 24015) and YR (SEQ ID N: 24013). The top panel shows read coverage from the TwinStrep-TasR pull-down and the bottom panel shows read coverage from the TwinStrep-YR pull-down, bothmapped to the native locus schematic depicting the YR and TasR genes (SEQ ID NO. 24012). A prominent peak of read coverage localizes to the intergenic region upstream of the TasR gene in both pull-downs, indicating robust bridge RNA expression and co-purifi cation with both protein components. FIG. 35C: Phylogenetic tree of the Tas- YR family with the position of ortholog #6 highlighted, showing its placement within the broader family diversity. FIG. 35D: Predicted secondary structure of the bridge RNA of ortholog #6, with nucleotide positions colored according to small RNA sequencing read coverage, ranging from low (grey) to high (red / yellow). The RNA transcription start site is indicated by an arrow, demonstrating that the highly expressed bridge RNA originates at a defined position upstream of the TasR coding sequence.
[0079] FIG. 36: Top row left: Schematic of target DNA recognition by IS621 (IS110 family), showing a 14-nucleotide target sequence engaged by TSA and TSB spacers, with three mismatches tolerated between the guide RNA and the target strand. Top row, center: Corresponding schematic of target DNA recognition by Tas-ortholog #6, showing engagement of a target sequence flanked by an approximately 3.5-nucleotide TAM with the consensus sequence RNHGAWA. The TAM is positioned adjacent to the target site and is recognized in addition to the spacer-mediated basepairing interactions, thereby providing an additional layer of targeting specificity relative to IS621. Top row, right: Sequence logo depicting the TAM consensus for Tas-ortholog #6, with the predominant motif RNHGAWA shown at positions 1-7, where position 4 (G) and positions 5-6 (AA) are the most highly conserved. Bottom row left: Ribbon diagram of the crystal structure of IS621 (PDB: 8WT7), showing the protein in complex with its guide RNA and target DNA, with structural elements colored to highlight domain architecture. Bottom row, right: Ribbon diagram of the AlphaF old-predicted structure of Tas-ortholog #6, overlaid with the IS621 complex structure (PDB: 8WT7). A structurally divergent region of Tas-ortholog #6, absent in IS621, is annotated as a potential TAM recognition domain, indicating that this insertion confers the TAM-dependent specificity characteristic of the Tas- YR family.
[0080] FIG. 37: Screening of Additional Tas-YR Orthologs Reveals Diverse TAM Recognition Sequences Providing Varying Degrees of Targeting Specificity. Top schematic: Diagram illustrating the TAM mapping strategy, showing the position of an 8-nucleotide TAM flanking the target site either on the left (8N-left) or right (8N-right) of the insertion site, used to determine TAM preference for each ortholog tested. Each row presents data for one of three Tas-YR orthologs — Ortholog #28 (SEQ ID NO: 23032), Ortholog #1 (SEQ ID NO: 24028), and Bob(SEQ ID NO: 24030) — comprising, from left to right: an AlphaFold-predicted ribbon structure of the Tas protein, an 8-position sequence logo for the left-flanking TAM (8N-left), and an 8-position sequence logo for the right-flanking TAM (8N-right). Ortholog #28: The predicted structure displays a multi-domain architecture with distinct helical and beta-strand regions. The 8N-left logo shows a moderately degenerate preference with T at position 4 as the most conserved residue, while the 8N-right logo displays a GA-rich motif with G and A most prominent at positions 3-4. Ortholog #1: The predicted structure shows an extended helical domain with pronounced betasheet content. The 8N-left logo reveals a strong single-position preference for T at position 1 with otherwise low information content, while the 8N-right logo shows a strong preference for A at position 1 with limited conservation elsewhere, indicating a relatively permissive TAM. Bob: The predicted structure shows a compact arrangement with a prominent helical bundle. The 8N-left logo shows modest preferences for A, T, and G at positions 1-3 and 2, while the 8N-right logo shows a strong A preference at position 1 and G at position 6, with otherwise low conservation, indicating broad TAM tolerance.
[0081] FIG. 38: _Bar graph showing recombination activity, measured as the percentage of mCherry-positive HEK293FT cells, for five conditions: no editor (negative control), YR-Tas fusion protein, Tas protein alone (SEQ ID NO: 24015), V5-tagged Tas (V5-Tas, SEQ ID NO: 24047), and mNeonGreen-tagged Tas (mNG-Tas, SEQ ID NO: 24049). Each bar represents the mean of technical replicates shown as individual data points with error bars. The no editor control produces background-level mCherry signal of approximately 0.1%, confirming assay specificity. All Tas-containing constructs produce recombination activity substantially above background, demonstrating that the Tas protein alone, without co-expression of a separate tyrosine recombinase, is necessary and sufficient to mediate site-specific recombination in this assay. The YR-Tas fusion construct produces the highest mean activity at approximately 0.9% mCherry-positive cells, while Tas alone produces approximately 0.5%. Addition of an N-terminal epitope tag modestly increases recombination efficiency relative to untagged Tas: V5-Tas produces approximately 0.7% and mNG-Tas approximately 0.65% mCherry-positive cells. The data establish that N-terminal fusion tags are tolerated by the Tas protein without loss of activity, and that the V5 epitope tag provides a slight enhancement in recombination efficiency. On the basis of these results, V5-Tas is selected as the preferred protein format for subsequent engineering and optimization experiments.
[0082] FIG. 39: Optimization of a Minimal Essential System for Efficient Genome Insertion in E. coli. FIG. 39A: Schematic of the minimal two-plasmid system used for genome insertion in E. coli. The donor plasmid (pDonor, 3 kb) contains a 14-nucleotide bridge RNA recognition site and a 14-nucleotide flanking sequence flanking the sequence to be inserted. The helper plasmid (pHelper) provides co-expression of the Tas transposase protein and the bridge RNA from a single construct. Both plasmids are co-transformed into E. coli and cultured for 16 hours, after which genomic DNA is extracted and insertion efficiency is quantified by droplet digital PCR (ddPCR). FIG. 39B: Bar graph showing genome insertion frequency (%) quantified across eight different chromosomal target sites in the E. coli genome (G2 (SEQ ID NO: 24050), G3 (SEQ ID NO: 24051), G4 (SEQ ID NO: 24052), G5 (SEQ ID NO: 24053), G11 (SEQ ID NO: 24054), G12 (SEQ ID NO: 24055), G13 (SEQ ID NO: 24053), and G16 (SEQ ID NO: 24056)). For each site, insertion efficiency is independently measured at both the 5' and 3' junctions, shown as paired bars, confirming the occurrence of precise, full-length insertion events at each locus. Insertion frequencies vary across the eight genomic sites tested, ranging from approximately 5% to 35%, demonstrating that the minimal two-component system achieves robust site-specific genome insertion across multiple chromosomal loci. Site G4 shows the highest insertion efficiency among those tested, reaching approximately 33% at the 3' junction. The data collectively establish that a pDonor encoding a 14-nucleotide bridge RNA recognition site with 14-nucleotide flanking sequences, combined with pHelper-driven Tas protein and bridge RNA expression, constitutes a minimal yet sufficient system for efficient programmable genome insertion at diverse sites in a bacterial host.
[0083] FIG. 40: Split Bridge RNA Increases Insertion Efficiency in the E. coli Genome. Left column, top: Schematic of the original (old) full-length bridge RNA design (SEQ ID NO: 24034), in which a single RNA molecule encodes both the target site (TS) and donor site (DS) recognition modules, expressed from the pHelper plasmid together with Tas. Left column, bottom: Schematic of the redesigned split bridge RNA (new RNA), in which the donor site (DS) module is grafted onto the structural scaffold of the target site (TS) RNA, yielding two separate RNA species — a TS RNA (SEQ ID NO: 24035) and a DS RNA (SEQ ID NO:24058) ) — both expressed from the modified pHelper plasmid together with Tas. Center top: Workflow schematic showing cotransformation of the pHelper plasmid and a 3 kb pDonor into E. coli, followed by 16 hours of culture, genomic DNA extraction, and ddPCR quantification of insertion efficiency at both the 5'1and 3' junctions of the inserted sequence. Bar graph: Comparison of insertion frequency (%) at genomic site G2 using the old full-length RNA design versus the new split RNA design, with 5' and 3' junction measurements shown as paired bars for each condition. The old RNA design achieves approximately 20% insertion frequency at both junctions. In contrast, the new split RNA design achieves approximately 80% insertion frequency at the 5' junction and approximately 68% at the 3' junction at the same genomic site, representing an approximately four-fold improvement in insertion efficiency. These data establish that the split bridge RNA architecture, in which the DS module is grafted onto the TS scaffold, substantially increases genome insertion efficiency in E. coli, achieving up to 80% insertion frequency.
[0084] FIG. 41: Tas and Bridge RNA Can Mediate Programmable Deletion. Top: Schematic illustrating the mechanism of programmable deletion of a large DNA sequence. A plasmid carrying both a donor site and a target site flanking an intervening sequence is shown. Upon expression of Tas and bridge RNA, recombination between the donor and target sites on the same plasmid molecule results in excision of the intervening sequence as a circular DNA product, leaving behind a plasmid from which the intervening sequence has been deleted. Bottom left: Gel electrophoresis showing deletion activity for four replicate reactions each using the old full-length RNA design and the split RNA design. Two bands are visible per lane: an upper band corresponding to the undeleted (no deletion) plasmid and a lower band corresponding to the deletion product. Reactions with both RNA designs produce clearly visible deletion bands, with the split RNA design producing proportionally more deletion product relative to the no-deletion band across all four replicates. Bottom center: Bar graph quantifying deletion efficiency (%) for the old RNA and split RNA designs across multiple conditions. The split RNA design achieves up to 100% deletion efficiency in the E. coli plasmid assay, representing complete recombination between the donor and target sites, compared to substantially lower efficiencies with the old RNA design. These data demonstrate that the split bridge RNA architecture mediates programmable 100% deletion efficiency in an E co / z plasmid. This capability has direct application to the programmable deletion of harmful genetic elements, such as antibiotic resistance genes and virulence factors, from gut microbiome bacteria, representing a potentially significant therapeutic use of the Tas- YR system.
[0085] FIG. 42: Testing System Efficiency in Human Cells Using Junction PCR and Reporter Readout. Schematic illustrating the assay strategy used to evaluate the TIGR recombinase systemin human cells, with the goal of achieving site-specific, programmable insertion of large DNA sequences in the human genome. Four plasmids — encoding the bridge RNA, the Tas protein, a donor construct, and a target construct — are co-transfected into HEK293FT cells by transient transfection to test recombination between donor and target plasmids. Following transfection, recombination activity is assessed by three complementary readout methods: junction PCR, which detects the formation of recombinant junctions between donor and target sequences on the recovered plasmid DNA; luminescence, measured at the cell population level; and fluorescence, measured by detection of reporter-positive individual cells. This multi-readout strategy enables both qualitative and quantitative assessment of insertion efficiency during system optimization in human cells.
[0086] FIG. 43: Two secondary structure diagrams of the bridge RNA are shown, with structural elements highlighted in red boxes to indicate the positions of modifications tested. Left structure: The bridge RNA in its native configuration, indicating the positions of G-U wobble base pairs distributed along the stem regions and the 5' end where an HDV-like ribozyme can be appended. The target spacers A and B are shown within the central loop region. Two modifications tested at this structure that decreased recombination efficiency are indicated: (1) substitution of the naturally occurring G-U wobble base pairs with canonical G-C Watson-Crick base pairs, and (2) attachment of a 5' HDV-like ribozyme to the 5' end of the bridge RNA. These results indicate that G-U wobble base pairs within the stem regions are functionally important for recombination activity, and that the native 5' terminus architecture is preferred. Right structure: The same bridge RNA scaffold with the tetraloop, bulge, and linker elements annotated. Three structural modifications tested at this configuration had no measurable effect on recombination efficiency: (1) decreasing the size of the tetraloop, (2) removing the bulge, and (3) removing the linker element at the 3' end. These results demonstrate that the tetraloop, bulge, and linker are structurally dispensable for recombination activity, identifying elements of the bridge RNA that can be modified without loss of function and thereby informing rational engineering of minimized or optimized guide RNA variants.
[0087] FIG. 44: Removal of the Pl Stem Loop in the Target RNA Improves Integration Efficiency. Left: Secondary structure diagram of the bridge RNA target site (TS) module, indicating two features relevant to the modifications tested. A DNA-RNA mismatch position is annotated within the stem region proximal to the target spacers A and B. The Pl stem loop, locatedat the 5' end of the RNA, is highlighted with a red box. Two modifications that improved recombination efficiency were identified: correcting a single base mismatch between the RNA and its DNA target, and removal of the Pl stem loop. Right: Bar graph quantifying recombination efficiency across ten conditions, measured as the percentage of mCherry-positive cells. The conditions tested are, from left to right: no RNA (negative control, ~0.1%); no TS (-0.1%); wildtype TS (SEQ ID NO: 24012) (WT TS, -0.3%); addition of a 5' ribozyme (SEQ ID NO: 24043) (+ribozyme, -0.1%); G-U wobble substitution (SEQ ID NO: 24041) (-wobble, -0.2%); tetraloop size reduction (SEQ ID NO: 24039) (tetraoloop, -0.2%); bulge removal (SEQ ID NO: 24040) (-bulge, -0.2%); linker removal (SEQ ID NO: 24037) (-linker, -0.25%); correction of the singlebase DNA-RNA mismatch (SEQ ID NO: 24042) (-mismatch, -0.35%); and removal of the Pl stem loop (SEQ ID NO: 24038) (-P1, -0.5%). The -Pl modification produces the highest recombination efficiency of all single modifications tested, approximately five-fold above background and representing a near two-fold improvement over the wild-type TS RNA, establishing the -Pl bridge RNA design as the preferred single RNA modification for subsequent system optimization in human cells.
[0088] FIG. 45: Splitting the Bridge RNA and Using the TS Scaffold Improves Recombination Efficiency in Human Cells. Left: Secondary structure diagram of the full-length bridge RNA showing the complete architecture of both the target-site (TS) module (blue) and the donor-site (DS) module (orange) as a single continuous molecule. Target spacers A and B are located within the TS central loop, and donor spacers A and B are located within the DS module. This structure illustrates the basis for the split RNA strategy, in which the two functional modules are separated into independent RNA species, with the DS module grafted onto the structural scaffold of the TS RNA (hereafter, DSg). Top right: Schematic diagrams of the four RNA configurations tested — full-length bridge RNA (full), TS module alone (TS), DSg module alone (DSg), and the combined split design (TS+DSg) — depicted as simplified secondary structure cartoons in blue (TS) and orange (DSg). Bottom right: Junction PCR gel evaluating recombination activity for each of the four RNA configurations, with 5' and 3' junction bands assessed for each condition. The full-length bridge RNA produces junction PCR bands at both the 5' and 3' junctions, confirming baseline recombination activity. Neither the TS alone nor the DSg alone produces detectable junction bands, demonstrating that both modules are individually insufficient for recombination. In contrast, the TS+DSg split RNA combination produces clear junction bands atboth the 5' and 3' junctions at approximately 200 bp, establishing that the split RNA design, in which the TS scaffold carries the DS spacers, is necessary and sufficient to restore and enhance recombination activity in human cells.
[0089] FIG. 46: Rational Mutagenesis of the Tas Protein Based on Ortholog Comparison and ISCro4 Mutations. Left: AlphaF old-predicted ribbon structure of the Tas protein, with candidate mutated residues highlighted in pink (positions identified by comparison with Tas-YR family orthologs) and blue (positions identified by analogy with mutations previously shown to improve activity in the related IS110-family transposase ISCro4, as reported by Perry et al., 2025, PMID: 40997214). The structure illustrates the spatial distribution of mutation candidates across the protein. Center: Schematic alignment of the ortholog comparison strategy used to identify candidate residues. Positively charged amino acids (arginine, lysine, and histidine; R / K / H) that differ between the reference ortholog and multiple other Tas-YR family orthologs are identified as mutation candidates (indicated by pink dots) at two positions. This ortholog-guided approach identifies positions where introduction of positively charged residues may enhance DNA binding or recombination activity.
[0090] FIG. 47:_Comprehensive Screen of Rational Tas Point Mutants Identifies TopPerforming Variants from Three Structural Regions of the Protein. Left: AlphaFold-predicted ribbon structure of the Tas protein with three structural regions color-coded to indicate the origin of the top-performing mutants: the dactylus region (pink / salmon), the wedge loop region (blue), and the RuvC domain (yellow). These three regions of the protein collectively account for the highest-activity single point mutants identified in the screen, establishing them as functionally important surfaces for recombination activity. Right: Ranked bar graph showing mCherry-positive cell percentage (%) for approximately 50 individual V5-Tas point mutants, YR-Tascontrol, and wild-type (WT) V5-Tas, tested in HEK293FT cells. Bars are color-coded by mutation origin: hatched bars denote mutations informed by ISCro4, solid pink bars denote mutations informed by Tas-YR ortholog comparisons, and open white bars denote wild-type. A dotted horizontal line indicates the wild-type efficiency threshold at approximately 1.3% mCherry-positive cells. The majority of single mutants tested produce activity at or below wild-type levels. Example single mutants, highlighted within a dashed box at the right end of the graph, are H294K, Y299K, V303K, A27K, T300K, and E356K, all derived from Tas-YR ortholog-guided positions and all producing mCherry-positive cell percentages substantially above wild-type, with the best single mutantsreaching approximately 2.5-3% mCherry -positive cells. These mutations map to the dactylus, wedge loop, and RuvC structural regions of the Tas protein.
[0091] FIG. 48: Combinatorial Mutagenesis of Top-Performing Tas Variants Further Improves Recombination Efficiency in Human Cells, with Triple Mutants Reaching Up to Approximately Seven-Fold Above Wild-Type. Left: AlphaF old-predicted ribbon structure of the Tas protein with the three structural regions from which the top-performing single mutants derive color-coded: the dactylus region (pink / salmon), the wedge loop region (blue), and the RuvC domain (yellow). This structural annotation identifies the spatial relationship between the three mutation-bearing regions and underscores that each region contributes independently to recombination efficiency, providing a structural rationale for additive effects when mutations from different regions are combined. Center bar graph: Systematic screen of double-mutation combinations of V5-Tas variants, tested in HEK293FT human cells using the mCherry fluorescent reporter assay, with recombination efficiency expressed as percentage mCherry-positive cells normalized to wild-type V5-Tas. Bars are color-coded by the structural origin of the constituent mutations (dactylus, wedge loop, and / or RuvC), enabling direct comparison of intra-region versus inter-region combinations. The no-editor control and wild-type V5-Tas are included as negative and reference controls, respectively. Among double mutants, A27K+E356K and Y299K+E356K achieve the highest activity, reaching approximately three- to four-fold above wild-type, substantially exceeding the best single mutants. Right bar graph: Extension of the combinatorial screen to triple-mutation combinations. The top-performing triple mutants, A27K+V303K+E356K and A27K+Y299K+E356K, achieve recombination efficiencies of approximately six- to seven-fold above wild-type. These data demonstrate that mutations drawn from distinct structural regions of the Tas protein act additively to improve recombination efficiency, establishing multi-site engineered Tas variants as high-performance TIGR recombinase system components for human genome editing applications.
[0092] FIG. 49: Engineered Tas Polypeptide Variants and Engineered Guide RNA Modifications Act Additively to Improve Recombination Efficiency in Human Cells, with Combined Optimized Systems Reaching Approximately Four-Fold Above Wild-Type. Left: AlphaFold-predicted ribbon structure of the Tas protein with the three structural regions from which the top-performing point mutants derive color-coded: the dactylus region (pink / salmon), the wedge loop region (blue), and the RuvC domain (yellow). This structural annotation, consistentwith the protein engineering data presented in the preceding figures, identifies the spatial context of the H294K and E356K substitutions tested in this figure. Right: Bar graph presenting mCherry-positive cell percentage (%) across two paired RNA conditions — wild-type (WT) bridge RNA and -Pl RNA (bridge RNA from which the Pl stem loop has been removed) — each tested in combination with four Tas protein variants: no editor control, YR-Tas, V5-WT, V5-H294K, and V5-E356K. Schematic diagrams below the x-axis illustrate the structural distinction between the WT and -Pl RNA architectures, with the Pl stem loop present in the WT RNA and absent in the -Pl RNA. Under WT RNA conditions, V5-H294K and V5-E356K each achieve approximately two- to three-fold above V5-WT, consistent with the protein engineering results described herein. Switching to -Pl RNA further increases efficiency for all Tas variants tested, and the combination of either V5-H294K or V5-E356K with -Pl RNA reaches approximately four-fold above V5-WT with WT RNA. These data demonstrate that protein engineering and guide RNA engineering contributions are additive, establishing that co-optimized TIGR recombinase systems comprising both an engineered Tas polypeptide and an engineered guide molecule achieve superior recombination efficiency in human cells relative to either optimization alone.
[0093] FIG. 50 provides a series of diagrams and bar graphs demonstrating that the Tas protein alone is sufficient for transposon excision and insertion. Panel A is a schematic illustration of a 4.5 kb mini-transposon construct and a series of sequential deletion mutants (Remove 1 -Removed), depicting insertion into a pTarget plasmid and excision as assayed in E. coli. Panels B and C are bar graphs showing insertion efficiency (right end junction rate) and excision rate, respectively, for each deletion mutant. Panels D and E are bar graphs showing insertion efficiency and excision rate, respectively, for a series of YR domain mutants, including a stop codon mutant, a catalytic residue mutant (Y— F), and a full YR ORF deletion, demonstrating that the YR recombinase domain is dispensable for transposon recombination activity. The min-transposon sequences are complete transposon (SEQ ID NO: 24016), Remove 1 (SEQ ID NO: 24017), Remove 2 (SEQ ID NO: 24018), Remove 3 (SEQ ID NO: 24019), and Remove 4 (SEQ ID NO: 24020).DETAILED DESCRIPTION OF THE EXAMPLE EMBODIMENTS OVERVIEW
[0094] The present disclosure relates to improved systems and methods for targeted genetic modification, and more particularly to a novel programmable gene editing system that address significant limitations of current programmable-nuclease-based approaches.
[0095] Current programmable editing systems, like CRISPR-Cas, face several notable constraints that limit their broad applicability in research and therapeutic contexts. These limitations include: (1) substantial size requirements that complicate delivery, particularly for therapeutic applications; (2) rigid structural architectures that become prohibitively large when modified with additional functional domains; (3) guide RNA sequences that typically interact with only one strand of target DNA, potentially reducing targeting specificity; and (4) strict protospacer adjacent motif (PAM) requirements that restrict the range of targetable genomic loci.
[0096] The novel programmable gene editing system described herein, Tandem Interspersed Guide RNA (TIGR) systems, addresses these limitations through several innovative features. The system comprises a split-spacer guide molecule, which may be referred to herein as a tigRNA, and a TIGR-associated (Tas) polypeptide. The system's compact size facilitates more efficient delivery across a broader range of applications while maintaining full editing functionality. Its modular architecture can accommodate the addition of diverse functional domains while maintaining a delivery-compatible size, thereby enabling expanded functionality without compromising cellular delivery efficiency - a significant advantage over current systems where additional domains often render the final engineered system too bulky for effective delivery.
[0097] A feature of TIGR systems is its unique guide sequence design, which enables simultaneous binding to both strands of the target DNA sequence. This dual-strand interaction may substantially enhance targeting specificity compared to current single-strand-targeting approaches, potentially reducing off-target effects that have hindered widespread therapeutic application of existing gene editing technologies.
[0098] Furthermore, the system may operate independently of PAM or TAM sequence requirements, a significant advancement over current gene editing systems. This PAM-independent targeting dramatically expands the range of accessible genomic loci, enabling modifications at previously inaccessible sites and providing greater flexibility in target selection for both research and therapeutic applications.
[0099] In another aspect, a TIGR system described herein can facilitate transfer of a polynucleotide from a donor construct and insertion of the polynucleotide into a target polynucleotide in a site-specific manner. This is facilitated by a tigRNA that can bind a donor construct and a target polynucleotide and a Tas polypeptide that is associated with a recombinase polypeptide.
[0100] In another aspect, a TIGR system described herein can facilitate epigentic modification of target DNA molecules by directing methylation of target loci. This is faciliated by a tigRNA that can bind to the target DNA molecule and direct methylation of the target DNA at a target loci via Tas polypeptide that is associated with a methyltransferase polypeptide.
[0101] The functionality of this system has been validated in human cells, demonstrating successful genome editing activity in therapeutically relevant contexts. The combination of these features - compact size, modular architecture enabling expanded functionality while maintaining deliverability, enhanced targeting specificity, PAM-independent operation, and validated human cell activity - represents a substantial improvement over the current state of the art in programmable gene editing systems. These advantages make the present invention particularly well-suited for applications ranging from basic research to therapeutic genome editing, where precise, flexible, and efficient genetic modification is essential.TIGR SYSTEMS
[0102] The engineered or non-naturally occurring TIGR systems disclosed herein comprise a Tas polypeptide and a tandem interspersed guide molecule (TIGR guide). The Tas polypeptide may comprise a Nop domain that may also optionally include a RuvC domain, a HNH domain, or both. In an embodiment, the Tas polypeptide is a protomer that forms a dimer with another Tas polypeptide protomer. The split-spacer guide molecule is capable of forming a complex with the Tas polypeptide and can direct sequence-specific binding with both sense and antisense strands of a double-stranded oligonucleotide, which can be subsequently cleaved by the Tas polypeptide.
[0103] In one embodiments, the engineered or non-naturally occurring TIGR system is an engineered or non-naturally occurring TIGR recombinase system that includes a Tas polypeptide, a recombinase capable of associating with the Tas polypeptide and directing site-specific insertion of a donor insertion sequence from a donor construct at a target insertion site; and a tandem-interspersed guide molecule (TIGR guide) capable of forming a complex with the Tas polypeptide and directing the Tas polypeptide and recombinase to the target insertion site. In one embodiment, the TIGR recombinase system includes a TIGR guide that includes target binding domain that recognizes the target insertion site in a target molecule and a donor binding domain that recognizes the donor construct. The donor insertion sequence can be a polynucleotide of interest (such as a gene of interest). In one embodiment, the Tas polypeptide lacks a DNA binding domain.
[0104] In one embodiment, the TIGR recombinase system includes one or more Tas polypeptides. In one embodiment, the TIGR recombinase system includes two Tas polypeptides. In one embodiment, the two Tas polypeptides are dimerized. In one embodiment, the two Tas polypeptides are dimerized head to tail. In one embodiment, the two Tas polypeptides are dimerized via a helix domain in each of the Tas polypeptides. In one embodiment, the TIGR recombinase system includes one or more Tas polypeptides and a recombinase polypeptide (also referred to herein as “a recombinase”), where the recombinase polypeptide is coupled with one or more of the one or more Tas polypeptides. In one embodiment, the TIGR recombinase system includes a multimer that includes one or more Tas prolylpeptides and a recombinase polypeptide. In one embodiment, the TIGR recombinase system includes a multimer that includes two or more Tas prolylpeptides and a recombinase polypeptide. In one embodiment, the TIGR recombinase system includes one or more Tas polypeptides and a recombinase polypeptide, where the recombinase polypeptide is recruited by the one or more Tas polypeptides and / or tigRNA and associates with one or more of the one or more Tas polypeptides. Tas polypeptides and recombinase polypeptides are described in greater detail elsewhere herein.
[0105] In one embodiment, the TIGR recombinase system comprises one or more TasR polypeptides. In one embodiment, the TIGR recombinase system comprises two or more TasR polypeptides. In one embodiment, the TIGR recombinase system includes two TasR polypeptides that are dimerized. In one embodiment the TIGR recombinase system includes two TasR polypeptides that are dimerized head to tail. In one embodiment, the TIGR recombinase system includes two TasR polypeptides that are dimerized via a helix domain. In one embodiment, one or more of the TasR polypeptides lacks a DNA binding domain. In one embodiment, two or more of the TasR polypeptides lack a DNA binding domain. In one embodiment, the TIGR recombinase system includes one or more TasR polypeptides and a recombinase polypeptide, where the recombinase polypeptide is coupled with one or more of the one or more TasR polypeptides. In one embodiment, the TIGR recombinase system includes a multimer that includes one or more TasR prolylpeptides and a recombinase polypeptide. In one embodiment, the TIGR recombinase system includes a multimer that includes two or more TasR prolylpeptides and a recombinase polypeptide. In one embodiment, the TIGR recombinase system includes one or more TasR polypeptides and a recombinase polypeptide, where the recombinase polypeptide is recruited by the one or more TasR polypeptides and / or tigRNA and associates with one or more of the one ormore TasR polypeptides. TasR polypeptides and recombinase polypeptides are described in greater detail elsewhere herein.
[0106] In an embodiment, the engineered or non-naturally occurring TIGR systems disclosed herein comprise a Tas polypeptide, a Tas-associated methyltranserase and a tandem interspersed guide molecule (TIGR guide). The Tas polypeptide may comprise a Nop domain that may also optionally include a RuvC domain, a HNH domain, or both. In an embodiment, the Tas polypeptide is a protomer that forms a dimer with another Tas polypeptide protomer. The split-spacer guide molecule is capable of forming a complex with the Tas polypeptide and can direct sequencespecific binding with both sense and antisense strands of a double-stranded oligonucleotide, which can be subsequently methylated by the Tas-associated methyltransferase.
[0107] In one embodiment, the target molecule can be a double or single-stranded nucleic acid. In one embodiment, the target molecule can be a double-stranded or single- stranded deoxyribonucleic acid (DNA). In one embodiment, the target molecule sequence can be either double-stranded or single-stranded ribonucleic acid (RNA). In one embodiment, the target molecule can be an RNA / DNA hybrid.
[0108] In one embodiment, the target molecule can be a double or single-stranded nucleic acid. In one embodiment, the target molecule can be a double-stranded or single- stranded deoxyribonucleic acid (DNA). In one embodiment, the target molecule sequence can be either double-stranded or single-stranded ribonucleic acid (RNA). In one embodiment, the target molecule can be an RNA / DNA hybrid.
[0109] In one embodiment, the target nucleotide sequence is a genomic DNA sequence of a eukaryote. In one embodiment, the target nucleotide sequence is a genomic DNA sequence of a prokaryote. In one embodiment, the target nucleotide sequence is a genomic DNA sequence of a mammal. In one embodiment, the target nucleotide sequence is a genomic DNA sequence of a primate. In one embodiment, the target nucleotide sequence is a genomic DNA sequence of a human. In one embodiment, the target nucleotide sequence is a genomic DNA sequence of a plant. In one embodiment, the target nucleotide sequence is a genomic DNA sequence of an insect. In one embodiment, the target nucleotide sequence is a genomic DNA sequence of a bacterium. In one embodiment, the target nucleotide sequence is a genomic DNA sequence of yeast. In one embodiment, the target nucleotide sequence is a mitochondrial DNA sequence. In one embodiment, the target nucleotide sequence is a chloroplast DNA sequence.Tandem Interspersed Guide Molecules
[0110] As used herein, a tandem interspersed guide molecule (also referred to herein as a “TIGR Guide,” “guide molecule” and “tigRNA”) refers to the dual-spacer configuration of a TIGR guide molecule, in which the spacer nucleotide sequence is split into a first spacer nucleotide sequence (Spacer A) that can base pair with a corresponding complementary nucleotide sequence on the sense strand of the double- stranded DNA target nucleotide sequence and a second spacer nucleotide sequence (Spacer B) that can base pair with a corresponding complementary nucleotide sequence on the antisense strand of the double-stranded DNA target nucleotide sequence (see FIG.4). In an embodiment, the guide molecule may further comprise an edge repeat-first spacer-loop repeat-second spacer-edge repeat architecture. In one embodiment, the guide molecule comprises nucleotide sequence motifs arranged in a 5’ to 3’ direction: a 5’ edge repeat, a first spacer, a loop repeat, a second spacer and a 3 ’edge repeat, where the 5’ and 3’ edge repeats do not base pair to form a stem (tandem array). In one embodiment, the guide molecule comprises nucleotide sequence motifs arranged in a 5’ to 3’ direction: a 5’ edge repeat, a first spacer, a loop repeat, a second spacer and a 3 ’edge repeat, where the 5’ and 3’ edge repeats base pair to form a stem (stem loop array). See e.g., FIG 7E. In one embodiment, the tandem interspersed guide molecule may not comprise a loop repeat.[OHl] In one embodiment, the guide molecule is 36 nucleotides in length. In one embodiment, the guide molecule can be 72 nucleotides in length. In one embodiment, the guide molecule can be 108 nucleotides in length. In one embodiment, the length of the guide molecule can be at least 36 nucleotides. In one embodiment, the length of the guide molecule can be from 36 to 72 nucleotides, e.g., 36, 37, 38, 3940, 41, 42, 43, 44, 45, 46, 4748, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, or 72 nucleotides.
[0112] In one embodiment, the length of the guide molecule can be from 36 to 108 nucleotides, e.g., 36, 37, 38, 3940, 41, 42, 43, 44, 45, 46, 4748, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, or 108 nucleotides.
[0113] In one embodiment, the length of the guide molecule can be from 36 to 144 nucleotides, e.g., 36, 37, 38, 3940, 41, 42, 43, 44, 45, 46, 4748, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86,87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 17, 138, 19, 140, 141, 142, 143, or 144 nucleotides.
[0114] In one embodiment, the engineered or non-naturally occurring TIGR system, including but not limited to an engineered or non-naturally occurring TIGR recombinase system, includes a tigRNA that includes target binding domain that recognizes the target insertion site and a donor binding domain that recognizes the donor construct. See e.g., FIGS. 25 and 26A-26B. In this way, the tigRNA can form a “bridge” that joins the target polynucleotide or oligonucleotide and the donor construct on which the Tas polypeptide and recombinase can act to facilitate removal of the donor insertion sequence from the donor construct and its insertion into the target polynucleotide or oligonucleotide. Other features of the stem loops are described elsewhere herein. The first stem loop and the second stem loop can be connected by about 20 to about 80 nucleotides. In one embodiment, the first stem loop and the second stem loop are connected by 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides. In one embodiment, the first stem loop and the second stem loop are connected by 20, to 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides. In one embodiment, the first stem loop and the second stem loop are connected by 30-70 nucleotides. In one embodiment, the first stem loop and the second stem loop are connected by 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 nucleotides. In one embodiment, the first stem loop and the second stem loop are connected by 30 to 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 nucleotides. In one embodiment, the first stem loop and the second stem loop are connected by 40-60 polynucleotides. In one embodiment, the first stem loop and the second stem loop are connected by 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides. In one embodiment, the first stem loop and the second stem loop are connected by 40 to 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides. In one embodiment, the first stem loop and the second stem loop are connected by 50 polynucleotides.
[0115] In one embodiment, the target binding domain includes a stem loop (“a first stem loop”) that includes a first spacer and a second spacer that bind both strands of a double stranded target polynucleotide or oligonucleotide. In one embodiment, the first spacer can bind a first strand of the double stranded polynucleotide, and the second spacer can bind the opposite strand of the double stranded polynucleotide or oligonucleotide. In one embodiment, the donor binding domain of a tigRNA includes a stem loop (“a second stem loop”) that includes a first spacer and a second spacer that bind both strands of a double stranded target polynucleotide or oligonucleotide. In one embodiment, the first spacer can bind a first strand of a double stranded donor construct, and the second spacer can bind the opposite strand of the double stranded donor construct. This binding can occur at one or more tigRNA recognition sites within the donor construct. The tigRNA recognition sites are further described elsewhere herein in connection with the donor binding constructs.Spacer
[0116] The spacer may also be referred to herein as a “guide sequence” within the context of the complete guide molecule and correspond to a variable sequence that can be reprogrammed to direct the TIGR complex to a given target sequence in a double-stranded oligonucleotide by engineering of the appropriate complementary sequences. As noted above the first spacer can bind to a complimentary sequence on the sense strand, and the second spacer can bind to a complimentary sequence on the anti-sense strand and define a nicking site on each strand, that when combined result in a double-strand cleavage of the target oligonucleotide. In an embodiment, the first and second spacer nick their respective stands in the same way. In an embodiment, the first and second spacer nick their respective stands in a different way. In an embodiment, each spacer specifies a stand cut site 3’ to the nucleotide complimentary to the fifth base of the spacer (FIG 4C). In an embodiment, the spacer is the variable sequence between a box D motif and the a C motif. In one embodiment, the box D motif has the sequence UG. In one embodiment, the box C motif has the sequence CCN where N can be any nucleotide. In one embodiment, the box C motif has the sequence CCA.
[0117] In one embodiment, the first or second spacer nucleotide sequences can be from 5 to 24 nucleotides in length. In one embodiment, the first or second spacer nucleotide sequences can be from 9 to 24 nucleotides in length. In one embodiment, the first or second spacer nucleotide sequences can be from 12 to 24 nucleotides in length. In one embodiment, the first or second spacernucleotide sequences can be from 15 to 24 nucleotides in length. In one embodiment, the first or second spacer nucleotide sequences can be from 18 to 24 nucleotides in length. In one embodiment, the first or second spacer nucleotide sequences can be from 21 to 24 nucleotides in length. In one embodiment, the first or second spacer nucleotide sequences can be from 6 to 21 nucleotides in length. In one embodiment, the first or second spacer nucleotide sequences can be from 6 to 18 nucleotides in length. In one embodiment, the first or second spacer nucleotide sequences can be from 6 to 15 nucleotides in length. In one embodiment, the first or second spacer nucleotide sequences can be from 6 to 9 nucleotides in length. In one embodiment, the first or second spacer nucleotide sequences can be from 6 to 15 nucleotides in length. In one embodiment, the first or second spacer nucleotide sequences can be from 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides in length. In one embodiment, the first or second spacer nucleotide sequences can be from 9 to 12 nucleotides in length.
[0118] In one embodiment, the first or second spacer nucleotide sequence can have 100% complementarity with a nucleotide sequence on one of the strands of a double-stranded DNA target nucleotide sequence. In one embodiment, the first or second spacer nucleotide sequence can have 90% complementarity with a nucleotide sequence on one of the strands of a double-stranded DNA target nucleotide sequence. In one embodiment, the first or second spacer nucleotide sequence can have 80% complementarity with a nucleotide sequence on one of the strands of a double-stranded DNA target nucleotide sequence.
[0119] In one embodiment, the first spacer binds a complementary sequence of one stand and the second spacer binds a complementary sequence of the other strand directly adjacent to the reverse-complement of the sequence that paired with the first spacer. See FIG. 6. C.Edge RepeatsIn an embodiment, the guide molecule has an edge repeat on either end of the guide molecule. The 5’ and 3’ edge repeats may also be referred herein as the first and second scaffold sequence respectively. The edge repeats interact with the Nop domain of the Tas proteins helping facilitate complex formation of the guide molecule with the Tas polypeptide. In an embodiment, the each edge repeats may comprise box C and box D motifs, and these box C / D motifs can bind to identical sites in the Nop domain of a Tas protein (Fig. 7E). In an embodiment, a 5’ first edge repeat may end with a box C motif and the 3’ edge repeat may begin with a box D motif. In the context of astem-loop array configuration, the 5’ edge repeat and the 3’ edge repeat may base pair to form a stem.Loop Repeat
[0120] The loop repeat is located between the first and second spacer sequence. In one embodiment, the loop repeat can be from 8-12 nucleotides in length. In one embodiment, the guide molecule loop repeat nucleotide sequence can be 5-25 nucleotides in length. In one embodiment, the guide molecule loop repeat nucleotide sequence can be 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In one embodiment, the guide molecule loop repeat nucleotide sequence can be 12-25 nucleotides in length. In one embodiment, the guide molecule loop repeat nucleotide sequence can be 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length.
[0121] In one embodiment, the loop repeat nucleotide sequence also starts on the 5’ end after a box D motif and ends on the 3’ end with the start of a box C motif as described for the edge repeats above. The sequence between the box D and box C motifs may vary but is often A rich.Gap sequence
[0122] In one embodiment, the first or second spacer nucleotide sequences can be separated by a gap sequence of 0, 1, or 2 nucleotides. In one embodiment, the first or second spacer nucleotide sequences can be separated by 0 to 25 nucleotides. In one embodiment, the first or second spacer nucleotide sequences can be separated by 5 to 25 nucleotides. In one embodiment, the first or second spacer nucleotide sequences can be separated by 10 to 25 nucleotides. In one embodiment, the first or second spacer nucleotide sequences can be separated by 15 to 25 nucleotides. In one embodiment, the first or second spacer nucleotide sequences can be separated by 20 to 25 nucleotides. In one embodiment, the first or second spacer nucleotide sequences can be separated by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 or more nucleotides.Guide Molecule Modifications
[0123] The guide molecules of the present invention may be modified to enhance their functionality, stability, and specificity.Truncations
[0124] In an embodiment, the guide molecules may be truncated to enhance targeting specificity while maintaining activity. For instance, whereas truncation of CRISPR-Cas9 guide RNA spacers from 20 to 17 nucleotides has been shown to retain on-target activity while reducingoff-target effects by up to 5,000-fold, the present invention's guide molecules may be similarly optimized through strategic truncation while preserving the dual-strand binding capability.Chemical Modifications
[0125] In an embodiment, the guide molecules may incorporate modified nucleotides to enhance stability and reduce immunogenicity. Such modifications may include 2'-deoxy substitutions at terminal positions. Modified nucleotides may include 2'-fluoro modifications at terminal positions, potentially reducing off-target effects while enhancing nuclease resistance. Additionally, 2'-O-methyl modifications may be employed, particularly at terminal positions, where incorporation at three 3'- and 5'-terminal positions has been demonstrated to increase activity while improving stability.
[0126] In an embodiment, the guide molecules may incorporate phosphorothioate (PS) linkage modifications. Full phosphorothioate modification has been shown to increase activity approximately four-fold in certain CRISPR systems, with further enhancement observed when combined with terminal 2'-O-methyl modifications. The present invention may advantageously employ such modifications while maintaining its characteristic dual-strand binding capability.
[0127] In an embodiment, Bridged nucleic acid (BNA) modifications are chemical alterations to the structure of nucleic acids that enhance their stability and binding properties. In the context of guide molecules, BNA modifications can significantly improve specificity and performance. These modifications, including bridged nucleic acids (BNAs) and locked nucleic acids (LNAs), involve constraining the ribose ring via a methylene-based bridge. This structural change increases the melting temperature of DNA / RNA duplexes, improves the affinity and stability of hybridization, and enhances resistance to nucleases.
[0128] When incorporated into guide RNAs, BNA modifications can have several important effects. BNANC (N-methyl carbamate) incorporation in crRNAs has been shown to increase Cas9 cleavage specificity, particularly when placed in the central region (positions 10-14) of the guide sequence. BNA NC modifications also slow down nuclease kinetics, which contributes to improved specificity. Structurally, BNA modifications may promote an extended A-form structure throughout the guide molecule, altering how it hybridizes with the target oligonucleotides. This structural change likely contributes to the improved discrimination against off-target sequences.
[0129] The improved specificity of BNA-modified guide RNAs appears to work through a conformational mechanism. BNA modifications alter the hybridization properties of the guideRNA with the target DNA, leading to more dynamic interactions, especially with off-target sequences. The modified structure may promote more stringent base pairing requirements, enhancing discrimination against mismatches. By fine-tuning the placement and number of BNA modifications, researchers can optimize guide RNAs for improved specificity while maintaining efficient on-target activity.Hybrid Structures
[0130] In an embodiment, the guide molecules may comprise DNA-RNA hybrid structures, wherein specific regions of the guide molecule incorporate DNA nucleotides while other regions maintain RNA character. Such hybrid guide molecules may provide advantages in stability, specificity, or cellular processing. Characteristic repeating DNA-RNA structures (charDNA) may be incorporated into the guide molecule design, potentially enhancing nuclease resistance while maintaining targeting functionality. The DNA portions of hybrid guide molecules may be strategically positioned to optimize guide performance, including modulation of flexibility, binding kinetics, or interactions with cellular machinery.Conjugation Strategies
[0131] In an embodiment, the guide molecules may incorporate RNA aptamer sequences capable of specifically recruiting additional sequences (such as recombination templates) and functional domains to the target site. Such aptamer sequences may be positioned at various locations within the guide molecule structure, including terminal regions or internal loops, while maintaining the guide's core targeting functionality. Multiple aptamer sequences may be incorporated into a single guide molecule to enable simultaneous recruitment of different sequences and functional domains, with spacing and orientation optimized to prevent steric interference while maintaining efficient recruitment. The aptamer sequences may be designed to recruit epigenetic modifiers, transcriptional regulators, fluorescent proteins, or other functional proteins of interest through direct aptamer-protein interactions or through aptamer binding to specific small molecule tags conjugated to the desired functional domains.
[0132] In an embodiment, the guide molecules may be modified with chemical end modifications including amine groups, azide groups, fluorescent dyes, or strained alkynes at either the 5' or 3' terminus. Such modifications may be particularly advantageous for tracking guide molecule localization or enabling specific conjugation strategies. These modifications may becombined with the system's dual-strand binding architecture to enable both enhanced targeting and improved visualization or delivery.
[0133] Click chemistry conjugation offers numerous avenues to enhance guide molecule function, providing a versatile toolset for researchers in drug delivery, gene therapy, and molecular imaging. This approach allows for precise and efficient attachment of functional groups to guide molecules, improving their targeting capabilities and delivery mechanisms. By conjugating specific ligands or targeting moieties, researchers can increase binding affinity to target cells or tissues, improve cellular uptake, and enable tissue-specific delivery. Click chemistry modifications can also enhance the stability and bioavailability of guide molecules by allowing the attachment of stabilizing groups like PEG, protective moieties to reduce enzymatic degradation, or cellpenetrating peptides to enhance cellular entry. The technique's versatility enables the creation of multifunctional guide molecules, combining imaging agents, therapeutic payloads, and multiple targeting ligands for improved specificity. The modular nature of click chemistry facilitates rapid optimization of guide molecule properties, allowing for easy screening of different functional groups and configurations. Additionally, click reactions can be used for site-specific labeling, enabling in vivo tracking, integration of affinity tags, and controlled modification in complex biological environments. By harnessing these advantages, researchers can develop more effective and adaptable guide molecules for a wide range of applications in biomedical research and therapeutic development.
[0134] The aforementioned modifications may be implemented individually or in combinations determined to be beneficial for specific applications. The optimal combination of modifications may vary depending on the intended use, target sequence, delivery method, and cellular context. For example, combinations of stability-enhancing modifications such as 2'-O-methyl and phosphorothioate modifications may be particularly advantageous for in vivo applications where enhanced durability is required. The modification strategies may be specifically optimized to complement the system's unique dual-strand binding mechanism, potentially offering enhanced control over targeting specificity and activity beyond what has been achieved with single-strand targeting approaches.Donor Construct
[0135] The engineered or non-naturally occurring TIGR recombinase system includes a donor construct. The donor construct can include a donor insertion sequence. The donor construct canprovide a donor insertion sequence that can be removed from the donor construct and inserted into a target polynucleotide by the engineered or non-naturally occurring TIGR recombinase system of the present description. The donor insertion sequence can be any sequence that is desired to be inserted into a target polynucleotide. In one embodiment, the donor construct includes one or more donor insertion sequences. In one embodiment, the donor construct includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more donor insertion sequences.
[0136] The donor construct can contain one or more recombinase recognition sites. In one embodiment, the donor insertion sequences can be operatively coupled to one or more recombinase recognition sites. In this context, “operatively coupled” means that the one or more recombinase recognition sites are included in the donor construct at a location that a recombinase will interact with the donor construct such that the recombinase can facilitate removal of the donor insertion sequence to which the one or more recombinase recognition sites are operatively coupled. In one embodiment, the donor insertion sequences can include one or more recombinase recognition sites in the donor construct. In one embodiment, each of the donor insertions sequences are flanked by one or more recombinase recognition sites. The one or more recombinase recognition sites can bind or otherwise interact, e.g., be cleaved, by one or more molecules, including a Tas protein of the present disclosure, so as to remove the donor insertion sequences from the donor construction. In one embodiment, the donor construct is cleaved at or in proximity to one or more of the one or more recombinase recognition sites. In one embodiment, the donor construct is cleaved within 1-10 nucleotides of one or more of the one or more recombinase recognition sites.
[0137] In one embodiment, the recombinase recognition sites are 2 to 30 nucleotides. In one embodiment, the recombinase recognition sites are 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 2 to 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 3 to 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 4 to 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 5 to 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 5 to 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17,18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 6 to 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 7 to 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 8 to 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 9 to 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 10 to 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 11 to 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 12 to 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 13 to 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 14 to 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 15 to 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 16 to 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 17 to 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 18 to 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 19 to 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 20 to 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 21 to 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 22 to 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 23 to 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 24 to 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 25 to 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 26 to 27, 28, 29, or 30 nucleotides. In one embodiment, the recombinase recognition sites are 27 to 28, 29, or 30nucleotides. In one embodiment, the recombinase recognition sites are 28 to 29 or 30 nucleotides. In one embodiment, the recombinase recognition sites are 29 or 30 nucleotides.
[0138] The donor construct can contain one or more tigRNA recognition sites. These are sites within the donor construct that can hybridize with a donor binding domain of a tigRNA. In one embodiment the donor binding domain of the tigRNA includes a first stem loop that includes first and / or a second spacer of a tigRNA that can hybridize with one or more tigRNA recognition sites in the donor construct. In one embodiment, the tigRNA recognition sites are located within the donor construct so that, when hybridized with a tigRNA, the system can interact with the donor construct and remove a donor insertion sequence from the construct. In one embodiment, the donor insertion sequences can be operatively coupled to one or more tigRNA recognition sites. In this context, “operatively coupled” means that the one or more tigRNA recognition sites are included in the donor construct at a location that a tigRNA will interact (such as hybridize) with the donor construct such that the tigRNA can facilitate removal of the donor insertion sequence to which the one or more tigRNA recognition sites are operatively coupled.
[0139] In one embodiment, the tigRNA recognition sites are 2 to 30 nucleotides. In one embodiment, the tigRNA recognition sites are 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 2 to 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 3 to 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 4 to 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 5 to 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 5 to 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 6 to 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 7 to 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 8 to 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 9 to 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21,22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 10 to 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 11 to 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 12 to 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 13 to 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 14 to 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 15 to 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 16 to 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 17 to 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 18 to 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 19 to 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 20 to 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 21 to 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 22 to 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 23 to 24, 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 24 to 25, 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 25 to 26, 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 26 to 27, 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 27 to 28, 29, or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 28 to 29 or 30 nucleotides. In one embodiment, the tigRNA recognition sites are 29 or 30 nucleotides.
[0140] In one embodiment, the donor construct is a linear nucleic acid. In one embodiment, the donor construct is a circular nucleic acid. In one embodiment, the donor construct is a single stranded nucleic acid. In one embodiment, the donor construct is double stranded. In one embodiment, the donor construct is DNA. In one embodiment, the donor construct is RNA.
[0141] In one embodiment, the size of the donor insertion sequence is 1 to 20kb. In one embodiment, the size of the donor insertion sequence is Ikb to 1.5kb, 2kb, 2.5kb, 3kb, 3.5kb, 4kb, 4.5kb, 5kb, 5.5kb, 6kb, 6.5kb, 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, 1 Ikb, 11.5kb,12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is I.5kb to 2kb, 2.5kb, 3kb, 3.5kb, 4kb, 4.5kb, 5kb, 5.5kb, 6kb, 6.5kb, 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 2kb to 2.5kb, 3kb, 3.5kb, 4kb, 4.5kb, 5kb, 5.5kb, 6kb, 6.5kb, 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 2.5kb, to 3kb, 3.5kb, 4kb, 4.5kb, 5kb, 5.5kb, 6kb, 6.5kb, 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 3kb to 3.5kb, 4kb, 4.5kb, 5kb, 5.5kb, 6kb, 6.5kb, 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 3.5kb to 4kb, 4.5kb, 5kb, 5.5kb, 6kb, 6.5kb, 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 4kb to 4.5kb, 5kb, 5.5kb, 6kb, 6.5kb, 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 4.5kb to 5kb, 5.5kb, 6kb, 6.5kb, 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 5kb to 5.5kb, 6kb, 6.5kb, 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, II.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 5.5kb to 6kb, 6.5kb, 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 6kb to 6.5kb, 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb,14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 6.5kb to 7kb, 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 7kb to 7.5kb, 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 7.5kb to 8kb, 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 8kb to 8.5kb, 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 8.5kb to 9kb, 9.5kb, lOkb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 9kb to 9.5kb, lOkb, 10.5kb, 1 Ikb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 9.5kb to 10kb, 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is lOkb to 10.5kb, llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 10.5kb to llkb, 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is llkb to 11.5kb, 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 11.5kb to 12kb, 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 12kb to 12.5kb, 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 12.5kb to 13kb, 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb,19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 13kb to 13.5kb, 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 13.5kb to 14kb, 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or20kb. In one embodiment, the size of the donor insertion sequence is 14kb to 14.5kb, 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 14.5kb to 15kb, 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 15kb to 15.5kb, 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 15.5kb to 16kb, 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 16kb to 16.5kb, 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 16.5kb to 17kb, 17.5kb, 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 17.5kb to 18kb, 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 18kb to 18.5kb, 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 18.5kb to 19kb, 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 19kb to 19.5kb, or 20kb. In one embodiment, the size of the donor insertion sequence is 19.5kb to 20kb.TIGR-associated (Tas) Polypeptides
[0142] Tas polypeptides are compact in size ranging from 200 to 400 amino acids in size. In one embodiment, the polypeptide is between 150 and 500 amino acids in size. In a specific embodiment, the polypeptide is between 150 and 400 amino acids in size. In another embodiment, the polypeptide is between 200 and 450 amino acids in size. In yet another embodiment, the polypeptide is between 250 and 400 amino acids in size. In a further embodiment, the polypeptide is between 300 and 350 amino acids in size. In still another embodiment, the polypeptide is between 150 and 300 amino acids in size. In an additional embodiment, the polypeptide is between 350 and 500 amino acids in size. In certain embodiments, the polypeptide is between 150 and 225 amino acids in size, between 225 and 300 amino acids in size, between 300 and 375 amino acids in size, or between 375 and 500 amino acids in size. In particular embodiments, the polypeptide consists of 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, or 500 amino acids.
[0143] Tas polypeptides may comprise a Nop domain. In an embodiment the Nop domain further comprises a coiled-coil region (“coiled-coil domain”, “coiled-coil structural motif’) In an embodiment, the Tas polypeptide consists only of a Nop domain. In an embodiment, the Tas polypeptide includes a nuclease domain. In another embodiment, the Tas polypeptide may further comprise a RuvC nuclease domain, or an HNH nuclease domain. In one embodiment, the nuclease domain is catalytically active. In one embodiment, the nuclease domain is catalytically inactive. In one embodiment, the nuclease domain includes one or more mutations that modify the activity of the nuclease domain. In one embodiment, the nuclease domain includes one or more mutations that increase activity of the nuclease domain. In one embodiment, the nuclease domain includes one or more mutations that decrease or eliminate activity of the nuclease domain.
[0144] In one embodiment, the Tas polypeptide can dimerize. In one embodiment, the Tas polypeptide can homodimerize. In one embodiment the Tas polypeptides homodimerizes head to tail. In one embodiment, the Tas polypeptide or Tas homodimer can form a multimer with one or more polypeptides.Nop Domains
[0145] The Nop domain (PF AM accession number: PF01798, InterPro ID: IPR002687) represents a structural and functional element in proteins involved in RNA processing, particularly within the context of ribonucleoprotein (RNP) complex formation and function. This domain, characterized by its compact globular sub-structure, typically comprises 100-200 amino acid residues and is predominantly located at the C-terminal end. Structurally, the Nop domain exhibits a densely packed, energy-efficient conformation that facilitates specific molecular interactions. Its globular morphology provides well-defined binding surfaces, crucial for its role as an RNP binding module. This structural configuration enables the domain to engage in both RNA and protein interactions with high specificity and affinity, in nucleolar proteins like Nop56 and Nop58, it associates with box C / D small nucleolar RNA complexes. The evolutionary conservation of Nop domains across all three domains of life - Archaea, Bacteria, and Eukarya - underscores their biological significance. This conservation suggests a fundamental role in RNA processing and RNP formatting processing that has been maintained throughout evolutionary history.Coiled-Coil Structural Motifs
[0146] Coiled-coil domains are ubiquitous structural motifs in proteins, characterized by 2-7 alpha-helices intertwined in a rope-like configuration. These domains, present in approximately 5-10% of proteins across all domains of life, exhibit a distinctive heptad repeat pattern (hxxhcxc) of hydrophobic (h) and charged (c) residues. This arrangement engenders an amphipathic structure with a hydrophobic stripe spiraling around the helix, facilitating the formation of left-handed supercoils in either parallel or anti-parallel orientations. Coiled-coils serve multifarious functions in cellular processes, including protein-protein interactions, molecular spacing, oligomerization, and signal transduction. In an embodiment, the coiled-coil motif is located on the N-terminal side of the Nop domain.RuvC Domains
[0147] As used herein, the term " RuvC domain" refers to a nuclease domain having structural and functional characteristics similar to those found in RuvC resolvase proteins involved in DNA repair and recombination. In various embodiments, a RuvC domain comprises one or more catalytic residues, including histidine and aspartic acid residues, capable of coordinating divalent metal ions, particularly Mg2+ions, in the active site to facilitate DNA cleavage activities. RuvC domains function as components of CRISPR-Cas nuclease systems, including but not limited to Cas9 and Casl2a (Cpfl) proteins. In the context of Cas9, the RuvC domain operates in conjunction with an HNH domain, wherein the RuvC domain specifically cleaves the non-target DNA strand while the HNH domain cleaves the target strand, resulting in a double-strand break. In contrast, in Cast 2a systems lacking an HNH domain, the RuvC domain may be responsible for cleaving both DNA strands through a series of conformational changes. In an embodiment, the RuvC domain correspondence to the protein family identified by Pfam identifier PF01548. In an embodiment, the RuvC domain is on the N-terminal side of the Nop domain. In an embodiment, the coiled coil motif is located between the RuvC and Nop domains.
[0148] The RuvC endonuclease domain is a critical component of various nucleases, including the CRISPR-Cas9 system and the E. coli RuvC protein. The RuvC domain features a typical RNase H fold, consisting of a central five-stranded 0-sheet surrounded by helices. In Cas9, the RuvC domain is divided into RuvC -I and RuvC-II subdomains, separated by a "stop" (STP) domain. The active form of the RuvC protein functions as a homodimer. The RuvC domain is responsible for cleaving DNA strands. In E. coli RuvC, it resolves Holliday junctions during DNA recombination and repair. In Cas9, it is a defining characteristic of Class 2 CRISPR systems, including Cas9 and Cast 2 proteins, where it cleaves the non-target strand (NTS) of the target DNA. The catalytic center typically involves four essential acidic residues, which are in Cas9: DIO, E762, H983, andD986. These residues coordinate with divalent metal ions (usually Mg2+or Mn2+) for catalysis. In CRISPR-Cas systems, the RuvC domain requires the HNH domain to be in an active conformation for optimal positioning and catalysis.
[0149] Mutations in the RuvC domain can significantly impact DNA cleavage activity, particularly of the non-target strand. The RuvC domain is highly conserved across different CRISPR-Cas systems, indicating its fundamental importance in DNA cleavage mechanisms. Its presence in both bacterial DNA repair systems (like E. coli RuvC) and CRISPR-Cas systems suggests an evolutionary link between these DNA-processing machineries.
[0150] TasR, which features a RuvC nuclease in its N-terminal region (see FIG. IB and TABLES IX-XIII below) is predicted to be either active (e.g., in the Thermoproteota archaeon isolate LB CRA l locus, TaTIGR) or inactivated (e.g., in the Candidatus Buchananbacteria bacterium RIFCSPHIGHO2_02_FULL_56_16 rifcsphigho2_02_scaffold_l 5087) (FIG. 10D). Detail of the RuvC domain of monomer A interacting with the spacer A RNA / DNA heteroduplex is depicted in Figure 10. The 5' phosphate of the nick is shown to be in proximity to the RuvC active site.
[0151] In one embodiment, RuvC’s active site comprises amino acid residues DI 1, W 12, Dll, E43, D87, or D90 (shown in bold in TABLE IX).
[0152] In one embodiment, one or more amino acid residues in the RuvC’s active site is mutated. In one embodiment, two or more amino acid residues in the RuvC’ s active site is mutated. In one embodiment, three or more amino acid residues in the RuvC’s active site is mutated. In one embodiment, the mutations in RuvC’s active site abolish RuvC’s catalytic activity.
[0153] In one embodiment, the RuvC’s active site includes a mutation of one or more amino acid residues chosen from Dll, W12, Dll, E43, D87, and D90.
[0154] In one embodiment, the RuvC’s active site includes a mutation of two or more amino acid residues chosen from DI 1, W12, Dll, E43, D87, and D90.
[0155] In one embodiment, the RuvC’s active site includes a mutation of three or more amino acid residues chosen from Dll, W12, Dll, E43, D87, and D90.TABLE IX:RuvC nuclease (TaTasR (Thermoproteota archaeon isolate LB CRA l QQVH01000006)) 1 M D N D I T I L A V D W S H E E R K L A 201 ATGGATAACGATATCACAATATTGGCGGTCGATTGGTCTCATGAAGAACGTAAACTTGCA 60 21 I F D G K K I R K K L P E P S S D V I I 4061 ATTTTTGATGGCAAGAAGATAAGGAAAAAACTTCCTGAGCCTTCTAGCGATGTAATAATT 120TABLE IX:RuvC nuclease (TaTasR (Thermoproteota archaeon isolate LB CRA l QQVHO 1000006)) 41 V A E N I P Q K Y A A P F I E V G A K V 60121 GTGGCTGAAAATATTCCTCAAAAATATGCCGCTCCGTTCATCGAGGTAGGGGCAAAAGTT 180 61 L R C S T N A T A D A R K N Y Q K K V D 80181 CTTAGATGTTCGACAAACGCCACAGCAGATGCAAGAAAAAACTATCAAAAGAAAGTCGAT 24081 A A F A K N D E N D S K V I W A L Y Q T 100241 GCTGCTTTTGCGAAAAACGACGAAAATGATTCTAAAGTGATTTGGGCGTTATATCAAACA 300101 H P E L F R E M K L E ( SEQ I D NO: 34 ) 111301 CACCCAGAATTGTTTCGTGAAATGAAGCTTGAG ( SEQ I D NO: 35 ) 333
[0156] The consensus sequence for the RUVC nuclease active site can be found in Table VI below and in FIG. 10:TABLE X: RUVC nuclease catalytic site (Branch 1)' Vjgg. AH F £ ' < V - U l JV ¥88 KA'AK F R C <laiu E g F vr m L P \ v£8 r®i A Q:< A 'M CGSS H r K VY F l 1 F' K 1 ^ K A F =■■&: CE K ¥ U x R KI K A I-3S KHi D Kc- K K IWx K; <;.iS RX® HK K FZ I K I H I® E Y < O RK ¥ K ¥ K 11 M 3 PX KcK K ¥ ¥ vil 1® A A‘ X<:-: » R€ X'ffiS I K < V K < F VgJ I® K A K H R C VWJS H K K \ F X ' V8 ® S A F I ¥ S R C| ®*N V P K K J K 33 ® K A r L A IS KI OT$ H D < ¥ K K FX vs IIS K P S f R O1 ®A \ 1 FX p S3 US b A H i. 5 B C1 MEA V P F P K 133 F A Q'AV: - ■■ A |jg FXREGION 1 | | REGION 2M A E &• • ’. « ■ ■.; IgA RSW. wvi WV WV" $3. 13E!s L K ISA OA. ISA ISE V ISA S3 13N A ($F.. P I SA 33Y5 i t ®AR »¥ 13AK & I M F ISA gg (A £ I3K - - _ E IS fl 1 33N A H - I E IS K S3W L 33 "N AK ¥ r 8 R 3IW I 33N A E HE 1 ■ ~ b SAK Ot L 130 AK HE L E IS K SS¥ 1 &K M F IS R IIP 1 13Q K H t SAQ OF 1 13D A A HK ¥! HSR 83 L H| REGION 3;g||TABLE X: RUVC nuclease catalytic site (Branch 1)CONSENSUS SEQUENCE SEQ IDvWftiV £ K V - Vffl W / A f ~ A 18 Rt64 W Arc FW ww* iK-B - ~ - x E w:*w i> V - K winAR I vwI" - vwi:8 tvw1 W E 10 16 20 CONSENSUS st X K L A 9 K K (SEQ ID NO: 65) □ * l ^ Catalytic residues~X K (SEQ ID NO: 66) X ■■■■ A / G / S O D / E / K R / K / N X X ix? MO(SEQ ID NO: 67) X X X X X X X X XHNH Domains
[0157] As used herein, the term " HNH domain" refers to a compact nuclease domain belonging to the HNH endonuclease superfamily, characterized by a PPa-metal fold structure. In various embodiments, an HNH domain comprises a conserved catalytic triad typically consisting of aspartic acid (D), histidine (H), and asparagine (N) residues, which coordinate divalent metal ions, particularly Mg2+, to facilitate DNA cleavage activities. In certain embodiments, HNH domains function as components of CRISPR-Cas nuclease systems, including but not limited to Cas9, wherein the HNH domain specifically cleaves the target DNA strand complementary to the guide RNA. In an embodiment, the HNH domain corresponds to the protein family identified by Pfam identifier PF01844. The domain may maintain its structural and functional characteristics across a range of physiological conditions suitable for gene editing applications. In an embodiment, theRuvC domain is on the N-terminal side of the Nop domain. In an embodiment, the coiled coil motif is located between the RuvC and Nop domains.
[0158] HNH endonuclease domains capable of nicking double-stranded DNA sites in the presence of a divalent metal ion possess a conserved catalytic HNH core motif, with a zinc-binding site [CxxC] found in many bacteriophages and prophages. The catalytic core motif of HNH endonucleases typically folds into a PPa-fold (see FIG. 10B), which consists of two antiparallel strands, a connecting loop of variable length ( -loop), an a helix, and a Zn metal binding site located between these elements (Zhang et al., Sei Rep 7, 42542, (2017)). This minimal catalytic core exists in site-specific homing endonucleases, restriction enzymes, pyocins, colicins, and anaredoxins. Bacterial HNH domains are known to function as both site-specific DNA nucleases (as seen in homing nucleases or the RNA-guided DNA endonuclease Cas9) and non-specific nucleases (such as colicins, which are a subgroup of bacterial toxins) (Jinek et al., Science 337, 816-21, (2012)).
[0159] TasH contains an N-terminal HNH nuclease domain (see Tables III and V). TasH cleaves target DNA in an HNH- and tigRNA-dependent fashion when the opposite strands of the DNA target are base paired with the two tigRNA spacers (see FIG. 15B).
[0160] In one embodiment, HNH’s active site comprises amino acid residues H(N), H(N+4), and G(N+5) (shown in bold in TABLE V, where N can be the number of any amino acid residue within the HNH domain).
[0161] In one embodiment, HNH’s active site comprises amino acid residues H(N) and H(N+4) (shown in bold in TABLE V, N can be the number of any amino acid residue within the HNH domain).
[0162] In one embodiment, HNH’s active site comprises amino acid residues HX1X2X3H (shown in bold in TABLE V), where Xi. X2. and X3 can be any amino acid.
[0163] In one embodiment, HNH’s active site comprises the amino acid residues HX1X2X3HX4 (shown in bold in TABLE V), where Xi, X2, and X3 can be any amino acid, and X4 can be G, N, or K.
[0164] In one embodiment, HNH’s active site comprises amino acid residues H54, H58, and G59 (shown in bold in TABLE V).
[0165] In one embodiment, HNH’s active site comprises amino acid residues H27, H59, and H63 (see FIG. 10B).
[0166] In one embodiment, HNH’s active site comprises amino acid residues H40, H64, and H68 (see FIG. 10B)..
[0167] In one embodiment, one or more amino acid residues in the HNH’s active site is mutated. In one embodiment, two or more amino acid residues in the HNH’s active site is mutated.In one embodiment, a mutation in HNH’s active site abolishes HNH’s catalytic activity.TABLE V: HNH nuclease domain (SpTasH (Salicola phage CGphi29 NC_020844))1 M N K Q V L K E Q A S H C E I T G A P L 201 ATGAACAAGCAAGTACTCAAAGAGCAAGCATCACACTGCGAAATCACCGGCGCACCGCTG 6021 A G L P E L V D V D R I T E R F Q G G T 4061 GCCGGCCTCCCTGAGCTGGTTGACGTTGATCGCATCACCGAGCGTTTCCAAGGCGGAACG 120 41 Y T P D N T R V L T P R A H M E R H G I 60121 TACACCCCTGACAATACTCGTGTATTGACACCACGAGCCCACATGGAGCGCCACGGCATT 18061 L R E R D ( SEQ ID NO: 9 ) 65181 TTGCGCGAGCGTGAC ( SEQ ID NO: 10 ) 195
[0168] The consensus sequence for the HNH nuclease active site can be found in Table VI below and in FIG. 10B:TABLE VI: HNH nuclease catalytic site- - - -............. W. H......... H...[nhhhiohhhhhhhhiNmsisa... iiiiiiiSSssss..CONSENSUS (SEQ ID NO:30) L C H A K X H G 1 E P * * Catalytic residuesL / M x K / E X ■■ G / K X X P (SEQ ID NO:31) x K / E x HH G / K / N X X P (SEQ ID NO:32)— X X X | H | G / N (SEQ ID N0:33)
[0169] In one embodiment, the HNH nuclease domain comprises 5 contiguous amino acids of the of an amino acid sequence selected from the group consisting of SEQ ID Nos: 11-33.
[0170] In one embodiment, the HNH nuclease domain comprises 10 contiguous amino acids of an amino acid sequence selected from the group consisting of SEQ ID Nos: 11-33.Example Tas Polypeptide Domain Architectures
[0171] Provided a below are example Tas polypeptides representative of the domain architectures seen in TIGR systems and demonstrating the modular nature of their domain designs.TasA
[0172] In one embodiment, the RNA-guided TIGR-associated (Tas) proteins provided herein can be a TasA polypeptide that comprise an N-terminal coiled-coil motif adjacent to a C-terminal Nop domain. TasA polypeptides do not comprise an N-terminal nuclease domain. In one embodiment, the TasA protein has an amino acid sequence of SEQ ID NO: 1 (see, for example, TABLE I). In one embodiment, the RNA-guided TIGR-associated TasA protein is encoded by the nucleotide sequence of SEQ ID NO: 2. In one embodiment, the RNA-guided TIGR-associated TasA protein or active fragments or variants may thereof retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence-specific manner.
[0173] In one embodiment, an active variant of a Tas A protein may comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the amino acid sequence set forth as SEQ ID NO: 1.
[0174] In one embodiment, an active fragment of the Tas R protein comprises at least 5, 6, 7, 8, 9, 10 contiguous amino acid residues of the amino acid sequence set forth as SEQ ID NO: 3.TABLE I: FpTIQR (TasA). JAQLWX010000047)1 M D I Q K N R I R N I V G G„ I Y D I Q K 20 130 ATGGACATTCAAAAGAACAGGATTCGCAACATCGTTGGTGGCATCT ACGACATCC AGAAA 189 21 L R I A T G N R I V A S L R P G L V E E 40 190 CTCCGTATCGCTACTGGCAATCGCATCGTAGCCAGCCTGCGGCCAGGCCTGGTGGAGGAG 249 41 V K E G E E D T K Y L P A I L S E Y R R 60 250 GTCAAGG AAGGTGAGGAGGACACC AAGTACCTTCCTGCTATTCTTTCTGAGTACAGACGC 309 61 I T D Y F V S E F E G R G S I E K A I T 80 310 ATCACCGATTACTTCGTTTCCGAGTTTGAGGGTCGTGGAAGCATTGAAAAGGCCAT CACC 369 81 P N N P E Y I K S R L D Y D L V T S Y K 100 370 CCTAACAATCCGGAGTACATCAAGAGCCGCCTTGACTATGATCTGGTGACATCCTACAAG 429 101 R L L. E T E E. G L T K V A E R E V K A H 120 430 CGGCTGCTGGAGACGGAGGAGGGGCTCACGAAGGTTGCTGAGCGTGAGGTCAAAGCACAT 489 121 P M W D A F F A G V K G G G P L M S A V 140 490 CCAATGTGGGA’rGCCTTCTTTGCAGGGGTAAAAGGCTGCGGACCGCTCATGTCCGCCGTG 549 141 C L A Y F D P Y K A R H A S S F W R Y A 160 550 TGCCTCGCGTACTTCGATCCCTACAAGGCGCGCCACGCTTCCTCCTTCTGGCGCTATGCC 609 161 G L D V Q R D P D K D K M R G V G K W Y 180 610 GGCCTGG ACGTGCAGCGCGACCCTGACAAGGAT AAGATGCGTGGCGTTGGCAAGTGGTAC 669 181 T E E R P Y I D K D G K E Q M K K S E T 200 670 ACTGAGGAGCGCCCTTACATCGATAAGGACGGGAAGGAGCAGATGAAGAAGTCGCTCACC 729 201 Y N P F L K T K L V G V L G S A F L R A 220 730 TACAACCCGTTCCTCAAAACCAAGCTGGTTGGTGTCCTCGGCTCTGCATTCCTGCGGGCC 789 221 K D S Y Y G K V Y Y D Y K N R L D N R D 240 790 AAAGACTCCTACTATGGTAAGGTTTACTAGGATTACAAAAATCGCCTGGATAACCGGGAT 849 241 E D s’ 0 P V K H R M A T R Y A V K M F 260 850 GAGGACTTCTCCCCTATTGTTAAGCACCGCATGGCCACCCGTTACGCCGTGAAGATGTTC 909 261 L R 0 M W V V W R E L E G L E V T E P Y 280 910281 E ' A K L G K K? H 8 8 5 *. (SEQ ID NO: 1 ) 294 970 GAG CGCTAAGTTGGGTCACAAGCCCCACCACTCATCTTAA (SEQ ID NO: 2 ) 1011 Coiled-coil domain: amino acids 1-121RNA binding domain: amino acids 122-271Kpp Protein domain: 1 -273Length: 2S6 amino acidsTasR
[0175] In one embodiment, the Tas polypeptide is a TasR polypeptide comprising a RuvC domain. In one embodiment, the Tas polypeptide is a TasR polypeptide comprising an N-terminal Ruv-C domain, followed by a coiled-coil structural motif, and a C-terminal Nop domain. In one embodiment, the TasR protein has an amino acid sequence of SEQ ID NO: 3 (see, for example, TABLE II). In one embodiment, the RNA-guided TIGR-associated TasR protein is encoded bythe nucleotide sequence of SEQ ID NO: 4. In one embodiment, the RNA-guided TIGR-associated TasR protein or active fragments or variants thereof retain the ability to bind and cleave a target nucleotide sequence in an RNA-guided sequence-specific manner. In an embodiment, the TasR polypeptide lacks a DNA binding domain.
[0176] In one embodiment, the RuvC domain(s) in the TasR polypeptide(s) is catalytically active. In one embodiment, the RuvC domain(s) in the TasR polypeptide(s) is not catalytically active (i.e., is catalytically inactive). In one embodiment, the RuvC domain includes one or more mutations that modify the activity of the RuvC domain. In one embodiment, the RuvC domain includes one or more mutations that decrease or eliminate an activity of the RuvC domain. In one embodiment, the RuvC domain includes one or more mutations that decrease or eliminate nuclease activity of the RuvC domain. In one embodiment, the RuvC domain includes one or more mutations that increase the activity of the RuvC domain.
[0177] In one embodiment, one or more Tas polypeptides, such as one or more TasR polypeptides, can interact with a recombinase polypeptide. In one embodiment one or more Tas polypeptides can recruit a recombinase polypeptide. In one embodiment, the Tas polypeptide contains a helix. In one embodiment, the helix of a first Tas polypeptide interacts with a helix of a second Tas polypeptide so as to form a dimer. In one embodiment, a Tas polypeptide dimer is formed with head to tail. In one embodiment, each RNA binding domain of the Tas polypeptides interacts with a recombinase polypeptide. In one embodiment, an RNA binding domain of a TasR polypeptide interacts with a tyrosine recombinase polypeptide. In one embodiment, the TasR polypeptide interacts with a tyrosine recombinase domain.
[0178] In one embodiment, one or more Tas polypeptides, such as one or more TasR polypeptides can couple with or are coupled to a recombinase polypeptide. This is described in greater details with respect to the recombinase polypeptides.
[0179] In one embodiment, an active variant of a Tas R protein may comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the amino acid sequence set forth as SEQ ID NO: 3. In one embodiment, an active fragment of the Tas R protein comprises at least 5, 6, 7, 8, 9, 10 contiguous amino acid residues of the amino acid sequence set forth as SEQ ID NO: 3.
[0180] In one embodiment, the TasR polypeptide is selected from any one of SEQ ID NOs: 21777 to 22022. In one embodiment, the TasR polypeptide is selected from anyone of SEQ ID NOs: 22023-22148.
[0181] The cryo-EM structure shows the TasR protein folds into a canonical Nop domain and forms a C2-symmetric dimer via a coiled-coil domain in a similar manner to Nop5 of the archaeal box C / D snoRNP (Fig. 6A, FIG. 19C). The TasR dimer binds one copy of the target DNA and one copy of the 36-nt tigRNA (Fig. 6A).TABLE II:archaeon isolate LB CRA 1 QQVH0100QQQ6Accession No.; RDJ35482.1I M D R D I T I L A V D W S H E E R K L A 201 ATGGATAAOGATAYCACAATATTGGCGGTCGATTGGTCTCATGAAGAAOGTAAACTTGCA 60 21 1 F D G K K I R K K L P E P S §. » V I J. 4061 ATrTT'TGATGGCAAGAAGATAAGGAAAAAACTTCCTGAGCCTTCTAGCGATGTAATAATT 120 41 V A E N I P fi K Y A ^ P F I E V G A K V 60121 GTGGCTGAAAA’TATTCCTCAAAAATATGCCGCTCCGTTCATCGAGGTAGGGGCAAAAGTT 180 61 A R C S T H A T A D A R K » Y Q K g V D 80181 CTTAGATGTTCGACAAACGCCACAGCAGATGCAAGAAAAAACTATCAAAAGAAAGTCGAT 24081 A & F A K K D E N D 8 K V I W A L Y 0 T 100241 &7TGCTTTTGCGAAAAACGACGAAAATGATTCTAAAGTGATTTGGGCGTrATATCAAACA 300101 H P E L F R E M K X, E g L S f. Y £ A I 120301 CACCCAGAATTG'n’TCGTGAAA^GAAGCTTGAGaCACCCOTTCTTC’rTAYTATGCTATT 300121 F K 8 Y Q E V K I R T G N R L Y S O R T 140361 TTTAAAGATTACCAAGAGGTAAGAAl'TAGAACTGGCAATCGACTCTAxTCAGATCGCACT 420141 0 A M E £ F E K I V K. £ G E H E L K & A 160421 GAT£CAATGGAARAGIT:TftTAAAATAGYYAAGAAAGGYGAACATGAACYAAAAW5GCT 480161 V D K R I, E K H F V Y T Q W L Q 0 I K G 180481 G1L'GAT> AAAGAA1:TAGAAAAATCATCCGG17J'1A1ACGCAGTGGTTACAGCATATTAAAGGT 840181 I G P V y A G §. L 1 5 L X G D I D R F » 200541 Al’l’GGCCCGGTCGTTGCCGGGGGATTGAlATCGTTxAATTGGAGATATAGATCGCTTGGAC 600201 0 V S K L W A Y A G Y S V E N G K V Q K 220601 AGCGTTTCTAAATT& TGGGCCTATGCAGGTTATAGTGTEGATAATGGAAAAGTGCAAAAG 660221 R K & G V A S R W K N K 1 R T H C Y K I 240661 CGlAAAAAAGGTGTAGCETCAAATTGGAAAAATAAAAEi’CGCACACAETGTTACAAOATT 720241 Y D S F X K Q R T S V Y R E L Y D A E K 2S0721 GTAGATTCATTTATCAAACAACGaACTTCAG'rTTATCGTGAGTTGTATGA-fGCCGAAAAA 780261 A R Q R P K V E S & G H A H K R A V R K 280781 GCTGGTCAACGTCCAAAGGYAGAATOCGAYGGTGATGCTCACAACCGTGC'rGTAAGAAAA 840281 V A K V F 1 Q H "i W V V S R £ L A G F S 300841 GTTGCTAAGGTATTTTTACAACATTATTGGGTAGTTTCTAGAGAATTAQCAGGGTTTTCT 900301 V S K P W I L E B G G H V D Y I K P P R 320901 GTGTCTAAACCTTGGATTTTAGAACATGGTGGGGACGFTGATTACATCAAACCACCAOAT 960321 W W K V E X K P C K P *. (SEQ ID NG: 3) 332961 TGGAATAAAGTAGAAATTAAGCCATGCAAGCCCTGA (SEC) TO NO: 4 ) 996 fexG domain; amino acids 1.-111 (in bold)Coiled-cail domain; amino acids 112-166RN A binding domain: amino acids 167- 298Non. domain: amino acids 112-298Length; 331 amino acids
[0182] In one embodiment, the Tas polypeptide includes or consists of a polypeptide that is about 90% to 100% identical to any one of SEQ ID NO. 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 90%, 90.5%, 91%, 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%,97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 90% to 90.5%, 91%, 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 90.5% to 91%, 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 91% to 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 91.5% to 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 92% to 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 92.5% to 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 93% to 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 93.5% to 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 94% to 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 94.5% to 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical toany one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 95% to 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 95.5% to 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 96% to 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 96.5% to 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 97% to 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 97.5% to 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 98% to 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 98.5% to 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 99% to 99.5%, or 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 99.5% to 100% identical to any one of SEQ ID NO: 22199-23909 or a fragment thereof.
[0183] In one embodiment, the Tas polypeptide includes or consists of a polypeptide that is about 90% to 100% identical to any one of SEQ ID NO. 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 90%, 90.5%, 91%, 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of apolypeptide having a sequence that is 90% to 90.5%, 91%, 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 90.5% to 91%, 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 91% to 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 91.5% to 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 92% to 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 92.5% to 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 93% to 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 93.5% to 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 94% to 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 94.5% to 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 95% to 95.5%, 96%,96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 95.5% to 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 96% to 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 96.5% to 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 97% to 97.5%, 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 97.5% to 98%, 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 98% to 98.5%, 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 98.5% to 99%, 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 99% to 99.5%, or 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof. In one embodiment, the Tas polypeptide includes or consists of a polypeptide having a sequence that is 99.5% to 100% identical to any one of SEQ ID NO: 23910-23938 or a fragment thereof.TasH
[0184] In one embodiment, the RNA-guided TIGR-associated (Tas) proteins provided herein can be a TasH protein comprising an N terminal RuvC nuclease domain. In one embodiment, the TasH protein has an amino acid sequence of SEQ ID NO: 5 (see, for example, TABLE III). In one embodiment, the RNA-guided TIGR-associated TasH protein is encoded by the nucleotide sequence of SEQ ID NO: 6. In one embodiment, the RNA-guided TIGR-associated TasH proteinor active fragments or variants thereof retain the ability to bind and cleave a target nucleotide sequence in an RNA-guided sequence-specific manner.
[0185] In one embodiment, the TasH polypeptide lacks a DNA binding domain. In one embodiment, the HNH domain includes one or more mutations that modify the activity of the HNH domain. In one embodiment, the HNH domain includes one or more mutations that decrease or eliminate an activity of the HNH domain. In one embodiment, the HNH domain includes one or more mutations that decrease or eliminate nuclease activity of the HNH domain. In one embodiment, the HNH domain includes one or more mutations that increase the activity of the HNH domain.
[0186] In one embodiment, an active variant of a Tas H protein comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the amino acid sequence set forth as SEQ ID NO: 5. In one embodiment, an active fragment of the Tas R protein comprises at least 5, 6, 7, 8, 9, 10 contiguous amino acid residues of the amino acid sequence set forth as SEQ ID NO: 5.TABLE III: TasH (Salicola phage CGphi29)Accession No.: YP_007673700.11M N K Q V L K E Q A S H C E I T G A P L1ATGAACAAGCAAGTACTCAAAGAGCAAGCATCACACTGCGAAATCACCGGCGCACCGCTG21 460 A G L P E L V D V D R I T E R F Q G G T61 120 GCCGGCCTCCCTGAGCTGGTTGACGTTGATCGCATCACCGAGCGTTTCCAAGGCGGAACG41Y T P D N T R V L T P R A |H] M E R |H] |G| I121TACACCCCTGACAATACTCGTGTATTGACACCACGAGCCCACATGGAGCGCCACGGCATT „ „ 61 o 0 181 L R E R D Q W L E E L K A M M D D R A Q2 4 QTTGCGCGAGCGTGACCAATGGCTGGAAGAACTTAAGGCCATGATGGACGACCGGGCGCAA81Q QT M K V V M K M N N Q L L A Y Q R Q T D2413 Q 0ACCATGAAGGTTGTGATGAAGATGAACAACCAGCTTCTTGCGTACCAGCGGCAGACCGAT1012 QH A R Q S T E Q F L Q D T L D A S N K R301 360 CACGCACGCCAGAGCACCGAGCAATTCCTGCAAGACACGCTGGACGCATCCAATAAGCGC121± 4 QL A Q I D R E V T K H I K H A K D P L A3614 2 QCTTGCTCAGATTGACCGGGAAGTGACCAAGCACATCAAGCACGCAAAAGACCCGCTTGCT, _. 141 16n0 Q A A M G V P G V G P I T V A G L Q T Y421 4n o8n0 CAGGCGGCCATGGGCGTGCCTGGTGTTGGCCCTATCACTGTGGCGGGTTTGCAAACGTAC.161 ioonl) V D L E K A K S A S A L W A Y I G I D K4815 4 QGTTGATCTCGAAAAGGCCAAGTCAGCCTCTGCGCTGTGGGCATACATCGGGATCGACAAG | 181Q 0541 P S H D R Y T K G E A G G G N K T L R T 600 CCTTCACACGACCGTTACACCAAAGGCGAGGCGGGCGGAGGAAACAAGACGCTGCGCACC2012 2 QM V W N M A N S M I K N R K C P Y R T V601 660221 ATGGTATGGAACATGGCAAACAGCATGATTAAGAATCGCAAGTGCCCTTACCGCACTGTG 240 661 Y E Q T K E R L A V S E K V T K S R N T 720 241 TATGAGCAGACCAAGGAGCGGCTTGCGGTATCCGAAAAGGTCACTAAGTCCCGCAACACG 2 60 721 Q G Q L I E C A W K D T K P S H R H G A 780 2 61 CAAGGGCAGTTGATCGAATGCGCGTGGAAGGACACCAAGCCGAGCCACCGGCACGGGGCA 280 781 A L R A V M K H F L A D Y W F V G R E L 840 281 GCGCTCCGGGCAGTGATGAAGCACTTTCTGGCAGATTACTGGTTCGTAGGCCGTGAGCTT 300 841 A G L D T R P L Y V Q E K L G H T G I V 900 301 GCCGGACTGGATACGCGCCCGCTGTACGTGCAGGAAAAACTGGGGCACACTGGCATTGTC 310 901 Q P Q E R G W E W * ( SEQ ID NO: 5 ) 930 CAGCCGCAAGAGCGCGGTTGGGAATGGTGA ( SEQ ID NO: 6 )HNH nuclease domain: amino acids 1-65 (in bold)Coiled-coil domain: amino acids 66-137DNA binding domain: amino acids 137-281X: denotes amino acid residues in the active siteNop domain: amino acids 66-281Length: 309 amino acidsTABLE IV: PITas (gut microbiome of Peromyscus leucopus JAAGKNO 10000977)1 M T T N K K T C A S Q C I T D T Q G S S 201 ATGACCACGAACAAAAAGACTTGCGCAAGCCAGTGCATCACCGATACCCAGGGTTCCAGC 6021 A C A P T L D S T H S A R D T Q F K R G 4061 GCTTGCGCACCAACACTTGACAGCACCCACTCAGCTCGCGATACCCAGTTCAAGCGCGGT 12041 A V T N L A G N Q Y C H D T Q P A I V S 60121 GCTGTCACCAATCTTGCCGGAAACCAGTACTGTCACGATACCCAGCCCGCTATCGTTTCC 180 61 G V A V V G A D E T T P T T I P I T Q S 80181 GGCGTTGCTGTGGTAGGAGCAGACGAGACCACGCCGACGACGATCCCCATTACGCAATCG 240 81 S R L A Y P I A D P M L F T L A Q T L Q 100241 TCTCGTCTGGCCTATCCCATTGCTGACCCCATGCTATTCACCCTCGCGCAAACACTCCAG 300101 D Y E T L R I A E E H R L R I F S T P S 120301 GATTACGAGACGCTGCGTATCGCGGAGGAACACCGGTTGCGTATCTTCTCAACGCCTTCC 360121 D V P D E D G V C R G F G Y A K D S N E 140361 GACGTGCCCGACGAGGACGGTGTATGCCGTGGCTTCGGTTACGCGAAGGATTCCAATGAA 420 141 V Q V V K G L I D P L K D L E H R T V L 160421 GTGCAGGTCGTCAAAGGGTTGATCGACCCGTTGAAGGACTTGGAACACCGCACGGTGCTC 480161 S L Q K R M R V N P I W P Y F K D V K G 180481 TCATTGCAGAAGCGTATGCGCGTGAACCCGATCTGGCCGTATTTCAAAGATGTGAAAGGC 540 181 V G E K T L A R L M A C I G D P Y L R P 200541 GTCGGTGAGAAAACACTGGCACGTCTCATGGCGTGCATCGGTGACCCGTATCTGCGTCCA 600201 L D D G S Y T P R T V S Q L W A Y C G M 220601 CTGGACGATGGTTCGTATACGCCTCGCACAGTGAGTCAACTGTGGGCGTACTGCGGTATG 660221 H T M P N K D G E I I A A K R M K G V Q 240661 CACACCATGCCGAACAAGGATGGTGAGATCATCGCGGCGAAACGCATGAAGGGCGTGCAG 720 241 A N W N T E A K T R L F L L S Q G L L R 260721 GCGAACTGGAACACGGAGGCGAAAACCAGACTGTTCCTCTTGTCGCAAGGATTGCTCAGG 780261 Q G I R K D K D G N Q F A V T P Y G Q L 280781 CAGGGGATTCGCAAGGACAAGGACGGCAACCAGTTTGCGGTAACACCTTACGGCCAGTTG 840281 Y L D R R A R T A V T H P E W N P G H G 300841 TATCTCGACCGTCGTGCCCGCACCGCTGTGACACACCCTGAATGGAATCCGGGCCATGGG 900301 L N D A L R I M G K E L L K Q L W R A A 320901 TTGAACGATGCGCTCAGGATCATGGGCAAGGAACTGCTCAAACAGTTGTGGCGTGCCGCC 960321 R E I H T G I P M D V D T S K V N E L E 340961 CGTGAAATCCATACGGGCATTCCCATGGACGTGGACACGTCCAAAGTCAACGAACTCGAA 1020 341 E T A *. ( SEQ ID NO: 7 ) 344 1021 GAAACCGCATGA ( SEQ ID NO: 8 ) 1032Tas Polypeptide Modifications
[0187] The Tas polypeptides described above may comprise one or more modifications that improve stability, reduce immunogenicity or improve localization of the polypeptide within the cell. The modular nature of the Tas polypeptides also enables swapping of domains between different Tas orthologs to unlock new functionalities. The compact and modular nature of the Tas polypeptides also creates the ability to covalently add additional functional domains to extend the gene editing capabilities and functionalities of the TIGR system.
[0188] Recombinant Tas endonucleases can be designed for expression in eukaryotic cells. The compactness and modular structure of Tas proteins enable the creation of recombinant Tas fusion proteins. For example, the Tas protein can be engineered to include a localization signal to facilitate the transport of the endonuclease to genomic nucleotide sequences in the nucleus or in an organelle, e.g., mitochondria or chloroplast. In a different form, the amino acid sequence of a Tas protein can be codon-optimized for enhanced expression in humans cells.Localization Modifications
[0189] Modifications that improve localization of the Tas polypeptide may include the addition of nuclear locations signal(s), localization signal sequences, and cell uptake sequences, and cell uptake signals.Subcellular Localization Sequences
[0190] In one embodiment, a heterologous polypeptide (a fusion partner) provides for subcellular localization, i.e., the heterologous polypeptide contains a subcellular localization sequence (e.g., a nuclear localization signal (NLS) for targeting to the nucleus, a sequence to keep the fusion protein out of the nucleus, e.g., a nuclear export sequence (NES), a sequence to keep the fusion protein retained in the cytoplasm, a mitochondrial localization signal for targeting to the mitochondria, a chloroplast localization signal for targeting to a chloroplast, an ER retention signal, and the like). In one embodiment, a Tas fusion polypeptide does not include an NLS so that the protein is not targeted to the nucleus (which can be advantageous, e.g., when the target nucleic acid is an RNA that is present in the cytosol). In one embodiment, the heterologous polypeptide can provide a tag (i.e., the heterologous polypeptide is a detectable label) for ease of tracking and / or purification (e.g., a fluorescent protein, e.g., green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), mCherry, tdTomato, and the like; a histidine tag, e.g., a 6XHis tag; a hemagglutinin (HA) tag; a FLAG tag; a Myc tag; and the like).Nuclear Localization Signal (NLS)
[0191] In one embodiment, a Tas protein (e.g., a wild type Tas protein, a variant Tas protein, a fusion Tas protein, a dTas protein, and the like) includes (is fused to) a nuclear localization signal (NLS) (e.g., in one embodiment 2 or more, 3 or more, 4 or more, or 5 or more NLSs). Thus, in one embodiment, a Tas polypeptide includes one or more NLSs (e.g., 2 or more, 3 or more, 4 or more, or 5 or more NLSs). In one embodiment, one or more NLSs (2 or more, 3 or more, 4 or more, or5 or more NLSs) are positioned at or near (e.g., within 50 amino acids of) the N-terminus and / or the C-terminus. In one embodiment, one or more NLSs (2 or more, 3 or more, 4 or more, or 5 or more NLSs) are positioned at or near (e.g., within 50 amino acids of) the N-terminus. In one embodiment, one or more NLSs (2 or more, 3 or more, 4 or more, or 5 or more NLSs) are positioned at or near (e.g., within 50 amino acids of) the C-terminus. In one embodiment, one or more NLSs (3 or more, 4 or more, or 5 or more NLSs) are positioned at or near (e.g., within 50 amino acids of) both the N-terminus and the C-terminus. In one embodiment, an NLS is positioned at the N-terminus and an NLS is positioned at the C-terminus.
[0192] In one embodiment, a Tas protein (e.g., a wild type Tas protein, a variant Tas protein, a fusion Tas protein, a dTas protein, and the like) includes (is fused to) between 1 and 10 NLSs (e.g., 1-9, 1-8, 1-7, 1-6, 1-5, 2-10, 2-9, 2-8, 2-7, 2-6, or 2-5 NLSs). In one embodiment, a Tas protein (e.g., a wild type Tas protein, a variant Tas protein, a fusion Tas protein, a dTas protein, and the like) includes (is fused to) between 2 and 5 NLSs (e.g., 2-4, or 2-3 NLSs).
[0193] Non-limiting examples of NLSs include an NLS sequence derived from: the NLS of the SV40 virus large T-antigen, having the amino acid sequence PKKKRKV (SEQ ID NO: 22150); the NLS from nucleoplasmin (e.g., the nucleoplasmin bipartite NLS with the sequence KRPAATKKAGQAKKKK (SEQ ID NO:22151)); the c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO:22152) or RQRRNELKRSP (SEQ ID NO:22153); the hRNPAl M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO:22154); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO:22155) of the IBB domain from importin-alpha; the sequences VSRKRPRP (SEQ ID NO:22156) and PPKKARED (SEQ ID NO:22157) of the myoma T protein; the sequence PQPKKKPL (SEQ ID NO:22158) of human p53; the sequence SALIKKKKKMAP (SEQ ID NO:22159) of mouse c-abl IV; the sequences DRLRR (SEQ ID NO:22160) and PKQKKRK (SEQ ID NO:22161) of the influenza virus NS1; the sequence RKLKKKIKKL (SEQ ID NO:22162) of the Hepatitis virus delta antigen; the sequence REKKKFLKRR (SEQ ID NO:22163) of the mouse Mxl protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO:22164) of the human poly(ADP-ribose) polymerase; and the sequence RKCLQAGMNLEARKTKK (SEQ ID NO:22165) of the steroid hormone receptors (human) glucocorticoid. In general, NLS (or multiple NLSs) are of sufficient strength to drive accumulation of the Tas protein in a detectable amount inthe nucleus of a eukaryotic cell. Detection of accumulation in the nucleus may be performed by any suitable technique. For example, a detectable marker may be fused to the Tas protein such that location within a cell may be visualized. Cell nuclei may also be isolated from cells, the contents of which may then be analyzed by any suitable process for detecting protein, such as immunohistochemistry, Western blot, or enzyme activity assay. Accumulation in the nucleus may also be determined indirectly.Chloroplast Transit Peptide (CTP)
[0194] In some case, a Tas fusion polypeptide of the present disclosure comprises: a) a Tas polypeptide of the present disclosure; and b) a chloroplast transit peptide. Thus, for example, a Tas polypeptide / guide RNA complex can be targeted to the chloroplast. In one embodiment, this targeting may be achieved by the presence of an N-terminal extension, called a chloroplast transit peptide (CTP) or plastid transit peptide. Chromosomal transgenes from bacterial sources must have a sequence encoding a CTP sequence fused to a sequence encoding an expressed polypeptide if the expressed polypeptide is to be compartmentalized in the plant plastid (e.g. chloroplast). Accordingly, localization of an exogenous polypeptide to a chloroplast is often 1 accomplished by means of operably linking a polynucleotide sequence encoding a CTP sequence to the 5' region of a polynucleotide encoding the exogenous polypeptide. The CTP is removed in a processing step during translocation into the plastid. Processing efficiency may, however, be affected by the amino acid sequence of the CTP and nearby sequences at the amino terminus (NH2 terminus) of the peptide. Other options for targeting to the chloroplast which have been described are the maize cab-m7 signal sequence (U. S. Pat. No. 7,022,896, WO 97 / 41228) a pea glutathione reductase signal sequence (WO 97 / 41228) and the CTP described in US2009029861.Mitochondrial Transit Peptide (MTP)
[0195] A “mitochondrial transit peptide” or “MTP” refers to a peptide or fragment of amino acids that can be attached to a separate molecule in order to transport the molecule in the mitochondria. For example, an MTP can be attached to a nuclease, such as an engineered Tas polypeptide, in order to transport the engineered Tas polypeptide into mitochondria (, see, for example, the published U. S. Patent Application No. 2023 / 029,559, the content of which is incorporated by reference herein in its entirety). MTPs can consist of an alternating pattern of hydrophobic and positively charged amino acids to form what is called amphipathic helix. In one embodiment, the mitochondrial transit peptide has the amino acid sequence of:(SEQ ID NO: 22166)MSVLTPLLLRGLTGSARRLPVPRAKIHSLPPEGKLMAS MTPTRVLASRLASQMAASAKVARPAVRVAQVSKRTIQTGS PLQTLKRTQMTSIVNATTRQAFQProtein Transduction Domain
[0196] In one embodiment, a Tas fusion polypeptide includes a “Protein Transduction Domain” or PTD (also known as a CPP — cell penetrating peptide), which refers to a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates traversing a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. A PTD attached to another molecule, which can range from a small polar molecule to a large macromolecule and / or a nanoparticle, facilitates the molecule traversing a membrane, for example going from extracellular space to intracellular space, or cytosol to within an organelle. In one embodiment, a PTD is covalently linked to the amino terminus a polypeptide (e.g., linked to a wild type Tas to generate a fusion protein, or linked to a variant Tas protein such as a dTas, a nickase Tas, or a fusion Tas protein, to generate a fusion protein). In one embodiment, a PTD is covalently linked to the carboxyl terminus of a polypeptide (e.g., linked to a wild type Tas to generate a fusion protein, or linked to a variant Tas protein such as a dTas, a nickase Tas, or a fusion Tas protein to generate a fusion protein). In one embodiment, the PTD is inserted internally in the Tas fusion polypeptide (i.e., is not at the N- or C-terminus of the Tas fusion polypeptide) at a suitable insertion site. In one embodiment, a subject Tas fusion polypeptide includes (is conjugated to, is fused to) one or more PTDs (e.g., two or more, three or more, four or more PTDs). In one embodiment, a PTD includes a nuclear localization signal (NLS) (e.g., in one embodiment 2 or more, 3 or more, 4 or more, or 5 or more NLSs). Thus, in one embodiment, a Tas fusion polypeptide includes one or more NLSs (e.g., 2 or more, 3 or more, 4 or more, or 5 or more NLSs). In one embodiment, a PTD is covalently linked to a nucleic acid (e.g., a Tas guide nucleic acid, a polynucleotide encoding a Tas guide nucleic acid, a polynucleotide encoding a Tas fusion polypeptide, a donor polynucleotide, etc.). Examples of PTDs include but are not limited to a minimal undecapeptide protein transduction domain (corresponding to residues 47-57 of HIV-1 TAT comprising YGRKKRRQRRR; SEQ ID NO:22167); a polyarginine sequence comprising a number of arginines sufficient to direct entry into a cell (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines); a VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96); an Drosophila Antennapediaprotein transduction domain (Noguchi et al. (2003) Diabetes 52(7): 1732-1737); atruncated human calcitonin peptide (Trehin et al. (2004) Pharm. Research 21: 1248-1256); polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008); RRQRRTSKLMKR (SEQ ID NO:22168); Transportan GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 22169); KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO:22170); and RQ1K1WFQNRRMKWKK (SEQ ID NO:22171). Exemplary PTDs include but are not limited to, YGRKKRRQRRR (SEQ ID NO:22167), RKKRRQRRR (SEQ ID NO:22172); an arginine homopolymer of from 3 arginine residues to 50 arginine residues; Exemplary PTD domain amino acid sequences include, but are not limited to, any of the following: YGRKKRRQRRR (SEQ ID NO:22167); RKKRRQRR (SEQ ID NO:22173); YARAAARQARA (SEQ ID NO:22174); THRLPRRRRRR (SEQ ID NO:22175); and GGRRARRRRRR (SEQ ID NO:22176). In one embodiment, the PTD is an activatable CPP (ACPP) (Aguilera et al. (2009) Integr Biol (Camb) June; 1(5-6): 371-381). ACPPs comprise a polycationic CPP (e.g., Arg9 or “R9”) connected via a cleavable linker to a matching polyanion (e.g., Glu9 or “E9”), which reduces the net charge to nearly zero and thereby inhibits adhesion and uptake into cells. Upon cleavage of the linker, the polyanion is released, locally unmasking the polyarginine and its inherent adhesiveness, thus “activating” the ACPP to traverse the membrane.Tas Domain Swapping
[0197] The present disclosure describes three primary Tas polypeptide domain architectures, Nop domain without a nuclease domain (TasA), Nop domain and a RuvC domain (TasR), and Nop domain with a HNH nuclease domain (TasH). The present disclosure likewise provides an extensive listing of representative Tas polypeptide species. In an embodiment, a reference Tas polypeptide is selected. Domains from other Tas polypeptide orthologs are then selected for modification of the Tas polypeptide and swapped out. For example, the reference Tas polypeptide may be a TasA to which a RuvC or HNH domain from another Tas ortholog. Likewise, the reference Tas polypeptide may be a TasR for which the native Nop domain is swapped out for a Nop domain from another TasA or TasH, or the reference Tas polypeptide is a TasH for which the native Nop domain is swapped out for a Nop domain from another TasA or TasH. The modifications may be more granular and involve swapping out only parts of another domain. For example, rather than swapping an entire Nop domain only the coiled-coil motif from one Tas polypeptide may be swapped into another. The proposed domain swaps are exemplary and any allpossible combinations of domain swaps are contemplated within the scope of this disclosure. Domain swaps may include the creation of new fusion proteins or the covalently linking of different domains to generated additional Tas polypeptidesFunctional Domain Modifications
[0198] The Tas polypeptides disclosed herein can be genetically engineered to generate recombinant Tas polypeptides that incorporate one or more heterologous polypeptides that confer additional functionality at the targeted nucleotide sequence. Complex formation of these recombinant Tas-associated polypeptides and an engineered TIGR guide molecule then can direct the sequence-specific binding of the complex and its associated heterologous polypeptide to a specific locus containing the targeted nucleotide sequence.
[0199] “Heterologous,” as used herein, means a nucleotide or polypeptide sequence that is not found in the native nucleic acid or protein, respectively. For example, in one embodiment, in a variant Tas protein of the present disclosure, a portion of naturally occurring Tas polypeptide (or a variant thereof) may be fused to a heterologous polypeptide (i.e. an amino acid sequence from a protein other than a Tas polypeptide or an amino acid sequence from another organism). As another example, a fusion Tas polypeptide can comprise all or a portion of a naturally occurring Tas polypeptide (or variant thereof) fused to a heterologous polypeptide, i.e., a polypeptide from a protein other than a Tas polypeptide, or a polypeptide from another organism. The heterologous polypeptide may exhibit an activity (e.g., enzymatic activity) that will also be exhibited by the variant Tas protein or the fusion Tas protein (e.g., biotin ligase activity; nuclear localization; etc.). A heterologous nucleic acid sequence may be linked to a naturally occurring nucleic acid sequence (or a variant thereof) (e.g., by genetic engineering) to generate a nucleotide sequence encoding a fusion polypeptide (a fusion protein).
[0200] In one embodiment, the Tas polypeptide can be an Nop domain of a Tas polypeptide. In one embodiment, the recombinant Tas polypeptide is noncovalently attached to one or more heterologous polypeptides. In one embodiment, the recombinant Tas polypeptide is covalently attached to one or more heterologous polypeptides to form a Tas fusion polypeptide where the N or C terminus of the Tas polypeptide is fused in frame with the one or more heterologous polypeptides. In one embodiment, the one or more heterologous polypeptides are inserted into the coding region of the recombinant Tas polypeptide.Tas Fusion Polypeptides
[0201] As noted above, in one embodiment, a Tas protein (e.g., a Tas protein with wild type cleavage activity; a variant Tas with reduced cleavage activity, e.g., a dTas or a nickase Tas; a chimeric Tas protein; etc.) is fused (conjugated) to a heterologous polypeptide (i.e., one or more heterologous polypeptides) that has an activity of interest (e.g., a catalytic activity of interest) to form a fusion protein. A heterologous polypeptide to which a Tas protein can be fused is referred to herein as a “fusion partner.” As used herein conjugation includes uses of linkers and other covalent attachment means as well as generation of straight recombinant fusion protein. In one embodiment, the Tas polypeptide can be covalently joined to a heterologous polypeptide using, for example, the SpyTag / SpyCatcher bioconjugation technology (see U. S. Pat. No. 9,547,003, the content of which is incorporated by reference herein in its entirety). Example linkers are described in a section below.
[0202] In one embodiment, the fusion partner can modulate transcription (e.g., inhibit transcription, increase transcription) of a target DNA. For example, in one embodiment the fusion partner is a protein (or a domain from a protein) that inhibits transcription (e.g., a transcriptional repressor, a protein that functions via recruitment of transcription inhibitor proteins, modification of target DNA such as methylation, recruitment of a DNA modifier, modulation of histones associated with target DNA, recruitment of a histone modifier such as those that modify acetylation and / or methylation of histones, and the like). In one embodiment, the fusion partner is a protein (or a domain from a protein) that increases transcription (e g., a transcription activator, a protein that acts via recruitment of transcription activator proteins, modification of target DNA such as demethylation, recruitment of a DNA modifier, modulation of histones associated with target DNA, recruitment of a histone modifier such as those that modify acetylation and / or methylation of histones, and the like). In one embodiment, the fusion partner is a reverse transcriptase. In one embodiment, the fusion partner is a base editor. In one embodiment, the fusion partner is a deaminase.
[0203] In one embodiment, a fusion Tas protein includes a heterologous polypeptide that has enzymatic activity that modifies a target nucleic acid (e g., nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimerforming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylase activity).
[0204] In one embodiment, a fusion Tas protein includes a heterologous polypeptide that has enzymatic activity that modifies a polypeptide (e.g., a histone) associated with a target nucleic acid (e.g., methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, derib osylati on activity, myristoylation activity or demyristoylation activity).
[0205] Examples of proteins (or fragments thereof) that can be used in increase transcription include, but are not limited to: transcriptional activators such as VP16, VP64, VP48, VP 160, p65 subdomain (e.g., from NFkB), and activation domain of EDLL and / or TAL activation domain (e.g., for activity in plants); histone lysine methyltransferases such as SET1A, SET1B, MLL1 to 5, ASH1, SYMD2, NSD1, and the like; histone lysine demethylases such as JHDM2a / b, UTX, JMJD3, and the like; histone acetyltransferases such as GCN5, PCAF, CBP, p300, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, SRC1, ACTR, Pl 60, CLOCK, and the like; and DNA demethylases such as Ten-Eleven Translocation (TET) dioxygenase 1 (TET1CD), TET1, DME, DML1, DML2, ROS1, and the like.
[0206] Examples of proteins (or fragments thereof) that can be used in decrease transcription include, but are not limited to: transcriptional repressors such as the Krtippel associated box (KRAB or SKD); K0X1 repression domain; the Mad mSIN3 interaction domain (SID); the ERF repressor domain (ERD), the SRDX repression domain (e.g., for repression in plants), and the like; histone lysine methyltransferases such as Pr-SET7 / 8, SUV4-20H1, RIZ1, and the like; histone lysine demethylases such as JMJD2A / JHDM3A, JMJD2B, IMJD2C / GASC 1, IMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID 1 C / SMCX, JARID1D / SMCY, and the like; histone lysine deacetylases such as HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HD AC 11, and the like; DNA methylases such as Hhal DNA m5c-methyltransf erase (M. Hhal), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3 a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plants), ZMET2, CMT1, CMT2 (plants), and the like; and periphery recruitment elements such as Lamin A, Lamin B, and the like.
[0207] In one embodiment, the fusion partner has enzymatic activity that modifies the target nucleic acid (e.g., ssDNA, or dsDNA). Examples of enzymatic activity that can be provided by the fusion partner include but are not limited to: nuclease activity such as that provided by a restriction enzyme (e.g., FokI nuclease), methyltransferase activity such as that provided by a methyltransferase (e.g., Hhal DNA m5c-methyltransferase (M. Hhal), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), MET1, DRM3 (plants), ZMET2, CMT1, CMT2 (plants), and the like); demethylase activity such as that provided by a demethylase (e.g., Ten-Eleven Translocation (TET) dioxygenase 1 (TET1CD), TET1, DME, DML1, DML2, ROS1, and the like), DNA repair activity, DNA damage activity, deamination activity such as that provided by a deaminase (e.g., a cytosine deaminase enzyme such as rat APOBEC1), dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer forming activity, integrase activity such as that provided by an integrase and / or resolvase (e.g., Gin invertase such as the hyperactive mutant of the Gin invertase, GinH106Y; human immunodeficiency virus type 1 integrase (IN); Tn3 resolvase; and the like), transposase activity, recombinase activity such as that provided by a recombinase (e.g., catalytic domain of Gin recombinase), polymerase activity, ligase activity, helicase activity, photolyase activity, and glycosylase activity).
[0208] In one embodiment, the fusion partner has enzymatic activity that modifies a protein associated with the target nucleic acid (e.g., ssDNA, or dsDNA) (e.g., a histone, an RNA binding protein, a DNA binding protein, and the like). Examples of enzymatic activity (that modifies a protein associated with a target nucleic acid) that can be provided by the fusion partner include but are not limited to: methyltransferase activity such as that provided by a histone methyltransferase (HMT) (e.g., suppressor of variegation 3-9 homolog 1 (SUV39H1, also known as KMT1A), euchromatic histone lysine methyltransferase 2 (G9A, also known as KMT1C and EHMT2), SUV39H2, ESET / SETDB1, and the like, SET1A, SET1B, MLL1 to 5, ASH1, SYMD2, NSD1, DOT1L, Pr-SET7 / 8, SUV4-20H1, EZH2, RIZ1), demethylase activity such as that provided by a histone demethylase (e.g., Lysine Demethylase 1A (KDM1A also known as LSD1), JHDM2a / b, JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, J ARID ID / SMC Y, UTX, JMJD3, and the like), acetyltransferase activity such as that provided by a histone acetylase transferase (e.g., catalytic core / fragment of the human acetyltransferase p300, GCN5, PCAF, CBP, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4,HB01 / MYST2, HM0F / MYST1, SRC1, ACTR, P160, CLOCK, and the like), deacetylase activity such as that provided by a histone deacetylase (e.g., HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, and the like), kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity, and demyristoylation activity.
[0209] In one embodiment, a suitable fusion partner is a dihydrofolate reductase (DHFR) destabilization domain (e.g., to generate a chemically controllable fusion Tas protein).
[0210] Additional suitable heterologous polypeptides include, but are not limited to, a polypeptide that directly and / or indirectly provides for increased or decreased transcription and / or translation of a target nucleic acid (e.g., a transcription activator or a fragment thereof, a protein or fragment thereof that recruits a transcription activator, a small molecule / drug-responsive transcription and / or translation regulator, a translation-regulating protein, etc.). Non-limiting examples of heterologous polypeptides to accomplish increased or decreased transcription include transcription activator and transcription repressor domains. In some such cases, a fusion Tas polypeptide is targeted by the guide nucleic acid (guide RNA) to a specific location (i.e., sequence) in the target nucleic acid and exerts locus-specific regulation such as blocking RNA polymerase binding to a promoter (which selectively inhibits transcription activator function), and / or modifying the local chromatin status (e.g., when a fusion sequence is used that modifies the target nucleic acid or modifies a polypeptide associated with the target nucleic acid). In one embodiment, the changes are transient (e.g., transcription repression or activation). In one embodiment, the changes are inheritable (e.g., when epigenetic modifications are made to the target nucleic acid or to proteins associated with the target nucleic acid, e.g., nucleosomal histones).
[0211] Non-limiting examples of heterologous polypeptides for use when targeting ssRNA target nucleic acids include (but are not limited to): splicing factors (e.g., RS domains); protein translation components (e.g., translation initiation, elongation, and / or release factors; e.g., eIF4G); RNA methylases; RNA editing enzymes (e.g., RNA deaminases, e.g., adenosine deaminase acting on RNA (ADAR), including A to I and / or C to U editing enzymes); helicases; RNA-binding proteins; and the like. It is understood that a heterologous polypeptide can include the entire protein or in one embodiment can include a fragment of the protein (e.g., a functional domain).
[0212] The heterologous polypeptide of a subject fusion Tas polypeptide can be any domain capable of interacting with ssRNA (which, for the purposes of this disclosure, includes intramolecular and / or intermolecular secondary structures, e.g., double-stranded RNA duplexes such as hairpins, stem-loops, etc.), whether transiently or irreversibly, directly or indirectly, including but not limited to an effector domain selected from the group comprising; Endonucleases (for example RNase 111, the CRR22 DYW domain, Dicer, and PIN (PilT N-terminus) domains from proteins such as SMG5 and SMG6); proteins and protein domains responsible for stimulating RNA cleavage (for example CPSF, CstF, CFIm and CFIIm); Exonucleases (for example XRN-1 or Exonuclease T); Deadenylases (for example HNT3); proteins and protein domains responsible for nonsense mediated RNA decay (for example UPF1, UPF2, UPF3, UPF3b, RNP Si, Y14, DEK, REF2, and SRml60); proteins and protein domains responsible for stabilizing RNA (for example PABP); proteins and protein domains responsible for repressing translation (for example Ago2 and Ago4); proteins and protein domains responsible for stimulating translation (for example Staufen); proteins and protein domains responsible for (e.g., capable of) modulating translation (e.g., translation factors such as initiation factors, elongation factors, release factors, etc., e.g., eIF4G); proteins and protein domains responsible for polyadenylation of RNA (for example PAP1, GLD-2, and Star-PAP); proteins and protein domains responsible for polyuridinylation of RNA (for example CI DI and terminal uridylate transferase); proteins and protein domains responsible for RNA localization (for example from IMP1, ZBP1, She2p, She3p, and Bicaudal-D); proteins and protein domains responsible for nuclear retention of RNA (for example Rrp6); proteins and protein domains responsible for nuclear export of RNA (for example TAP, NXF1, THO, TREX, REF, and Aly); proteins and protein domains responsible for repression of RNA splicing (for example PTB, Sam68, and hnRNP Al); proteins and protein domains responsible for stimulation of RNA splicing (for example Serine / Arginine-rich (SR) domains); proteins and protein domains responsible for reducing the efficiency of transcription (for example FUS (TLS)); and proteins and protein domains responsible for stimulating transcription (for example CDK7 and HIV Tat). Alternatively, the effector domain may be selected from the group comprising Endonucleases; proteins and protein domains capable of stimulating RNA cleavage; Exonucleases; Deadenylases; proteins and protein domains having nonsense mediated RNA decay activity; proteins and protein domains capable of stabilizing RNA; proteins and protein domains capable of repressing translation; proteins and protein domains capable of stimulating translation; proteins and protein'lldomains capable of modulating translation (e.g., translation factors such as initiation factors, elongation factors, release factors, etc., e.g., eIF4G); proteins and protein domains capable of polyadenylation of RNA; proteins and protein domains capable of polyuridinylation of RNA; proteins and protein domains having RNA localization activity; proteins and protein domains capable of nuclear retention of RNA; proteins and protein domains having RNA nuclear export activity; proteins and protein domains capable of repression of RNA splicing; proteins and protein domains capable of stimulation of RNA splicing; proteins and protein domains capable of reducing the efficiency of transcription; and proteins and protein domains capable of stimulating transcription. Another suitable heterologous polypeptide is a PUF RNA-binding domain, which is described in more detail in WO2012068627, which is hereby incorporated by reference in its entirety.
[0213] Some RNA splicing factors that can be used (in whole or as fragments thereof) as heterologous polypeptides for a fusion Tas polypeptide have modular organization, with separate sequence-specific RNA binding modules and splicing effector domains. For example, members of the Serine / Arginine-rich (SR) protein family contain N-terminal RNA recognition motifs (RRMs) that bind to exonic splicing enhancers (ESEs) in pre-mRNAs and C-terminal RS domains that promote exon inclusion. As another example, the hnRNP protein hnRNP Al binds to exonic splicing silencers (ESSs) through its RRM domains and inhibits exon inclusion through a C-terminal Glycine-rich domain. Some splicing factors can regulate alternative use of splice site (ss) by binding to regulatory sequences between the two alternative sites. For example, ASF / SF2 can recognize ESEs and promote the use of intron proximal sites, whereas hnRNP Al can bind to ESSs and shift splicing towards the use of intron distal sites. One application for such factors is to generate ESFs that modulate alternative splicing of endogenous genes, particularly disease associated genes. For example, Bcl-x pre-mRNA produces two splicing isoforms with two alternative 5' splice sites to encode proteins of opposite functions. The long splicing isoform Bcl-xL is a potent apoptosis inhibitor expressed in long-lived postmitotic cells and is up-regulated in many cancer cells, protecting cells against apoptotic signals. The short isoform Bcl-xS is a pro-apoptotic isoform and expressed at high levels in cells with a high turnover rate (e.g., developing lymphocytes). The ratio of the two Bcl-x splicing isoforms is regulated by multiple elements that are located in either the core exon region or the exon extension region (i.e., between the twoalternative 5' splice sites). For more examples, see W02010075303, which is hereby incorporated by reference in its entirety.
[0214] Further suitable fusion partners include, but are not limited to, proteins (or fragments thereof) that are boundary elements (e.g., CTCF), proteins and fragments thereof that provide periphery recruitment (e.g., Lamin A, Lamin B, etc.), protein docking elements (e.g., FKBP / FRB, Pill / Abyl, etc.).Nucleases
[0215] In one embodiment, a subject fusion Tas polypeptide comprises: i) a Tas polypeptide of the present disclosure; and ii) a heterologous polypeptide (a “fusion partner”), where the heterologous polypeptide is a nuclease. Suitable nucleases include, but are not limited to, a homing nuclease polypeptide; a FokI polypeptide; a transcription activator-like effector nuclease (TALEN) polypeptide; a MegaTAL polypeptide; a meganuclease polypeptide; a zinc finger nuclease (ZFN); an ARCUS nuclease; and the like. The meganuclease can be engineered from an LADLIDADG homing endonuclease (LHE). A megaTAL polypeptide can comprise a TALE DNA binding domain and an engineered meganuclease. See, e.g., WO 2004 / 067736 (homing endonuclease); Urnov et al. (2005) Nature 435:646 (ZFN); Mussolino et al. (2011) Nucle. Acids Res. 39:9283 (TALE nuclease); Boissel et al. (2013) Nucl. Acids Res. 42:2591 (MegaTAL).Reverse Transcriptases
[0216] In one embodiment, a subject fusion Tas polypeptide comprises: i) a Tas polypeptide of the present disclosure; and ii) a heterologous polypeptide (a “fusion partner”), where the heterologous polypeptide is a reverse transcriptase polypeptide. In one embodiment, the Tas polypeptide is catalytically inactive. Suitable reverse transcriptases include, e.g., a murine leukemia virus reverse transcriptase; a Rous sarcoma virus reverse transcriptase; a human immunodeficiency virus type I reverse transcriptase; a Moloney murine leukemia virus reverse transcriptase; and the like.Base Editors
[0217] In one embodiment, a Tas fusion polypeptide of the present disclosure comprises: i) a Tas polypeptide of the present disclosure; and ii) a heterologous polypeptide (a “fusion partner”), where the heterologous polypeptide is a base editor.
[0218] By “base editor (BE),” or “nucleobase editor (NBE)” is meant an agent that binds a polynucleotide and has nucleobase modifying activity. In various embodiments, the base editorcomprises a nucleobase modifying polypeptide (e.g., a deaminase) and a Tas polypeptide in conjunction with a guide polynucleotide (e.g., guide RNA). In various embodiments, the agent is a biomolecular complex comprising a protein domain having base editing activity, i.e., a domain capable of modifying a base (e.g., A, T, C, G, or U) within a nucleic acid molecule (e.g., DNA). In one embodiment, the Tas polypeptide is fused or linked to a deaminase domain. In one embodiment, the agent is a fusion protein comprising one or more domains having base editing activity. In another embodiment, the protein domains having base editing activity are linked to the guide RNA (e.g., via an RNA binding motif on the guide RNA and an RNA binding domain fused to the deaminase). In one embodiment, the domains having base editing activity are capable of deaminating a base within a nucleic acid molecule. In one embodiment, the base editor is capable of deaminating one or more bases within a DNA molecule. In one embodiment, the base editor is capable of deaminating a cytosine (C) or an adenosine (A) within DNA. In one embodiment, the base editor is capable of deaminating a cytosine (C) and an adenosine (A) within DNA. In one embodiment, the base editor is a cytidine base editor (CBE). In one embodiment, the base editor is an adenosine base editor (ABE). In one embodiment, the base editor is an adenosine base editor (ABE) and a cytidine base editor (CBE). In one embodiment, the base editor is a nuclease-inactive Tas fused to an adenosine deaminase. In one embodiment, the base editor is fused to an inhibitor of base excision repair, for example, a UGI domain, or a dISN domain. In one embodiment, the fusion protein comprises a Tas nickase fused to a deaminase and an inhibitor of base excision repair, such as a UGI or dISN domain. In other embodiments the base editor is an a basic base editor. Details of base editors are described in International PCT Application Nos. PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), each of which is incorporated herein by reference for its entirety. Also see Komor, A. C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N. M., et al., “Programmable base editing of A T to G C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A. C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C: G-to-T: A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017), and Rees, H. A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 December; 19(12):770-788. doi: 10.1038 / s41576-018-0059-1, the entire contents of which are hereby incorporated by reference.
[0219] In one embodiment, the base editor comprises a nuclease-inactive Tas polypeptide fused to a cytidine deaminase. In one embodiment, the base editor comprises a Tas polypeptide having nickase activity fused to a cytidine deaminase. In one embodiment, the base editor is fused to an inhibitor of base excision repair, for example, a UGI domain.
[0220] The term "uracil glycosylase inhibitor" or " UGI," as used herein, refers to a protein that is capable of inhibiting a uracil-DNA glycosylase base-excision repair enzyme.
[0221] In one embodiment, the base editor is capable of deaminating an adenosine (A) in DNA. In one embodiment, the base editor is a fusion protein comprising a Tas polypeptide fused to an adenine deaminase domain. In one embodiment, the base editor is a fusion protein comprising a TasA polypeptide fused to an adenine deaminase domain. In one embodiment, the base editor is a fusion protein comprising a Tas nickase polypeptide fused to an adenine deaminase domain.
[0222] Base Editing Activity: By “base editing activity” is meant acting to chemically alter a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, e.g., converting target C-Gto T-A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity, e.g., converting A-T to G-C. In another embodiment, the base editing activity is cytosine or cytidine deaminase activity, e.g., converting target C-G to T-A and adenosine or adenine deaminase activity, e.g., converting A-T to G-C.
[0223] Base Editor System: The term “base editor system” refers to a system for editing a nucleobase of a target nucleotide sequence. In various embodiments, the base editor (BE) system comprises (1) a polynucleotide programmable nucleotide binding domain (e.g., Tas polypeptide), a deaminase domain for deaminating nucleobases in the target nucleotide sequence; and (2) one or more guide polynucleotides (e.g., TIGR guide RNA) in conjunction with the Tas polypeptide. In various embodiments, the base editor (BE) system comprises a nucleobase editor domains selected from an adenosine deaminase or a cytidine deaminase, and a domain having nucleic acid sequence specific binding activity (Tas polypeptide). In one embodiment, the base editor system comprises (1) a base editor (BE) comprising a Tas polypeptide and a deaminase domain for deaminating one or more nucleobases in a target nucleotide sequence; and (2) one or more TIGR guide RNAs in conjunction with the Tas polypeptide. In one embodiment, the base editor is a cytidine base editor (CBE). In one embodiment, the base editor is an adenine or adenosine base editor (ABE). In oneembodiment, the base editor is an adenine or adenosine base editor (ABE) or a cytidine base editor (CBE).Adenosine deaminase
[0224] As used herein, the term “adenosine deaminase” or “adenosine deaminase protein” refers to a protein, polypeptide, or one or more functional domains of a protein or polypeptide that can catalyze hydrolytic deamination reaction to convert adenine (or the adenine portion of a molecule) to hypoxanthine (or the hypoxanthine portion of a molecule). In one embodiment, the adenine-containing molecule is adenosine (A) and the hypoxanthine-containing molecule is inosine (I). A suitable adenosine deaminase is any enzyme that is capable of deaminating adenosine in DNA. In one embodiment, the deaminase is a TadA deaminase.
[0225] In one embodiment, a suitable adenosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following amino acid sequence:(SEQ ID NO: 22177)MS EVE FS HE YWMRHAL T LAKRAWDE RE VP VGAVL V HNNRVI GEGWNRP I GRHDPTAHAE IMALRQGGLVM QN YRL I DAT L YVT LE PC VMCAGAM I H S R I GRWFG ARDAKTGAAGS LMDVLHHPGMNHRVE I TEG I LADE CAALL S D F FRMRRQE I KAQKKAQ S S T D
[0226] In one embodiment, a suitable adenosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following amino acid sequence:(SEQ ID NO: 22178)MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWD EREVPVGAVLVHNNRVI GEGWNRP I GRHDPTAHAE IMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAM IHSRIGRWFGSSTD
[0227] In one embodiment, a suitable adenosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following Staphylococcus aureus TadA amino acid sequence: (SEQ ID NO: 22179)MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIIT KDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVL GSWRLEGCTLYVTLEPCVMCAGT I VMSRI PRWYG ADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEA CSTLLTTFFKNLRANKKSTN
[0228] In one embodiment, a suitable adenosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following Bacillus subtilis TadA amino acid sequence:(SEQ ID NO: 22180)MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGE I IARAHNLRETEQRSIAHAEMLVIDEACKALGTWR LE GAT L YVT LE P C PMCAGAWL S RVE KWFGAFD P KGGCSGTLMNLLQEERFNHQAEWSGVLEEECGGM L S AF FRE LRKKKKAARKNL S E
[0229] In one embodiment, a suitable adenosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following Salmonella typhimurium TadA:(SEQ ID NO: 22181)MPPAFITGVTSLS DVE L DHE YWMRHAL T LAKRAWD EREVPVGAVLVHNHRVI GEGWNRP I GRHDPTAHAE IMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAM VHSRI GRWFGARDAKTGAAGSL I DVLHHPGMNHR VE I IEGVLRDECATLLSDFFRMRRQE IKALKKADR AEGAGPAV
[0230] In one embodiment, a suitable adenosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following Shewanella putrefaciens TadA amino acid sequence: (SEQ ID NO: 22182))MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQI ATGYNLS I SQHDPTAHAE I LCLRSAGKKLENYRLL DATLYITLEPCAMCAGAMVHSRIARWYGARDEKTGAAGTWNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE
[0231] In one embodiment, a suitable adenosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following Haemophilus influenzae F3031 TadA amino acid sequence:(SEQ ID NO: 22183)MDAAKVRSE FDEKMMRYALELADKAEALGE I PVGA VLVDDARNI IGEGWNLS IVQSDPTAHAE I IALRNG AKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKR LVFGASDYKTGAI GSRFHFFDDYKMNHTLE I TSGV LAEECSQKLS TFFQKRREEKKIEKALLKSLSDK
[0232] In one embodiment, a suitable adenosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following Caulobacter crescentus TadA amino acid sequence: (SEQ ID NO: 22184)MRT DE S E DQDHRMMRLALDAARAAAEAGE T PVGAV I LDPS TGEVI ATAGNGP I AAHDPTAHAE I AAMRAA AAKLGNYRLTDLTLWTLEPCAMCAGAISHARIGR WFGADDPKGGAWHGPKFFAQPTCHWRPEVTGGV LADESADLLRGFFRARRKAKI
[0233] In one embodiment, a suitable adenosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following Geobacter sulfurreducens TadA amino acid sequence:(SEQ ID NO: 22185)MS S LKKT P I RDDAYWMGKAI REAAKAAARDE VP I G AVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQA ARRSANWRLTGATLYVTLEPCLMCMGAI ILARLER WFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGV CQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP
[0234] Additional adenosine base editors can be found, for example, in U. S. Patent Nos.11,999,947, 11,795,443 and published U. S. Patent Application No. 2022 / 0307003, the contents of which are incorporated by reference herein in their entireties.Cytidine deaminases
[0235] Cytidine deaminases suitable for inclusion in a Tas polypeptide fusion polypeptide include any enzyme that is capable of deaminating cytidine in DNA.
[0236] Cytidine deaminases encompass enzymes in the cytidine deaminase superfamily, and in particular, enzymes of the APOBEC family (APOBEC1, APOBEC2, APOBEC4, and APOBEC3 subgroups of enzymes), activation-induced cytidine deaminase (AID or AICDA) and CMP deaminases (see, e.g., Conticello et al., Mol. Biol. Evol. 22:367-77, 2005; Conticello, Genome Biol. 9:229, 2008; Muramatsu et al., J. Biol. Chem. 274: 18470-6, 1999); and Carrington et al., Cells 9: 1690 (2020)). APOBEC is a family of evolutionarily conserved cytidine deaminases. Members of this family are C-to-U editing enzymes. The N-terminal domain of APOBEC like proteins is the catalytic domain, while the C-terminal domain is a pseudocatalytic domain. More specifically, the catalytic domain is a zinc dependent cytidine deaminase domain and is important for cytidine deamination.
[0237] In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of an APOBECI deaminase. In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of APOBEC2 deaminase. In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of is an APOBEC3 deaminase. In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of an APOBEC3A deaminase. In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of APOBEC3B deaminase. In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of APOBEC3C deaminase. In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of APOBEC3D deaminase. In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of APOBEC3E deaminase. In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of APOBEC3F deaminase. In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of APOBEC3G deaminase. In one embodiment, a deaminase incorporated into a fusion proteincomprises all or a portion of APOBEC3H deaminase. In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of APOBEC4 deaminase. In one embodiment, a deaminase incorporated into a fusion protein comprises all or a portion of activation-induced deaminase (AID). In some embodiments a deaminase incorporated into a fusion protein comprises all or a portion of cytidine deaminase 1 (CDA1). It should be appreciated that a fusion protein can comprise a deaminase from any suitable organism (e.g., a human or a rat). In one embodiment, a deaminase domain of a fusion protein is from a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In one embodiment, the deaminase domain of the fusion protein is derived from rat (e.g., rat APOBEC1). In one embodiment, the deaminase domain is human APOBEC1. In one embodiment, the deaminase domain is pmCDA1.
[0238] Sequences of exemplary cytidine deaminases are provided below.
[0239] In one embodiment, a suitable cytidine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following amino acid sequence:Human AID:(SEQ ID NO: 22186)MDSLLMNRRKFLYQFKNVRWAKGRRETYLC YWKRRDSAT S FS LD FG YLR NKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRG NPNLSLRI FTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKAPVUnderline: nuclear localization sequence.
[0240] In one embodiment, a suitable cytidine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following amino acid sequence:(SEQ ID NO: 22187)clAID (Canis lupus familiaris):MDSLLMKQRKFLYHFKNVRWAKGRHETYLC YWKRRDSAT S FS LD FGHLR NKSGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRG YPNLSLRI FAARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNT FVENREKTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGLUnderline: nuclear localization sequence.
[0241] In one embodiment, a suitable cytidine deaminase is an AID and comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%, amino acid sequence identity to the following amino acid sequence:(SEQ ID NO: 22188 )rAPOBEC-1 (Rattus norvegicus):MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSI WRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAI TEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESG YCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCI ILGLPPCLNILRRKQ PQLTFFTIALQSCHYQRLPPHILWATGLKUracil Glycosylase Inhibitor (UGI)
[0242] In one embodiment, a nucleic acid is provided, the nucleic acid comprising an open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., A3 A) and an RNA-guided Tas nickase, wherein the polypeptide does not comprise a uracil glycosylase inhibitor (UGI). In one embodiment, the nucleic acid is DNA or RNA. In one embodiment, the nucleic acid is mRNA. In one embodiment, a polypeptide encoded by the mRNA is provided.
[0243] In one embodiment, a polypeptide or an mRNA encoding the polypeptide, are provided, the polypeptide comprising a cytidine deaminase and an RNA-guided Tas nickase, wherein the polypeptide does not comprise a UGI. In one embodiment, the cytidine deaminase is A3 A. In one embodiment, the RNA-guided Tas nickase does not comprise a uracil glycosylase inhibitor (UGI).
[0244] In one embodiment, a composition is provided comprising a first polypeptide, or an mRNA encoding a first polypeptide, comprising a cytidine deaminase and an RNA-guided Tas nickase; and a second polypeptide, or an mRNA encoding a second polypeptide, comprising a uracil glycosylase inhibitor (UGI), wherein the second polypeptide is different from the first polypeptide.
[0245] In one embodiment, a composition is provided comprising a first nucleic acid comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase and an RNA-guided nickase, and a second nucleic acid comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), wherein the second nucleic acid is different from the first nucleic acid. In one embodiment, the first nucleic acid encodes a polypeptide that does not comprise a UGI.
[0246] In one embodiment, methods of modifying a target gene are provided comprising administering the compositions described herein. In one embodiment, the method comprises delivering to a cell a first nucleic acid comprising a first open reading frame encoding a first polypeptide comprising a cytidine deaminase and an RNA-guided Tas nickase, and a second nucleic acid comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), wherein the second nucleic acid is different from the first nucleic acid.
[0247] In one embodiment, the methods comprise delivering to a cell a polypeptide comprising a cytidine deaminase and an RNA-guided Tas nickase, or a nucleic acid encoding the polypeptide, and separately (e.g., not via the same nucleic acid construct) delivering to the cell a uracil glycosylase inhibitor (UGI), or a nucleic acid encoding the UGI.
[0248] Without being bound by any theory, providing a UGI together with a polypeptide comprising a deaminase may be helpful in the methods described herein by inhibiting cellular DNA repair machinery (e.g., UDG and downstream repair effectors) that recognize a uracil in DNA as a form of DNA damage or otherwise would excise or modify the uracil and / or surrounding nucleotides. It should be understood that the use of a UGI may increase the editing efficiency of an enzyme that is capable of deaminating C residues.
[0249] Suitable UGI protein and nucleotide sequences are known to those in the art, and include, for example, those published in Wang et al., Uracil-DNA glycosylase inhibitor gene of bacteriophage PBS2 encodes a binding protein specific for uracil-DNA glycosylase. J. Biol. Chem.264: 1163-1171(1989); Lundquist et al., Site-directed mutagenesis and characterization of uracil-DNA glycosylase inhibitor protein. Role of specific carboxylic amino acids in complex formation with Escherichia coli uracil-DNA glycosylase. J. Biol. Chem. 272:21408-21419(1997); Ravishankar et al., X-ray analysis of a complex of Escherichia coli uracil DNA glycosylase (EcUDG) with a proteinaceous inhibitor. The structure elucidation of a prokaryotic UDG. Nucleic Acids Res. 26:4880-4887(1998); and Putnam et al., Protein mimicry of DNA from crystal structures of the uracil-DNA glycosylase inhibitor protein and its complex with Escherichia coli uracil-DNA glycosylase. J. Mol. Biol. 287:331-346(1999), the entire contents of each are incorporated herein by reference. It should be appreciated that any proteins that are capable of inhibiting a uracil-DNA glycosylase base-excision repair enzyme are within the scope of the present disclosure. Additionally, any proteins that block or inhibit base-excision repair are also within the scope of this disclosure. In one embodiment, a uracil glycosylase inhibitor is a proteinthat binds uracil. In one embodiment, a uracil glycosylase inhibitor is a protein that binds uracil in DNA. In one embodiment, a uracil glycosylase inhibitor is a single-stranded binding protein. In one embodiment, a uracil glycosylase inhibitor is a catalytically inactive uracil DNA-glycosylase protein. In one embodiment, a uracil glycosylase inhibitor is a catalytically inactive uracil DNA-glycosylase protein that does not excise uracil from the DNA. In one embodiment, a uracil glycosylase inhibitor is a catalytically inactive UDG.Glycosylase-Based Base Editors
[0250] In one embodiment, a Tas fusion partner can be a base excising domain, for example, a glycosylase. In one embodiment, the glycosylase can be N-methylpurine DNA glycosylase (MPG), 8-oxoguanine DNA glycosylase (OGGI), methyl-CpG binding domain 4, DNA glycosylase (MBD4), thymine DNA glycosylase (TDG), uracil DNA glycosylase (UNG), singlestrand-selective monofunctional uracil-DNA glycosylase 1 (SMUG1), mutY DNA glycosylase (MUTYH), nth like DNA glycosylase 1 (NTHL1), nei like DNA glycosylase 1 (NEIL1), nei like DNA glycosylase 2 (NEIL2), nei like DNA glycosylase 3 (NEIL3), or mutants thereof capable of recognizing and excising a base from a nucleotide of a nucleic acid.
[0251] In one embodiment, a Tas fusion polypeptide is provided in which a Tas polypeptide is fused with an adenine base editor (ABE) with hypoxanthine excision protein N-methylpurine DNA glycosylase (MPG). This adenine transversion base editor, AYBE, allows for A-to-C and A-to-T transversion editing (see Tong, H. et al. Programmable A-to-Y base editing by fusing an adenine base editor with an N-methylpurine DNA glycosylase. Nat. Biotechnol. 41, 1080-1084 (2023)).
[0252] Other examples of DNA glycosylases are described in the International Patent Publication No. WO2024222812, the content of which is incorporated by reference herein in its entirety.Transcription Factors
[0253] In one embodiment, a Tas fusion polypeptide of the present disclosure comprises: i) a Tas polypeptide of the present disclosure; and ii) a heterologous polypeptide (a “fusion partner”), where the heterologous polypeptide is a transcription factor. A transcription factor can include: i) a DNA binding domain; and ii) a transcription activator. A transcription factor can include: i) a DNA binding domain; and ii) a transcription repressor. Suitable transcription factors include polypeptides that include a transcription activator or a transcription repressor domain (e g., theKruppel associated box (KRAB or SKD); the Mad mSIN3 interaction domain (SID); the ERF repressor domain (ERD), etc.); zinc-finger-based artificial transcription factors (see, e.g., Sera (2009) Adv. Drug Deliv. 61:513); TALE-based artificial transcription factors (see, e.g., Liu et al. (2013) Nat. Rev. Genetics 14:781); and the like. In one embodiment, the transcription factor comprises a VP64 polypeptide (transcriptional activation). In one embodiment, the transcription factor comprises a Kriippel-associated box (KRAB) polypeptide (transcriptional repression). In one embodiment, the transcription factor comprises a Mad mSIN3 interaction domain (SID) polypeptide (transcriptional repression). In one embodiment, the transcription factor comprises an ERF repressor domain (ERD) polypeptide (transcriptional repression). For example, in one embodiment, the transcription factor is a transcriptional activator, where the transcriptional activator is GAL4-VP16.Epigenetic Modifiers
[0254] In one embodiments, the engineered TIGR systems described herein may further comprise an epigenetic modification domain such that binding of the engineered TIGR at target sequence on genomic DNA (e.g., chromatin) results in one or more epigenetic modifications by the epigenetic modification domain that increases or decreases expression of the one or more polypeptides. As used herein, “linked to or otherwise capable of associating with” refers to a fusion protein or a recruitment domain or an adaptor protein, such as an aptamer (e.g., MS2) or an epitope tag. The recruitment domain or an adaptor protein can be linked to an epigenetic modification domain or the DNA binding domain (e.g., an adaptor for an aptamer). The epigenetic modification domain can be linked to an antibody specific for an epitope tag fused to the engineered TIGR. An aptamer can be linked to a guide sequence.
[0255] In example embodiments, the DNA binding domain is a programmable DNA binding protein linked to or otherwise capable of associating with an epigenetic modification domain. In example embodiments, the DNA binding domain is a nuclease-deficient RNA-guided DNA endonuclease enzyme or a nuclease-deficient endonuclease enzyme. In example embodiments, a CRISPR system having an inactivated nuclease activity (e.g., dCas) is used as the DNA binding domain.
[0256] In example embodiments, the epigenetic modification domain is a functional domain and includes, but is not limited to a histone methyltransferase (HMT) domain, histone demethylase domain, histone acetyltransferase (HAT) domain, histone deacetylation (HDAC) domain, DNAmethyltransferase domain, DNA demethylation domain, histone phosphorylation domain (e.g, serine and threonine, or tyrosine), histone ubiquitylation domain, histone sumoylation domain, histone ADP ribosylation domain, histone proline isomerization domain, histone biotinylation domain, histone citrullination domain (see, e.g., Epigenetics, Second Edition, 2015, Edited by C. David Allis; Marie-Laure Caparros; Thomas Jenuwein; Danny Reinberg; Associate Editor Monika Lachlan; Dawson MA, Kouzarides T. Cancer epigenetics: from mechanism to therapy. Cell.2012;150(l):12-27; Syding LA, Nickl P, Kasparek P, Sedlacek R. CRISPR / Cas9 Epigenome Editing Potential for Rare Imprinting Diseases: A Review. Cells. 2020;9(4):993; and Zhang Y. Transcriptional regulation by histone ubiquitination and deubiquitination. Genes Dev.2003;17(22):2733-2740). Example epigenetic modification domains can be obtained from, but are not limited to chromatin modifying enzymes, such as, DNA methyltransferases (e.g., DNMT1, DNMT3a and DNMT3b), TET1, TET2, thymine-DNA glycosylase (TDG), GCN5-related N-acetyltransferases family (GNAT), MYST family proteins (e.g., MOZ and MORE), and CBP / p300 family proteins (e.g., CBP, p300), Class I HDACs (e.g., HD AC 1-3 and HDAC8), Class II HDACs (e.g., HDAC 4-7 and HDAC 9-10), Class III HDACs (e.g., sirtuins), HDAC11, SET domain containing methyltransferases (e.g., SET7 / 9 (KMT7, NCBI Entrez Gene: 80854), KMT5A (SET8), MMSET, EZH2, and MLL family members), DOT1L, LSD1, Jumonji demethylases (e.g., KDM5A (JARID1A), KDM5C (JARID1C), and KDM6A (UTX)), kinases (e.g. Haspin, VRK1, PKCa, PKCP, PIM1, IKKa, Rsk2, PKB / Akt, Aurora B, MSK1 / 2, JNK1, MLTKa, PRK1, Chkl, Dlk / ZIP, PKC8, MST1, AMPK, JAK2, Abl, BMK1, CaMK, S6K1, SIK1), Ubp8, ubiquitin C-terminal hydrolases (UCH), the ubiquitin-specific processing proteases (UBP), and poly(ADP-ribose) polymerase 1 (P ARP-1). See, also, US Patent US 11001829B2 for additional domains. In example embodiments, the epigenetic modification domain is a catalytically active TIGR polypeptide described herein.
[0257] In example embodiments, histone acetylation is targeted to a target sequence using a engineered TIGR polypeptide (see, e.g., Hilton IB, et al. Epigenome editing by a CRISPR-Cas9-based acetyltransferase activates genes from promoters and enhancers. Nat Biotechnol. 2015). In example embodiments, histone deacetylation is targeted to a target sequence (see, e.g, Cong et al, 2012; and Konermann S, et al. Optical control of mammalian endogenous transcription and epigenetic states. Nature. 2013;500:472-476). In example embodiments, histone methylation is targeted to a target sequence (see, e.g, Snowden AW, Gregory PD, Case CC, Pabo CO. Gene-specific targeting of H3K9 methylation is sufficient for initiating repression in vivo. Curr Biol.2002;12:2159-2166; and Cano-Rodriguez D, Gjaltema RA, Jilderda LJ, et al. Writing of H3K4Me3 overcomes epigenetic silencing in a sustained but context-dependent manner. Nat Commun. 2016;7: 12284). In example embodiments, histone demethylation is targeted to a target sequence (see, e.g., Kearns NA, Pham H, TabakB, et al. Functional annotation of native enhancers with a Cas9-histone demethylase fusion. Nat Methods. 2015;12(5):401-403). In example embodiments, histone phosphorylation is targeted to a target sequence (see, e.g., Li J, Mahata B, Escobar M, et al. Programmable human histone phosphorylation and gene activation using a CRISPR / Cas9-based chromatin kinase. Nat Commun. 2021; 12(1):896). In example embodiments, DNA methylation is targeted to a target sequence (see, e.g., Rivenbark AG, et al. Epigenetic reprogramming of cancer cells via targeted DNA methylation. Epigenetics. 2012;7:350-360; Siddique AN, et al. Targeted methylation and gene silencing of VEGF-A in human cells by using a designed Dnmt3a-Dnmt3L single-chain fusion protein with increased DNA methylation activity. J Mol Biol. 2013;425:479-491; Bernstein DL, Le Lay JE, Ruano EG, Kaestner KH. TALE-mediated epigenetic suppression of CDKN2A increases replication in human fibroblasts. I Clin Invest. 2015;125:1998-2006; Liu XS, Wu H, li X, et al. Editing DNA Methylation in the Mammalian Genome. Cell. 2016;167(l):233-247.el7; StepperP, Kungulovski G, JurkowskaRZ, et al. Efficient targeted DNA methylation with chimeric dCas9-Dnmt3a-Dnmt3L methyltransferase. Nucleic Acids Res. 2017;45(4): 1703-1713; and Pflueger C., Tan D., Swain T., Nguyen T., Pflueger I., Nefzger C., Polo I. M., Ford E., Lister R. A modular dCas9-SunTag DNMT3A epigenome editing system overcomes pervasive off-target activity of direct fusion dCas9-DNMT3A constructs. Genome Res. 2018;28:l 193-1206). In example embodiments, DNA demethylation is targeted to a target sequence using a CRISPR system (see, e.g., TET1, see Xu et al, Cell Discov. 2016 May 3;2: 16009; Choudhury et al, Oncotarget. 2016 Jul 19;7(29):46545-46556; and Kang JG, Park JS, Ko JH, Kim YS. Regulation of gene expression by altered promoter methylation using a CRISPR / Cas9-mediated epigenetic editing system. Sci Rep.2019;9(l): 11960). In example embodiments, DNA demethylation is targeted to a target sequence (see, e.g., TDG, see, Gregory DJ, Zhang Y, Kobzik L, Fedulov AV. Specific transcriptional enhancement of inducible nitric oxide synthase by targeted promoter demethylation. Epigenetics.2013;8:1205-1212).
[0258] Example epigenetic modification domains can be obtained from, but are not limited to transcription activators, such as, VP64 (see, e.g., Ji Q, et al. Engineered zinc-finger transcription factors activate OCT4 (POU5F1), SOX2, KLF4, c-MYC (MYC) and miR302 / 367. Nucleic Acids Res. 2014;42:6158-6167; Perez-Pinera P, et al. Synergistic and tunable human gene activation by combinations of synthetic transcription factors. Nat Methods. 2013;10:239-242; Farzadfard F, Perli SD, Lu TK. Tunable and multifunctional eukaryotic transcription factors based on CRISPR / Cas. ACS Synth Biol. 2013;2:604-613; Black JB, Adler AF, Wang HG, et al. Targeted Epigenetic Remodeling of Endogenous Loci by CRISPR / Cas9-Based Transcriptional Activators Directly Converts Fibroblasts to Neuronal Cells. Cell Stem Cell. 2016;19(3):406-414; and Maeder ML, Linder SJ, Cascio VM, Fu Y, Ho QH, Joung JK. CRISPR RNA-guided activation of endogenous human genes. Nat Methods. 2013;10(10):977-979), p65 (see, e.g., Liu PQ, et al. Regulation of an endogenous locus using a panel of designed zinc finger proteins targeted to accessible chromatin regions. Activation of vascular endothelial growth factor A. J Biol Chem.2001;276:11323-11334; and Konermann S, et al. Genome-scale transcriptional activation by an engineered CRISPR-Cas9 complex. Nature. 2015;517:583-588), HSF1, and RTA (see, e.g., Chavez A, et al. Highly efficient Cas9-mediated transcriptional programming. Nat Methods.2015;12:326-328). Example epigenetic modification domains can be obtained from, but are not limited to transcription repressors, such as, KRAB (see, e.g., Beerli RR, Segal DJ, Dreier B, Barbas CF., 3rd Toward controlling gene expression at will: specific regulation of the erbB-2 / HER-2 promoter by using polydactyl zinc finger proteins constructed from modular building blocks. Proc Natl Acad Sci U S A. 1998;95:14628-14633; Cong L, Zhou R, Kuo YC, Cunniff M, Zhang F. Comprehensive interrogation of natural TALE DNA-binding modules and transcriptional repressor domains. Nat Commun. 2012;3:968; GilbertLA, et al. CRISPR-mediated modular RNA-guided regulation of transcription in eukaryotes. Cell. 2013;154:442-451; and Yeo NC, Chavez A, Lance-Byrne A, et al. An enhanced CRISPR repressor for targeted mammalian gene regulation. Nat Methods. 2018; 15(8):611-616).
[0259] In example embodiments, the epigenetic modification domain linked to a DNA binding domain recruits an epigenetic modification protein to a target sequence. In example embodiments, a transcriptional activator recruits an epigenetic modification protein to a target sequence. For example, VP64 can recruit DNA demethylation, increased H3K27ac and H3K4me. In example embodiments, a transcriptional repressor protein recruits an epigenetic modification protein to atarget sequence. For example, KRAB can recruit increased H3K9me3 (see, e.g., Thakore PI, D'Ippolito AM, Song L, et al. Highly specific epigenome editing by CRISPR-Cas9 repressors for silencing of distal regulatory elements. Nat Methods. 2015; 12(12): 1143-1149). In an example embodiment, methyl-binding proteins linked to a DNA binding domain, such as MBD1, MBD2, MBD3, and MeCP2 recruits an epigenetic modification protein to a target sequence. In an example embodiment, Mi2 / NuRD, Sin3A, or Co-REST recruit HDACs to a target sequence.
[0260] In example embodiments, the epigenetic modification domain can be a eukaryotic or prokaryotic (e.g., bacteria or Archaea) protein. In example embodiments, the eukaryotic protein can be a mammalian, insect, plant, or yeast protein and is not limited to human proteins (e.g., a yeast, insect, plant chromatin modifying protein, such as yeast HATs, HDACs, methyltransferases, etc.
[0261] In one aspect of the invention, is provided a fusion protein (epigenetic modification polypeptide) comprising from N-terminus to C-terminus, an epigenetic modification domain, an XTEN linker, and a nuclease-deficient RNA-guided DNA endonuclease enzyme or a nuclease-deficient endonuclease enzyme.
[0262] In aspects, the epigenetic modification polypeptide further comprises a transcriptional activator. In aspects, the transcriptional activator is VP64, p65, RTA, or a combination of two or more thereof. In another aspect, the epigenetic modification polypeptide further comprises one or more nuclear localization sequences. In embodiments, the epigenetic modification polypeptide comprises the nuclease-deficient RNA-guided DNA endonuclease enzyme. In embodiments, the fusion protein comprises the nuclease-deficient DNA endonuclease enzyme.
[0263] In some embodiments, the functional domains associated with the adaptor protein or the CRISPR enzyme is a transcriptional activation domain comprising VP64, p65, MyoDl, HSF1, RTA or SET7 / 9. Other references herein to activation (or activator) domains in respect of those associated with the adaptor protein(s) include any known transcriptional activation domain and specifically VP64, p65, MyoDl, HSF1, RTA or SET7 / 9 (see, e.g., US Patent, US11001829B2).
[0264] In certain embodiments, the present invention provides a fusion protein comprising from N-terminus to C-terminus, an RNA-binding sequence, an XTEN linker, and a transcriptional activator. In aspects, the transcriptional activator is VP64, p65, RTA, or a combination of two or more thereof. In aspects, the fusion protein further comprises a demethylation domain, a nuclease-deficient RNA-guided DNA endonuclease enzyme or a nuclease-deficient endonuclease enzyme,a nuclear localization sequence, or a combination of two or more thereof. In embodiments, the fusion protein comprises the nuclease-deficient RNA-guided DNA endonuclease enzyme. In embodiments, the fusion protein comprises the nuclease-deficient DNA endonuclease enzyme.
[0265] Example epigenome modification systems that can be adapted for use with the engineered TIGR systems disclosed herein include US 2020 / 0003761; WO 2018 / 053035 (“Targeted DNA Demethylation and Methylation”, Jackson Laboratory); WO 2018 / 148667 (“Reprogramming Cell Aging”, Memorial Sloan Kettering Cancer Center); WO 2022 / 140577 (“Compositions and Methods for Epigenetic Editing”, Chroma Medicine); WO 2017 / 090724 (“DNA Methylation Editing Kit and DNA Methylation Editing Method”, Gunma University NUC); WO 2019 / 0136229 (“Compositions and Methods of Improving Specificity in Genomic Engineering Using RNA-guided Nucleases”, Duke University); WO 2014 / 059255 (Transcription Activator-like Effector (TALE)- Lysine-specific Demthylase 1 (LSD 1) Fusion Proteins”, General Hospital Corp.); WO 2014 / 152432 (“Increasing Specificity for RNA-guided Genome Editing”, General Hospital Corp.).Tas-associatedMethyltransferases
[0266] The TIGR methyltransferase systems described herein are inspired by, and in certain embodiments derived from, naturally occurring TIGR system loci in which a Tas polypeptide gene is found in genomic proximity to a gene encoding a putative methyltransferase or methylase. Structural mining of TasA-containing TIGR loci has revealed candidate genomic architectures in which Tas genes are consistently proximal to putative methyltransferases and methylases, as illustrated in FIGS. 28A and 28B, and RNA-seq data mapping to these loci has confirmed that tigRNA-encoding arrays are actively expressed at these sites (FIG. 28C). Critically, the locusspecific methyltransferase physically associates with the TasA polypeptide, as demonstrated by copurification experiments (FIG. 29C), establishing that the TasA-tigRNA ribonucleoprotein complex and the cognate methyltransferase form a naturally occurring functional unit. This biochemical association provides strong evidence that, in certain naturally occurring TIGR systems, the Tas polypeptide does not function solely as an RNA-guided nuclease or transposase scaffold, but rather cooperates with a physically associated methyltransferase to direct site-specific DNA methylation at target loci identified by the tigRNA guide. The naturally occurring coassociation of TasA and a locus-specific methyltransferase thus served as the biological prototype for the engineered TIGR methyltransferase systems of the present disclosure.
[0267] In one embodiment, the present disclosure provides an engineered or non-naturally occurring TIGR methyltransferase system. As used herein, a " TIGR methyltransferase system" refers to a multicomponent, RNA-guided molecular system that achieves sequence-specific DNA methylation at a programmable target locus. The system comprises three principal components: a Tandem Interspaced Guide RNA (TIGR)-associated (Tas) polypeptide; a methyltransferase (MTase) capable of associating with the Tas polypeptide; and a tandem-interspersed guide molecule (TIGR guide, or tigRNA) capable of forming a complex with the Tas polypeptide and directing sequence-specific base pairing with both the sense and antisense strands of a doublestranded target oligonucleotide sequence. The dual-strand targeting feature of the TIGR guide is a defining characteristic of the system and distinguishes it from single-strand-targeting programmable methylation approaches; by engaging both strands of the target DNA duplex, the tigRNA confers substantially enhanced specificity at the target locus. The Tas polypeptide, as described in greater detail elsewhere herein, serves as the RNA-binding scaffold that simultaneously engages the tigRNA and recruits the associated MTase to the target site, thereby enabling methylation to be precisely delivered to any desired genomic or extrachromosomal locus simply by programming the spacer sequences of the tigRNA.
[0268] The physical association between the Tas polypeptide and the MTase may take several forms. In one embodiment, the Tas polypeptide and the MTase are covalently linked, either as a direct genetic fusion or through a peptide linker of defined length and composition. The linker polypeptide is between 1 and 100 or more amino acids in length. In practice, linkers of different lengths and flexibilities can be employed to optimize the geometry and activity of the fusion protein; flexible linkers composed predominantly of glycine and serine residues are suitable in many embodiments due to their conformational freedom, while rigid or semi-rigid linkers may be employed where maintaining a specific spatial relationship between the Tas polypeptide and the MTase is desired. Those of ordinary skill in the art will appreciate that the optimal linker length and composition for a given application may be determined empirically, and that linkers of 1-5, 5-15, 15-30, 30-50, or 50-100 amino acids are all within the scope of the present disclosure. In one embodiment, the MTase is covalently bound to the N-terminus of the Tas polypeptide. In another embodiment, the MTase is covalently bound to the C-terminus of the Tas polypeptide. The choice between N-terminal and C-terminal fusion may influence activity, and both orientations are expressly contemplated herein. In certain embodiments, the MTase is linked to an internal regionof the Tas polypeptide, for example via insertion at a surface-exposed loop or domain boundary. In still other embodiments, the Tas polypeptide and the MTase associate non-covalently, for example through protein-protein interaction domains, thereby preserving the natural modularity of the system while still enabling coordinated delivery of the MTase to the target site.
[0269] The MTase component of the TIGR methyltransferase system may be any methyltransferase capable of modifying a DNA base upon delivery to the target site. In one embodiment, the MTase is a DNA methyltransferase. As used herein, " DNA methyltransferase" refers to an enzyme that transfers a methyl group from S-adenosyl-L-methionine (SAM) to a base within a DNA sequence, resulting in a methylated base at the target position. DNA methyltransferases broadly include adenosine methyltransferases, which methylate the N6 position of adenine (producing N6-methyladenine, or 6mA), and cytosine methyltransferases, which methylate the C5 position of cytosine (producing 5 -methyl cytosine, or 5mC) or the N4 position of cytosine (producing N4-methylcytosine, or 4mC). In one embodiment, the DNA methyltransferase is an adenosine methyltransferase. In one embodiment, the DNA methyltransferase is a cytosine methyltransferase. Cytosine methylation at CpG dinucleotides is the dominant epigenetic DNA modification in mammalian cells and plays central roles in gene silencing, genomic imprinting, and X-chromosome inactivation, making targeted cytosine methyltransferases particularly relevant for therapeutic epigenome editing applications. In one embodiment, the cytosine methyltransferase is selected from Dnmtl (DNA (cytosine-5-)-methyltransferase 1), Dnmt3a (DNA (cytosine-5-)-methyltransferase 3 alpha), and Dnmt3b (DNA (cytosine-5-)-methyltransferase 3 beta), or a catalytic domain or active variant thereof. Dnmtl functions primarily as a maintenance methyltransferase that copies methylation patterns from hemimethylated DNA following replication, while Dnmt3a and Dnmt3b function primarily as de novo methyltransferases capable of establishing new methylation marks on previously unmethylated substrates; both activities are useful depending on the intended application. In one embodiment, the MTase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence selected from SEQ ID NOs corresponding to positions 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, and 104, each of which represents a naturally occurring or previously characterized methyltransferase sequence suitable for use in the TIGR methyltransferase system.Tas-Associated Methyltransferases (Tas-TMases)
[0270] As used herein, the term " Tas-associated methyltransferase" or " Tas-TMase" refers to a naturally occurring or engineered polypeptide or protein complex comprising a Tas polypeptide that is physically associated with a methyltransferase activity, reflecting the naturally observed coassociation of TasA proteins with cognate locus-specific methyltransferases in certain TIGR system loci. The Tas-TMase concept encompasses both naturally derived complexes, in which a Tas polypeptide and a methyltransferase are encoded in the same or adjacent genomic locus and are co-expressed and co-assembled in their native host, and engineered fusion or non-covalent complexes in which a Tas polypeptide and a heterologous or cognate MTase are combined to generate a functional targeting and methylation unit. The Tas-TMase is distinct from other Tas fusion polypeptides described herein in that it specifically pairs the RNA-guided targeting capability of the Tas polypeptide with enzymatic DNA methylation activity, enabling site-specific epigenetic modification of a target locus in a programmable manner.
[0271] In one embodiment, the Tas-TMase comprises a Tas polypeptide having an amino acid sequence with at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of any one of SEQ ID NOs: 23997, 23998, 23999, 24000, 24001, 24002, 24003, 24004, 24005, 24006, 24007, 24008, 24009, 24010, and 24011. These sequences represent naturally occurring Tas polypeptides identified from TIGR loci that are genomically associated with methyltransferase genes, and they therefore constitute well-validated Tas polypeptide components for inclusion in a Tas-TMase. The sequence identity ranges recited herein encompass individual identity percentages of 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, and 100%, provided the resulting Tas polypeptide retains the ability to bind a tigRNA, be directed to a target sequence, and support methylation of the target by the associated MTase component. In one embodiment, the Tas-TMase further comprises a DNA methyltransferase. In one embodiment, the DNA methyltransferase of the Tas-TMase is an adenosine methyltransferase. In one embodiment, the DNA methyltransferase is a cytosine methyltransferase, including but not limited to Dnmtl, Dnmt3a, and Dnmt3b, or catalytic domains thereof. Both the naturally co-encoded methyltransferases associated with TasA-containing TIGR loci and heterologous mammalian or prokaryotic DNA methyltransferases are contemplated as MTase components of the Tas-TMase.
[0272] The Tas polypeptide and the MTase within a Tas-TMase may be associated covalently or non-covalently. In one embodiment, the Tas polypeptide is non-covalently bound to the MTase,reflecting the naturally occurring mode of association observed between TasA and its cognate locus-specific methyltransferase. In this non-covalent configuration, the Tas polypeptide and the MTase are distinct polypeptide chains that are brought together through protein-protein interactions, optionally mediated by the tigRNA or by interaction surfaces on the Tas polypeptide itself. In one embodiment, the Tas polypeptide is covalently bound to the MTase, as in a direct genetic fusion or a chemically cross-linked complex, as described above in the context of the engineered TIGR methyltransferase system of claim 99. In certain embodiments, the covalent Tas-TMase is expressed as a single polypeptide chain from a single open reading frame, simplifying delivery and ensuring consistent stoichiometry of the Tas and MTase components.Methods of Site-Specific DNA Methylation Using the TIGR Methyltransferase System
[0273] In one embodiment, the present disclosure provides a method for methylating a base of a double-stranded DNA target nucleotide sequence at a locus of interest. The method comprises delivering to the locus of interest an engineered composition comprising: (a) a Tandem Interspaced Guide RNA (TIGR)-associated (Tas) polypeptide, or a nucleotide sequence encoding the Tas polypeptide; (b) a methyltransferase bound to the Tas polypeptide; and (c) a split spacer guide RNA (also referred to herein as a TIGR guide or tigRNA), or a nucleotide sequence encoding the guide molecule, wherein the split spacer guide RNA comprises a programmable nucleotide sequence capable of base pairing with both the sense and antisense strands of the double-stranded DNA target nucleotide sequence, and wherein formation of a complex between the Tas polypeptide, the methyltransferase, and the split spacer guide RNA facilitates binding of the complex to the target nucleotide sequence and triggers methylation of a base of the target nucleotide sequence at the locus of interest. The term "delivering to said locus" encompasses all modalities of introducing the composition to the biological context in which the target nucleotide sequence resides, including but not limited to transfection, transduction, lipid nanoparticle-mediated delivery, viral vector delivery, ribonucleoprotein (RNP) electroporation, and direct injection, as described in greater detail in the delivery vehicle section herein.
[0274] In an embodiment, the tigRNA, once expressed or introduced into the relevant cellular or in vitro context, associates with the Tas polypeptide through interactions between the tigRNA repeat sequences and the forming a Tas-tigRNA ribonucleoprotein complex. In another embodiment, the The programmable spacer sequences of the tigRNA then direct the complex to the target locus by base pairing with both strands of the target DNA duplex in the manner describedin detail elsewhere herein, forming an R-loop-like structure in which both the sense and antisense target strands are engaged. The MTase component, which is pre-associated with the Tas polypeptide either covalently or non-covalently, is thereby delivered in proximity to the target base, and catalyzes transfer of a methyl group from SAM to the target base. The result is sitespecific methylation at the intended locus of interest, without requiring the introduction of a DNA double-strand break. The programmability of the system — determined entirely by the spacer sequences within the tigRNA — allows the method to be applied to any target nucleotide sequence accessible to the TIGR guide, enabling flexible and modular epigenome editing across diverse genomic contexts and cell types.Recombinases
[0275] In one embodiment, a Tas fusion polypeptide of the present disclosure comprises: i) a Tas polypeptide of the present disclosure; and ii) a heterologous polypeptide (a “fusion partner”), where the heterologous polypeptide is a recombinase. Suitable recombinases include, e.g., a Cre recombinase; a Hin recombinase; a Tre recombinase; a FLP recombinase; tyrosine recombinase, and the like.
[0276] Examples of various additional suitable heterologous polypeptide (or fragments thereof) for a subject fusion Tas polypeptide include, but are not limited to, those described in the following applications (which publications are related to other CRISPR endonucleases such as Cas9, but the described fusion partners can also be used with a Tas protein instead): PCT patent applications: W02010075303, WO2012068627, and WO2013155555, and can be found, for example, in U. S. patents and patent applications: U. S. Pat. Nos. 8,906,616; 8,895,308; 8,889,418; 8,889,356; 8,871,445; 8,865,406; 8,795,965; 8,771,945; 8,697,359; 20140068797; 20140170753; 20140179006; 20140179770; 20140186843; 20140186919; 20140186958; 20140189896; 20140227787; 20140234972; 20140242664; 20140242699; 20140242700; 20140242702; 20140248702; 20140256046; 20140273037; 20140273226; 20140273230; 20140273231; 20140273232; 20140273233; 20140273234; 20140273235; 20140287938; 20140295556; 20140295557; 20140298547; 20140304853; 20140309487; 20140310828; 20140310830; 20140315985; 20140335063; 20140335620; 20140342456; 20140342457; 20140342458; 20140349400; 20140349405; 20140356867; 20140356956; 20140356958; 20140356959; 20140357523; 20140357530; 20140364333; and 20140377868; all of which are hereby incorporated by reference in their entirety.Prime Editors
[0277] In one embodiment, the present disclosure provides compositions and systems may comprise a engineered TIGR system or a catalytically inactive form, one or more TIGR guide or guide molecules, and a reverse transcriptase. The systems may be used to insert a donor polynucleotide to a target polynucleotide. In some examples, the composition or system comprises a catalytically inactive engineered TIGR system, a reverse transcriptase associated with or otherwise capable of forming a complex with the engineered TIGR systems, and a TIGR guide or guide molecule capable of forming a complex with the engineered TIGR systems and directing site-specific binding of the complex to a target sequence of a target polynucleotide, the TIGR guide or guide molecule further comprising a donor sequence for insertion into the target polynucleotide.
[0278] In some cases, the catalytically inactive engineered TIGR systems may be a nickase, e.g., a DNA nickase. In some cases, the engineered TIGR system has one or more mutations. In some examples, the engineered TIGR system comprises mutations corresponding to the mutations in the RuvC or HNH nuclease. The engineered TIGR systems may be associated with a reverse transcriptase. As used in this context “associated” means covalently linked, e.g. via a linker, or otherwise capable of complexing with the reverse transcriptase. A reverse transcriptase domain may be a reverse transcriptase or a fragment thereof.
[0279] In some examples, the compositions and systems may comprise the engineered TIGR system disclosed herein; a reverse transcriptase (RT) polypeptide connected to or otherwise capable of forming a complex with the engineered TIGR system; and a TIGR guide or guide molecule capable of forming a complex with the engineered TIGR system and comprising: a TIGR guide or guide sequence capable of directing site-specific binding of the engineered TIGR system complex to a target sequence of a target polynucleotide; a 3’ binding site region capable of binding to a cleaved upstream strand of the target polynucleotide; and a RT template sequence encoding an extended sequence, wherein the extended sequence comprises a variant region and a 3’ homologous sequence capable of hybridization to the downstream cleaved strand of the target polynucleotide.
[0280] A reverse transcriptase domain may be a reverse transcriptase or a fragment thereof. A wide variety of reverse transcriptases (RT) may be used in alternative embodiments of the present invention, including prokaryotic and eukaryotic RT, provided that the RT functions within the host to generate a donor polynucleotide sequence from the RNA template. If desired, the nucleotidesequence of a native RT may be modified, for example using known codon optimization techniques, so that expression within the desired host is optimized. A reverse transcriptase (RT) is an enzyme used to generate complementary DNA (cDNA) from an RNA template, a process termed reverse transcription. Reverse transcriptases are used by retroviruses to replicate their genomes, by retrotransposon mobile genetic elements to proliferate within the host genome, by eukaryotic cells to extend the telomeres at the ends of their linear chromosomes, and by some nonretroviruses such as the hepatitis B virus, a member of the Hepadnaviridae, which are dsDNA-RT viruses. Retroviral RT has three sequential biochemical activities: RNA-dependent DNA polymerase activity, ribonuclease H, and DNA-dependent DNA polymerase activity. Collectively, these activities enable the enzyme to convert single-stranded RNA into double-stranded cDNA. In an embodiment, the RT domain of a reverse transcriptase is used in the present invention. The domain may include only the RNA-dependent DNA polymerase activity. In one embodiment, the RT domain is non-mutagenic, i.e., does not cause mutation in the donor polynucleotide (e.g., during the reverse transcriptase process). In example embodiments, the RT domain may be non-retron RT, e g., a viral RT or a human endogenous RTs. In some examples, the RT domain may be retron RT or DGRs RT. In some examples, the RT may be less mutagenic than a counterpart wildtype RT. In one embodiment, the RT herein is not mutagenic. In one embodiment, the reverse transcriptase is Human immunodeficiency virus (HIV) RT, Avian myoblastosis virus (AMV) RT, Moloney murine leukemia virus (M-MLV) RT a group II intron RT, a group II intron-like RT, or a chimeric RT. In an embodiment, the RT comprises modified forms of these RTs, such as, engineered variants of Avian myoblastosis virus (AMV) RT, Moloney murine leukemia virus (M-MLV) RT, or Human immunodeficiency virus (HIV) RT (see, e.g., Anzalone, et al., Search-and-replace genome editing without double-strand breaks or donor DNA, Nature. 2019 Dec;576(7785): 149-157).
[0281] The reverse transcriptase may be fused to the C-terminus of an engineered TIGR system. Alternatively or additionally, the reverse transcriptase may be fused to the N-terminus of an engineered TIGR system. The fusion may be via a linker and / or an adaptor protein. In some examples, the reverse transcriptase may be an M-MLV reverse transcriptase or variant thereof. The M-MLV reverse transcriptase variant may comprise one or more mutations. For the examples, the M-MLV reverse transcriptase may comprise D200N, L603W, and T33OP. In another example, the M-MLV reverse transcriptase may comprise D200N, L603W, T33OP, T306K, and W313F. Ina particular example, the fusion of engineered TIGR systems and reverse transcriptase is an engineered TIGR system (with a mutation corresponding to H840A of SpCas9) fused with M-MLV reverse transcriptase (D200N+L603W+T330P+T306K+W313F).
[0282] In one embodiment, the engineered TIGR systems herein may target DNA using a TIGR guide or guide RNA containing a binding sequence that hybridizes to the target sequence on the DNA. The TIGR guide or guide RNA may further comprise an editing sequence that contains new genetic information that replaces target DNA nucleotides. The small sizes of the engineered TIGR systems herein may allow easier packaging and delivery of the prime editing system, e.g., with a viral vector, e.g., AAV or lentiviral vector.
[0283] A single-strand break (a nick) may be generated on the target DNA by the engineered TIGR systems at the target site to expose a 3 ’-hydroxyl group, thus priming the reverse transcription of an edit-encoding extension on the TIGR guide or guide directly into the target site. These steps may result in a branched intermediate with two redundant single-stranded DNA flaps: a 5’ flap that contains the unedited DNA sequence, and a 3’ flap that contains the edited sequence copied from the ®RNA. The 5’ flaps may be removed by a structure-specific endonuclease, e.g., FEN122, which excises 5’ flaps generated during lagging-strand DNA synthesis and long-patch base excision repair. The non-edited DNA strand may be nicked to induce bias DNA repair to preferentially replace the non-edited strand. Examples of prime editing systems and methods include those described in Anzalone AV et al., Search-and-replace genome editing without doublestrand breaks or donor DNA, Nature. 2019 Oct 21. doi: 10.1038 / s41586-019-1711-4, which is incorporated by reference herein in its entirety.
[0284] The engineered TIGR system (e.g., the nickase form) may be used to prime-edit a single nucleotide on a target DNA. Alternatively or additionally, the engineered TIGR systems may be used to prime-edit at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 nucleotides on a target DNA.
[0285] Examples of prime editing systems and methods that may be adapted for use with the engineered TIGRs described herein include those described in Anzalone x., “Search-and-replacegenome editing without double-strand breaks or donor DNA”, Nature. 576, 149-157 (2019); Chen et al. “Enhanced prime editing systems by manipulating cellular determinants of editing outcomes” Cell 184(22):5635-5652.e29 (2021); WO 2020 / 191233; WO 2020 / 191234; WO 2020 / 191239; WO 2020 / 191241; WO 2020 / 191242; WO 2020 / 191243; WO 2020 / 191245; WO 2020 / 191246; WO 2020 / 191248; WO 2020 / 191249; WO 2021 / 0823284, each of which is incorporated by reference herein in its entirety. In such cases, the system comprises an engineered T1GR system with nickase activity, a reverse transcriptase domain, and a DNA polymerase, and a TIGR guide or guide molecule comprising a binding sequence capable of hybridizing to the target polynucleotide and an editing sequence. The generated region may be further extended on a DNA template as described herein. The latter may allow generation of a target-independent sequence, compatible with a generic donor sequence.
[0286] The engineered TIGR system is capable of generating a first cleavage of in the target sequence and a second cleavage outside the target sequence on the target polynucleotide. In some variations, a second engineered TIGR system-mediated cleavage in vicinity to the target site may be made, which may enable more efficient invasion of the extended DNA.
[0287] In one example embodiment, multiple engineered TIGR prime editing systems may be used in combination to facilitate larger insertions and deletions. In one embodiment, the compositions and systems of the engineered TIGR system herein comprise: a reverse transcriptase (RT) polypeptide connected to or otherwise capable of forming a complex with the engineered TIGR system; a first TIGR guide or guide molecule capable of forming a first engineered TIGR system-Reverse transcriptase complex with the engineered TIGR system and comprising: a TIGR guide or guide sequence capable of directing site-specific binding of the first engineered TIGR system-Reverse transcriptase complex to a first target sequence of a target polynucleotide; a first binding site region capable of binding to a cleaved or nicked strand of the target polynucleotide; and an RT template sequence encoding a first extended sequence; a second TIGR guide or guide molecule capable of forming a second engineered TIGR system-Reverse transcriptase complex with the engineered TIGR system and comprising: a TIGR guide or guide sequence capable of directing site specific binding of the second engineered TIGR system-Reverse transcriptase complex to a second target sequence of the target polynucleotide; a second binding site region capable of binding to a cleaved or nicked strand of the target polynucleotide; and a RT template sequence encoding a second extended sequence.
[0288] In some cases, the compositions and systems may further comprise: a donor template; a third TIGR or guide sequence capable of forming a engineered TIGR system-Reverse transcriptase complex capable of directing site-specific binding to a target sequence on the donor template; a third binding region capable of binding to a cleaved or nicked strand of the donor template; and a RT template encoding a third extended region complementary to the first extended region generated on the target polynucleotide: and a fourth TIGR guide or guide sequence capable of forming a engineered TIGR system-Reverse transcriptase complex with the TIGR polypeptide or engineered TIGR system and comprising: a TIGR guide or guide sequence capable of directing site-specific binding to a second target sequence on the donor template; a fourth binding region capable of binding to a cleaved or nicked strand of the donor template; and a RT template encoding a fourth extended region complementary to the second extended region generated on the target polynucleotide.
[0289] The use of two engineered TIGR systems (which may also be referred to as doubleflap prime editing or twinPE) may be used to insert, delete, or replace larger sequences. Examples of CRISPR-Cas based prime editing systems that may be adapted for use with the engineered TIGRs described herein are disclosed in WO 2021 / 138469; Anzalone et al. “Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing” Nature Biotechnology 40(5):731-740 (2021); WO 20221 / 226558; WO 2021 / 226558, which are incorporated herein in their entirety by reference.
[0290] In another embodiment, the chimerc TIGR prime editing compositions and systems may further comprise a site-specific recombinase. The recombinase is connected to or otherwise capable of forming a complex with the engineered TIGR prime editing system. In an embodiment, the complex is capable of inserting a recombination site in the DNA loci of interest by extension of RT templates that encode for the recombination site on the 3’ extension of the TIGR guide or guide sequences by the reverse transcriptase. In an embodiment, a donor template comprising a compatible recombination site is provided that can recombine unidirectionally with the inserted recombination site when a recombinase specific for the recombination site is also provided. In an embodiment, the donor template is a plasmid comprising the complementary recombination site and any sequence for insertion at the DNA loci of interest. In an embodiment, the recombinase is connected to or capable of forming a complex with the engineered TIGR systems, such that all of the enzymatic proteins are brought into contact at the loci of interest. In an embodiment, therecombinase is codon optimized for eukaryotic cells (described further herein). In an embodiment, the recombinase includes a NLS (described further herein). In an embodiment, the recombinase is provided as a separate protein. The separate recombinase may form a dimer and bind to the donor template recombination site. The recombinase may be targeted to the loci of interest as a result of the insertion of the compatible recombination site that is also recognized by the recombinase. Thus, the recombinase may recognize the recombination site inserted at the DNA loci of interest and the recombination site on the donor and be targeted to the DNA loci of interest without any additional modifications to the recombinase.
[0291] In an embodiment, a second TIGR complex connected to a recombinase is targeted to the DNA loci of interest. In an embodiment, the second TnpB complex comprises a dead TIGR protein (dTIGR, described further herein), such that the recombinase is targeted to the DNA loci of interest, but the target sequence is not further cleaved. In an embodiment, the dTIGR targets a sequence generated only after the insertion of the recombination site. In an embodiment, the recombinase recognizes and binds to the donor template recombination site and the inserted recombination site. In an embodiment, the recombinase forms a dimer with a recombinase provided as a separate protein.
[0292] As used herein, the term “Recombinase” refers to an enzyme that catalyzes recombination between two or more recombination sites (e.g., an acceptor and donor site). Recombinases useful in the present invention catalyze recombination at specific recombination sites which are specific polynucleotide sequences that are recognized by a particular recombinase. “Uni-directional recombinases” or “integrases” refer to recombinase enzymes whose recognition sites are destroyed after the recombination has taken place. The term “integrase” refers to a type of recombinase. In other words, the sequence recognized by the recombinase is changed into one that is not recognized by the recombinase upon recombination. As a result, once a sequence is subjected to recombination by the uni -directional recombinase, the continued presence of the recombinase cannot reverse the previous recombination event.
[0293] “Recombination sites” are specific polynucleotide sequences that are recognized by the recombinase enzymes described herein. Typically, two different sites are involved (in regards to recombination termed “complementary sites”), one present in the target nucleic acid (e.g., a chromosome or episome of a eukaryote) and another on the nucleic acid that is to be integrated at the target recombination site. The terms “attB” and “attP,” which refer to attachment (orrecombination) sites originally from a bacterial target (attachment site of bacteria) and a phage donor (attachment site of phage), respectively, are used herein although recombination sites for particular enzymes may have different names. The two attachment sites can share as little sequence identity as a few base pairs. The recombination sites typically include left and right arms separated by a core or spacer region. Thus, an attB recombination site consists of BOB', where B and B' are the left and right arms, respectively, and O is the core region. Similarly, attP is POP', where P and P' are the arms and O is again the core region. Upon recombination between the attB and attP sites, and concomitant integration of a nucleic acid at the target, the recombination sites that flank the integrated DNA are referred to as “attL” and “aatR.” The attL and attR sites, using the terminology above, thus consist of BOP' and POB', respectively. In some representations herein, the “O” is omitted and attB and attP, for example, are designated as BB' and PP', respectively. Example CRISPR-Cas prime editing / recombinase compositions and systems that may be adapted for use with the engineered TIGRs disclosed herein are described in WO 2021 / 138469, Anzalone 2021; WO 2021 / 226558; Yarnall et al. “Drag-and-drop genome insertion of large sequences without double-strand DNA cleavage using CRISPR-directed integrases” 41, 500-512 (2023); and WO 2022 / 087235Prime Editor Nucleotide Polymerase Domain
[0294] In one embodiment, a prime editor comprises a nucleotide polymerase domain, e.g., a DNA polymerase domain. The DNA polymerase domain may be a wild-type DNA polymerase domain, a full-length DNA polymerase protein domain, or may be a functional mutant, a functional variant, or a functional fragment thereof. In one embodiment, the polymerase domain is a template dependent polymerase domain. For example, the DNA polymerase may rely on a template polynucleotide strand, e.g., the editing template sequence, for new strand DNA synthesis. In one embodiment, the prime editor comprises a DNA-dependent DNA polymerase. For example, a prime editor having a DNA-dependent DNA polymerase can synthesize a new single-stranded DNA using a PEgRNA editing template that comprises a DNA sequence as a template. In such cases, the PEgRNA is a chimeric or hybrid PEgRNA, and comprises an extension arm comprising a DNA strand. As used herein, an “extension arm” is a polynucleotide portion of a PEgRNA that comprises an editing template and a primer binding site sequence (PBS). In one embodiment, an extension arm further comprises additional components, for example, a 3' modifier. The chimeric or hybrid PEgRNA may comprise an RNA portion (including the spacer and the gRNA core) and a DNA portion (the extension arm comprising the editing template that includes a strand of DNA).
[0295] The DNA polymerases can be wild-type polymerases from eukaryotic, prokaryotic, archael, or viral organisms, and / or the polymerases may be modified by genetic engineering, mutagenesis, or directed evolution-based processes. The polymerases can be a T7 DNA polymerase, T5 DNA polymerase, T4 DNA polymerase, Klenow fragment DNA polymerase, DNA polymerase III and the like. The polymerases can be thermostable, and can include Taq, Tne, Tma, Pfu, Tfl, Tth, Stoffel fragment, VENT® and DEEPVENT® DNA polymerases, KOD, Tgo, JDF3, and mutants, variants and derivatives thereof.
[0296] In one embodiment, the DNA polymerase is a bacteriophage polymerase, for example, a T4, T7, or phi29 DNA polymerase. In one embodiment, the DNA polymerase is an archaeal polymerase, for example, pol I type archaeal polymerase or a pol II type archaeal polymerase. In one embodiment, the DNA polymerase comprises a thermostable archaeal DNA polymerase. In one embodiment, the DNA polymerase comprises a eubacterial DNA polymerase, for example, Pol I, Pol II, or Pol III polymerase. In one embodiment, the DNA polymerase is a Pol I family DNA polymerase. In one embodiment, the DNA polymerase is an E. coli Pol I DNA polymerase. In one embodiment, the DNA polymerase is a Pol II family DNA polymerase. In one embodiment, the DNA polymerase is a Pyrococcusfuriosus (Pfu) Pol II DNA polymerase. In one embodiment, the DNA polymerase is a Pol IV family DNA polymerase. In one embodiment, the DNA polymerase is an E. coli Pol IV DNA polymerase.
[0297] In one embodiment, the DNA polymerase comprises a eukaryotic DNA polymerase. In one embodiment, the DNA polymerase is a Pol-beta DNA polymerase, a Pol-lambda DNA polymerase, a Pol-sigma DNA polymerase, or a Pol-mu DNA polymerase. In one embodiment, the DNA polymerase is a Pol-alpha DNA polymerase. In one embodiment, the DNA polymerase is a POLA1 DNA polymerase. In one embodiment, the DNA polymerase is a POLA2 DNA polymerase. In one embodiment, the DNA polymerase is a Pol-delta DNA polymerase. In one embodiment, the DNA polymerase is a POLDI DNA polymerase. In one embodiment, the DNA polymerase is a POLD2 DNA polymerase. In one embodiment, the DNA polymerase is a human POLDI DNA polymerase. In one embodiment, the DNA polymerase is a human POLD2 DNA polymerase. In one embodiment, the DNA polymerase is a POLD3 DNA polymerase. In one embodiment, the DNA polymerase is a POLD4 DNA polymerase. In one embodiment, the DNA polymerase is a Pol-epsilon DNA polymerase. In one embodiment, the DNA polymerase is a POLE1 DNA polymerase. In one embodiment, the DNA polymerase is a POLE2 DNApolymerase. In one embodiment, the DNA polymerase is a POLE3 DNA polymerase. In one embodiment, the DNA polymerase is a Pol-eta (POLH) DNA polymerase. In one embodiment, the DNA polymerase is a Pol-iota (POLI) DNA polymerase. In one embodiment, the DNA polymerase is a Pol-kappa (POLK) DNA polymerase. In one embodiment, the DNA polymerase is a Revl DNA polymerase. In one embodiment, the DNA polymerase is a human Revl DNA polymerase. In one embodiment, the DNA polymerase is a viral DNA-dependent DNA polymerase. In one embodiment, the DNA polymerase is a B family DNA polymerase. In one embodiment, the DNA polymerase is a herpes simplex virus (HSV) UL30 DNA polymerase. In one embodiment, the DNA polymerase is a cytomegalovirus (CMV) UL54 DNA polymerase.
[0298] In one embodiment, the DNA polymerase is an archaeal polymerase. In one embodiment, the DNA polymerase is a Family B / pol I type DNA polymerase. For example, in one embodiment,, the DNA polymerase is a homolog of Pfu from Pyrococcus furiosus. In one embodiment, the DNA polymerase is a pol II type DNA polymerase. For example, in one embodiment,, the DNA polymerase is a homolog of P. furiosus DP1 / DP22-subunit polymerase. In one embodiment, the DNA polymerase lacks 5' to 3' nuclease activity. Suitable DNA polymerases (pol I or pol II) can be derived from archaea with optimal growth temperatures that are similar to the desired assay temperatures.
[0299] In one embodiment, the DNA polymerase comprises a thermostable archaeal DNA polymerase. In one embodiment, the thermostable DNA polymerase is isolated or derived from Pyrococcus spp. (furiosus, GB-D, woesii, abysii, horikoshii), Thermococcus spp. (kodakaraensis KOD1, litoralis, species 9 degrees North-7, species JDF-3, gorgonarius), Pyrodictium occultum, and Archaeoglobus fulgidus.
[0300] Polymerases may also be from eubacterial species. In one embodiment, the DNA polymerase is a Pol I family DNA polymerase. In one embodiment, the DNA polymerase is an E. coli Pol I DNA polymerase. In one embodiment, the DNA polymerase is a Pol II family DNA polymerase. In one embodiment, the DNA polymerase is a Pyrococcus furiosus (Pfu) Pol II DNA polymerase. In one embodiment, the DNA polymerase is a Pol III family DNA polymerase. In one embodiment, the DNA polymerase is a Pol IV family DNA polymerase. In one embodiment, the DNA polymerase is an E. coli Pol IV DNA polymerase. In one embodiment, the Pol I DNA polymerase is a DNA polymerase functional variant that lacks or has reduced 5' to 3' exonuclease activity.
[0301] Suitable thermostable pol I DNA polymerases can be isolated from a variety of thermophilic eubacteria, including Thermus species and Thermotoga maritima such as Thermus aquaticus (Taq), Thermus thermophilus (Tth) and Thermotoga maritima (Tma UlTma).
[0302] In one embodiment, a prime editor comprises an RNA-dependent DNA polymerase domain, for example, a reverse transcriptase (RT). A RT or an RT domain may be a wild-type RT domain, a full-length RT domain, or may be a functional mutant, a functional variant, or a functional fragment thereof. An RT or an RT domain of a prime editor may comprise a wild type RT, or may be engineered or evolved to contain specific amino acid substitutions, truncations, or variants. An engineered RT may comprise sequences or amino acid changes different from a naturally occurring RT. In one embodiment, the engineered RT may have improved reverse transcription activity over a naturally occurring RT or RT domain. In one embodiment, the engineered RT may have improved features over a naturally occurring RT, for example, improved thermostability, reverse transcription efficiency, or target fidelity. In one embodiment, a prime editor comprising the engineered RT has improved prime editing efficiency over a prime editor having a reference naturally occurring RT.
[0303] In one embodiment, a prime editor comprises a virus RT, for example, a retrovirus RT. Non-limiting examples of virus RT include Moloney murine leukemia virus (M-MLV or MLVRT); human T-cell leukemia virus type 1 (HTLV-1) RT; bovine leukemia virus (BLV) RT; Rous Sarcoma Virus (RSV) RT; human immunodeficiency virus (HIV) RT, M-MFV RT, Avian Sarcoma-Leukosis Virus (ASLV) RT, Rous Sarcoma Virus (RSV) RT, Avian Myeloblastosis Virus (AMV) RT, Avian Erythroblastosis Virus (AEV) Helper Virus MCAV RT, Avian Myelocytomatosis Virus MC29 Helper Virus (MCAV RT, Avian Reticuloendotheliosis Virus (REV-T) Helper Virus REV-A RT, Avian Sarcoma Virus UR2 Helper Virus (UR2AV) RT, Avian Sarcoma Virus Y73 Helper Virus (YAV) RT, Rous Associated Virus (RAV) RT, and Myeloblastosis Associated Virus (MAV) RT, all of which may be suitably used in the methods and composition described herein.
[0304] In one embodiment, the prime editor comprises a wild-type M-MLV RT. An exemplary sequence of a reference M-MLV RT is provided in SEQ ID NO: 22189:TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLI IPLKATSTPVS IKQY PMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPT VPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFDEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASA KKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGTAGFCRLWIPGFAEM AAPLYPLTKTGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQK LGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPP DRWLSNARMTHYQALLLDTDRVQFGPWALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQ PLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGK KLNVYTDSRYAFATAHIHGEIYRRRGLLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQK GHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP (SEQ ID NO: 22189)
[0305] In one embodiment, the prime editor comprises a M-MLV RT comprising one or more of amino acid substitutions P51X, S67X, E69X, L139X, T197X, D200X, H204X, F209X, E302X, T306X, F309X, W313X, T330X, L345X, L435X, N454X, D524X, E562X, D583X, H594X, L603X, E607X, or D653X as compared to the reference M-MLV RT as set forth in SEQ ID NO: 22189, where X is any amino acid other than the wild-type amino acid.
[0306] In one embodiment, the prime editor comprises a M-MLV RT comprising one or more of amino acid substitutions P51L, S67K, E69K, L139P, T197A, D200N, H204R, F209N, E302K, E302R, T306K, F309N, W313F, T330P, L345G, L435G, N454K, D524G, E562Q, D583N, H594Q, L603W, E607K, and D653N as compared to the reference M-MLV RT as set forth in SEQ ID NO: 22189.
[0307] In one embodiment, the prime editor comprises an M-MLV RT comprising one or more amino acid substitutions D200N, T330P, L603W, T306K, and W313F as compared to the reference M-MLV RT as set forth in SEQ ID NO: 295. In one embodiment, the prime editor comprises an M-MLV RT comprising amino acid substitutions D200N, T330P, L603W, T306K, and W313F as compared to the reference M-MLV RT as set forth in SEQ ID NO: 295. In one embodiment, a prime editor comprises an M-MLV RT variant having the sequence as set forth in SEQ ID NO: 22189.
[0308] In one embodiment, an RT variant may be a functional fragment of a reference RT that has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or up to 100, or up to 200, or up to 300, or up to 400, or up to 500 or more amino acid changes compared to a reference RT, e.g., a wild-type RT. In one embodiment, the RT variant comprises a fragment of a reference RT, e.g., a wild-type RT, such that the fragment is about 70% identical, about 80%Illidentical, about 90% identical, about 95% identical, about 96% identical, about 97% identical, about 98% identical, about 99% identical, about 99.5% identical, or about 99.9% identical to the corresponding fragment of the reference RT. In one embodiment, the fragment is 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% of the amino acid length of a corresponding reference M-MLV RT, e.g., SEQ ID NO: 22189.
[0309] In one embodiment, the RT functional fragment is at least 100 amino acids in length. In one embodiment, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, or up to 600 or more amino acids in length. In still other embodiments, the functional RT variant is truncated at the N-terminus or the C-terminus, or both, by a certain number of amino acids which results in a truncated variant which still retains sufficient DNA polymerase function.
[0310] In one embodiment, a prime editor comprises a eukaryotic RT, for example, a yeast, drosophila, rodent, or primate RT. In one embodiment, the prime editor comprises a Group II intron RT, for example, a. Geobacillus stearothermophilus Group II Intron (GsI-IIC) RT or a Eubacterium rectale group II intron (Eu.re. I2) RT. In one embodiment, the prime editor comprises a retron RT.TIGR Retrotransposon Systems
[0311] The systems and compositions herein may comprise a engineered TIGR system, one or more ®RNAs or guide RNAs, and one or more components of a retrotransposon, e.g., a non-LTR retrotransposon. The one or more components of a retrotransposon include a retrotransposon protein and retrotransposon RNA. The systems and compositions may be used to insert a donor polynucleotide to a target polynucleotide. The systems and compositions may further comprise a donor polynucleotide.
[0312] In some examples, the present disclosure provides an engineered, non-naturally occurring composition comprising: a engineered TIGR system, a non-LTR retrotransposon protein associated with or otherwise capable of forming a complex with the engineered TIGR systems; a single coRNA or guide capable of forming a complex with the engineered TIGR systems and directing site-specific binding to a target sequence of a target polynucleotide. The composition may further comprise a donor construct comprising a donor polynucleotide for insertion to the target polynucleotide and located between two binding elements capable of forming a complex with the non-LTR retrotransposon protein. In some cases, the engineered TIGR system is engineered to have nickase activity.
[0313] In some examples, the engineered TIGR system is fused to the N-terminus of the non-LTR retrotransposon protein. In some examples, the engineered TIGR system is fused to the C-terminus of the non-LTR retrotransposon protein.
[0314] The guides may direct the fusion protein to a target sequence 5 ’ of the targeted insertion site, and wherein the engineered TIGR system generates a double-strand break at the targeted insertion site. The guides may direct the fusion protein to a target sequence 3’ of the targeted insertion site, and wherein the engineered TIGR system generates a double-strand break at the targeted insertion site.
[0315] The donor polynucleotide may further comprise a polymerase processing element to facilitate 3’ end processing of the donor polynucleotide sequence. The polymerase may be a DNA polymerase, e.g., DNA polymerase I. In some examples, the polymerase may be an RNA polymerase.
[0316] In some examples, the donor polynucleotide further comprises a homology region to the target sequence on the 5’ end of the donor construct, the 3’ end of the donor construct, or both. In some examples, the homology region is from 1 to 50, from 5 to 30, from 8 to 25, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 base pairs in length.
[0317] Native or wild-type non-LTR retrotransposons encode the protein machinery necessary for their self-mobilization. The non-LTR retrotransposon element comprises a DNA element integrated into a host genome. This DNA element may encode one or two open reading frames (ORFs). For example, the R2 element of Bombyx mori encodes a single ORF containing reverse transcriptase (RT) activity and a restriction enzyme-like (REL) domain. LI elements encode two ORFs, ORF1 and ORF2. ORF1 contains a leucine zipper domain involved in protein-protein interactions and a C-terminal nucleic acid binding domain. ORF2 has a N-terminal apurinic / apyrimidinic endonuclease (APE), a central RT domain, and a C-terminal cysteine histidine rich domain. An example replicative cycle of a non-LTR retrotransposon may comprise transcription of the full-length retrotransposon element to generate an mRNA active element (retrotransposon RNA). The active element mRNA is translated to generate the encoded retrotransposon proteins or polypeptides. A ribonucleoprotein complex comprising the active element and retrotransposon protein, or polypeptide, is formed and this RNP facilitates integration of the active element into the genome. The RNA-transposase complex nicks the genome. The 3’end of the nicked DNA serves as a primer to allow the reverse transcription of the transposon RNA into cDNA. Fourth, the transposase proteins integrate the cDNA into the genome.
[0318] Elements of these systems may be engineered to work within the context of the invention. For example, a non-LTR retrotransposon polypeptide may be fused to a site-specific nuclease. The binding elements that allow a non-LTR retrotransposon polypeptide to bind to the native retrotransposon DNA element, may be engineered into a donor construct to facilitate entry of a donor polynucleotide sequence into a target polypeptide.
[0319] In the present invention the protein component of the non-LTR retrotransposon may be connected to or otherwise engineered to form a complex with a site-specific nuclease. The retrotransposon RNA may be engineered to encode a donor polynucleotide sequence. Thus, in certain example embodiments, the engineered TIGR system, via formation of a engineered TIGR system complex with a guide sequence, directs the retrotransposon complex (e.g., the retrotransposon polypeptide(s) and retrotransposon RNA to a target sequence in a target polynucleotide, where the retrotransposon RNP complex facilitates integration of the donor polynucleotide sequence into the target polynucleotide. Accordingly, the one or more non-LTR retrotransposon components may comprise retrotransposon polypeptides, or function domains thereof, that facilitate binding of the retrotransposon RNA, reverse transcription of the retrotransposon RNA into cDNA, and / or integration of the donor polynucleotide into the target polynucleotide, as well as retrotransposon RNA elements modified to encode the donor polynucleotide sequence.
[0320] Examples of non-LTR retrotransposons include CRE, R2, R4, LI, RTE, Tad, Rl, LOA, I, Jockey, CR1. In one example, the non-LTR retrotransposon is R2. In another example, the non-LTR retrotransposon is LI. Examples of non-LTR retrotransposons may include those described in Christensen SM et al., RNA from the 5' end of the R2 retrotransposon controls R2 protein binding to and cleavage of its DNA target site, Proc Natl Acad Sci U S A. 2006 Nov 21;103(47):17602-7; Eickbush TH et al, Integration, Regulation, and Long-Term Stability of R2 Retrotransposons, Microbiol Spectr. 2015 Apr;3(2): MDNA3-0011-2014. doi: 10.1128 / microbiolspec. MDNA3-0011-2014; Han JS, Non-long terminal repeat (non-LTR) retrotransposons: mechanisms, recent developments, and unanswered questions, Mob DNA. 2010 May 12;1(1): 15. doi: 10.1186 / 1759-8753-1-15; Malik HS et al., The age and evolution of non-LTR retrotransposable elements, Mol Biol Evol. 1999 Jun;16(6):793-805, which are incorporated by reference herein in their entireties.
[0321] Examples of the non-LTR retrotransposon polypeptides also include R2 from Cion orchis sinensis or Zonotrichia albicollis.
[0322] A non-LTR retrotransposon may comprise multiple retrotransposon polypeptides or polynucleotides encoding same. In one embodiment, the retrotransposon polypeptides may form a complex. For example, a non-LTR retrotransposon is a dimer, e.g., comprising two retrotransposon polypeptides forming a dimer. The dimer subunits may be connected or form a tandem fusion. A engineered TIGR system may be associate with (e g., connected to) one or more subunits of such complex. In some examples, the non-LTR retrotransposon is a dimer of two retrotransposon polypeptides; one of the retrotransposon polypeptides comprises nuclease or nickase activity and is connected with a engineered TIGR system.
[0323] The retrotransposon polypeptides may comprise one or more modifications to, for example, enhance specificity or efficiency of donor polynucleotide recognition, target-primed template recognition (TPTR). The retrotransposon polypeptides may also comprise one or more truncations or excisions to remove domains or regions of wild-type protein to arrive at a minimal polypeptide that retain donor polynucleotide recognition and TPTR. In some example embodiments, the native endonuclease activity may be mutated to eliminate endonuclease activity.
[0324] In certain example embodiments, the modifications or truncations of the non-LTR retrotransposon peptide may be in a zinc finger region, a Myb region, a basic region, a reverse transcriptase domain, a cysteine-histidine rich motif, or an endonuclease domain.
[0325] A non-LTR retrotransposon may comprise polynucleotide encoding one or more retrotransposon RNA molecules. The polynucleotide may comprise one or more regulatory elements. The regulatory elements may be promoters. The regulatory elements and promoters on the polynucleotides include those described throughout this application. For example, the polynucleotide may comprise a pol2 promoter, a pol3 promoter, or a T7 promoter.
[0326] In some cases, the polynucleotide encodes a retrotransposon RNA with at least a portion of its sequence complementary to a target sequence. For example, the 3’ end of the retrotransposon RNA may be complementary to a target sequence. The RNA may be complementary to a portion of a nicked target sequence. In one embodiment, a retrotransposonRNA may comprise one or more donor polynucleotides. In certain cases, a retrotransposon RNA may encode one or more donor polynucleotides.
[0327] A retrotransposon RNA may be capable of binding to a retrotransposon polypeptide. Such retrotransposon RNA may comprise one or more elements for binding to the retrotransposon polypeptide. Examples of binding elements include hairpin structures, pseudoknots (e.g., a nucleic acid secondary structure containing at least two stem-loop structures in which half of one stem is intercalated between the two halves of another stem), stem loops, and bulges (e.g., unpaired stretches of nucleotides located within one strand of a nucleic acid duplex). In certain examples, the retrotransposon RNA comprises one or more hairpin structures. In some examples, the retrotransposon RNA comprises one or more pseudoknots. In certain examples, a retrotransposon RNA comprises a sequence encoding a donor polynucleotide and one or more binding elements for forming a complex with the retrotransposon polypeptide. The binding elements may be located on the 5’ end or the 3’ end.
[0328] In one embodiment, a retrotransposon RNA comprises a region capable of hybridizing with an overhang of a target polynucleotide at the target site. The overhang may be a stretch of single-stranded DNA. The overhang may function as a primer for reverse transcription of at least a portion of the retrotransposon RNA to a cDNA. In some cases, a region of the cDNA may be capable of hybridizing a second overhang of the target polynucleotide. The second overhang may function as a primer for the synthesis of a second strand to generate a double-stranded cDNA. The cDNA may comprise a donor polynucleotide sequence. The two overhangs may be from different strands of the target polynucleotide.
[0329] Example CRISPR-Cas, or other programmable nuclease, systems that may be adapted for use with the engineered TIGR systems disclosed herein are disclosed in WO 2021 / 102042; WO 2022 / 17380; and Wilkinson et al. “Structure of the R2 non-LTR retrotransposon initiating a target-primed reverse transcription” Science. 380(6642), 301-308 (2023), which are incorporated herein by reference in their entirety.TIGR Diversity Generating Retroelement System
[0330] In an embodiment, the chimer TIGR systems disclosed herein may further comprise one or more diversity generating retroelement(s) (e.g., DGR described in US20100041033A1). In one embodiment, the DGR may insert a donor polynucleotide with its homing mechanism. For example, the DGR may be associated with a catalytically inactive TIGR protein (e.g., a deadTIGR), and integrate the single-strand DNA using a homing mechanism. In some examples, the DGR may be less mutagenic than a counterpart wild type DGR. In some examples, the DGR is not error-prone. In one embodiment, the DGR herein is not mutagenic. The non-mutagenic DGR may be a mutant of a wild type DGR. As used herein, the term “DGR” encompasses both diversity generating retroelement polynucleotides and proteins encoded by diversity generating retroelement polynucleotides. In some examples, DGR may be proteins encoded by diversity generating retroelement polynucleotides having reverse transcriptase activity. In some examples, DGR may be proteins encoded by diversity generating retroelement polynucleotides having reverse transcriptase activity and integrase activity. In some cases, the template or donor polynucleotide may be encoded by a diversity generating retroelement polynucleotide. In certain cases, the template may be a polynucleotide different from the diversity generating retroelement polynucleotide, e.g., provided as a separate construct or molecule.
[0331] In one embodiment, the DGR herein may also include a Group II intron (and any proteins and polynucleotides encoded), which are mobile ribozymes that self-splice from precursor RNAs to yield excised intron lariat RNAs, which then invade new genomic DNA sites by reverse splicing. Examples of Group II intron include those described in Lambowitz AM et al., Group II Introns: Mobile Ribozymes that Invade DNA, Cold Spring Harb Perspect Biol. 2011 Aug; 3(8): a003616.
[0332] In one embodiment, the diversity-generating retroelements (DGRs) are genetic elements that can produce targeted, massive variations in the genomes that carry these elements. In one embodiment, the DGR systems rely on error-prone reverse transcriptases to produce mutagenized cDNA (containing A-to-N mutations) from a template region (TR), to replace a segment called a variable region (VR) that is similar to the TR region — thi...
Claims
CLAIMSWhat is claimed is:
1. An engineered or non-naturally occurring Tandem Interspaced Guide RNA (TIGR) system, comprisinga Tandem Interspaced Guide RNA (TIGR)-associated (Tas) polypeptide, and a tandem interspersed guide molecule (TIGR guide) capable of forming a complex with the Tas polypeptide and directing sequence specific base pairing with both sense and antisense strands of a double-stranded target oligonucleotide.
2. The system of claim 1, wherein the Tas polypeptide comprises a Nucleolar Protein (Nop) domain.
3. The system of claim 2, wherein the Nucleolar Protein (Nop) domain comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of SEQ ID NO: 2.
4. The system of claim 3, wherein the Nucleolar Protein (Nop) domain has an amino acid sequence of SEQ ID NO: 3.
5. The system of claim 4, wherein the Tas polypeptide further comprises a nuclease domain.
6. The system of claim 5, wherein the nuclease domain comprises an HNH nuclease domain.
7. The system of claim 6, wherein the HNH nuclease domain is covalently linked to a N terminus of the Nop domain.
8. The system of claim 6, wherein the HNH nuclease domain comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of SEQ ID NO: 4.
9. The system of claim 6, wherein the HNH nuclease domain has an amino acid sequence of SEQ ID NO: 5.
10. The system of claim 6, wherein the Tas polypeptide comprising the HNH nuclease domain comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of SEQ ID NO: 6.
11. The system of claim 6, wherein the Tas polypeptide comprising the HNH nuclease domain has an amino acid sequence of SEQ ID NO: 7.
12. The system of claim 6, wherein the Tas polypeptide comprising the HNH nuclease domain is a SpTIGR protein from Salicola phage CGphi29.
13. The system of claim 5, wherein the nuclease domain comprises a RuvC nuclease domain.
14. The system of claim 13, wherein the RuvC nuclease domain is covalently linked to a N terminus of the Nop domain.
15. The system of claim 13, wherein the RuvC nuclease domain comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of SEQ ID NO: 8.
16. The system of claim 13, wherein the RuvC nuclease domain has an amino acid sequence of SEQ ID NO: 9.
17. The system of claim 13, wherein the Tas polypeptide comprising the RuvC nuclease domain comprises an amino acid sequence having at least 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of SEQ ID NO: 10.
18. The system of claim 13, wherein the Tas polypeptide comprising the RuvC nuclease domain has an amino acid sequence of SEQ ID NO: 11.
19. The system of claim 13, wherein the Tas polypeptide a TaTasR polypeptide derived from Thermoproteota archaeon isolate LB CRA l.
20. The system of claim 13. wherein the Tas polypeptide is a ParTasR polypeptide derived from Paracubacteria.
21. The system of claim 1, wherein the Tas polypeptide that lacks a nuclease domain comprises an amino acid sequence having at least 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of SEQ ID NO: 12.
22. The system of claim 21, wherein the Tas polypeptide that lacks a nuclease domain has an amino acid sequence of SEQ ID NO: 13.
23. The system of claim 21, wherein the Tas polypeptide without a nuclease domain is a FpTIGR protein from Flavonifractor plautii.
24. The system of any one of claims 1 to 23, wherein the Tas polypeptide forms a symmetric dimer.
25. The system of any one of claims 1 to 24, wherein the guide molecule comprises nucleotide sequence motifs arranged in a 5’ to 3’ direction: a first edge repeat, a C motif, a first spacer, a D motif, a loop repeat, a C motif, a second spacer a D motif and a second edge repeat.
26. The system of claim 25, wherein the C motif comprises CCA, and the D motif comprises UG27. The system of claim 25, wherein the loop repeat is from 8-12 nucleotides in length.
28. The system of claim 25, wherein the first and the second spacer nucleotide sequences are from 9 to 12 nucleotides in length.
29. The system of claim 25, wherein the first and the second spacer nucleotide sequences are 9 nucleotides in length.
30. The system of claim 25, wherein the first spacer nucleotide sequence is separated from the second spacer nucleotide sequence by a gap sequence.
31. The system of claim 25, wherein the guide molecule is 36 nucleotides in length.
32. The system of any one of claims 1 to 24, wherein the guide molecule comprises 2 or more stem loop unites, each stem loop unit comprising in a 5’ to 3’ direction a first stem sequence, a C motif, the first spacer, a D motif, a variable loop, a C motif, the second spacer, a D motif, and a second stem sequence, wherein the first and second stem sequences base pair to form a stem.
33. The system of claim 32, wherein the C motif comprises CCN and the D motif comprises UG.
34. The system of claim 32, where the guide molecule is 50 to 100 nucleotides in length.
35. The system of claim 34, wherein the guide molecule is 74 nucleotides in length.
36. The system of claim 32, wherein the nucleotide sequence of the first or second spacer is at least 6-15 nucleotides in length.
37. The system of claim 32, wherein each stem loop units are covalently linked in tandem.
38. The system of anyone of claims 1 to 37, wherein the second spacer can base pair with a corresponding complementary nucleotide sequence on the antisense strand of the doublestranded DNA target nucleotide sequence.
39. The system of claim 38, wherein every nucleotide of the first and second spacer nucleotide sequences can base pair with the corresponding complementary nucleotide sequences on the sense or antisense strand of the double-stranded DNA target nucleotide sequence.
40. The system of claim 38, wherein the first spacer can base pair with a corresponding complementary nucleotide sequence on a first stand and the second spacer can base pair with a corresponding complementary nucleotide sequence on the second strand directly adjacent to the reverse-complement of the sequence the first spacer based paired with.
41. The system of claim 38, wherein the guide molecule complex can cleave each strand of the double-stranded DNA target nucleotide sequence immediately 3’ to the nucleotide that is complementary to a 5th base from the 5’ end of the guide molecule first or second spacer nucleotide sequence.
42. The system of claim 38, wherein mutations of two nucleotides adjacent to and 5’ of one of the DNA strand’s cleavage sites inhibit cleavage of that DNA strand without affecting the cleavage of the other DNA strand.
43. The system of claim 42, wherein cleavage of the other DNA strand generates a nicked DNA strand.
44. The system of claim 43, wherein the guide molecule complex can cleave a single-stranded DNA target nucleotide sequence having nucleotide complementarity to either the first or second spacer nucleotide sequence.
45. The system of any one of claims 1 to 23, sequence specific base pairing with both sense and antisense strands does not require a protospacer-adjacent motif (PAM) or target- adjacent motif (TAM).
46. A vector system comprising one or more polynucleotides each encoding one or more components of the engineered systems as set forth in any one of claims 1-44, wherein oneor more polynucleotides are each optionally included in one or more recombinant expression vectors.
47. The vector system of claim 46, wherein the one or more polynucleotides encoding the Tas polypeptide is operably linked to an RNA polymerase II promoter or an inducible promoter.
48. The vector system of claim 46, wherein the one or more polynucleotides encoding the split guide molecule is operably linked to a U snRNA promoter.
49. A delivery system configured to deliver one or more components of the engineered systems of any one of claims 1-44 or a vector system thereof, comprising:a) a Tandem Interspaced Guide RNA (TIGR)-associated (Tas) protein or a nucleotide sequence encoding the Tandem Interspaced Guide RNA (TIGR)-associated (Tas) protein; andb) a split spacer guide RNA or a nucleotide sequence encoding the split spacer guide RNA (guide molecule),wherein the split spacer guide RNA can base pair with both sense and antisense strands of a double-stranded DNA target nucleotide sequence, andwherein the formation of a complex between the Tas protein and the guide molecule facilitates binding of the complex to the target nucleotide sequence.
50. The delivery system of claim 49, further comprising a delivery vehicle comprising a lipid nanoparticle, liposome, a particle, a nanoparticle, an exosome, a microvesicle, a gene gun, or one or more viral vectors.
51. The delivery system of claim 50, wherein the viral vector is an adeno-associated virus (AAV).
52. The delivery system of claim 50, wherein the delivery vehicle is a lipid nanoparticle comprising a ribonucleoprotein complex.
53. The delivery system of claim 50, wherein the delivery vehicle is a lipid nanoparticle comprising one or more mRNAs encoding the TIGR system.
54. A cell comprising any one of the engineered or non-naturally occurring DNA-targeting systems of claims 1-44.
55. The cell of claim 54, wherein the cell is a human cell.
56. A method for modifying a double-stranded target oligonucleotide sequence at a specified locus of interest, comprising delivering to said locus an engineered composition comprising:a) a Tandem Interspaced Guide RNA (TIGR)-associated (Tas) polypeptide or a nucleotide sequence encoding the Tas polypeptide; andb) a split spacer guide RNA (guide molecule) or a nucleotide sequence encoding the guide moleculewherein the split spacer guide RNA comprises a programmable nucleotide sequence that can base pair with both sense and antisense strands of a double-stranded target nucleotide sequence, andwherein formation of a complex between the Tas polypeptide and the guide molecule facilitates binding of the complex to the target nucleotide sequence and triggers a modification of the target nucleotide sequence at the locus of interest.
57. The method of claim 56, wherein the modification comprisesa) one or more nucleotide insertions, deletions, or substitutions,b) correction or introduction of a premature stop codon;c) insertion or deletion of one or more splicing sites;d) insertion of a functional gene;e) disruption or deletion of all or a portion of a gene;f) insertion or deletion of a sequence encoding a post-translational modification site; or g) a combination thereof.
58. The method of claim 56, wherein said modification comprises cleavage of the target nucleotide sequence.
59. The method of claim 56, wherein the target locus of interest is within a cell.
60. The method of claim 56, wherein the cell is a eukaryotic or prokaryotic cell.
61. The method of claim 60, wherein the cell is chosen from a bacterial cell, a plant cell, a fungal cell, a mammalian cell, a reptile cell, an insect cell, an avian cell, a fish cell, a parasite cell, an arthropod cell, a cell of an invertebrate, a cell of a vertebrate, a rodent cell, a mouse cell, a rat cell, a primate cell, and a non-human primate cell.
62. The method of claim 61, wherein the cell is a human cell.
63. The method of claim 62, wherein the cell is a cancer cell.
64. The method of claim 56, wherein the target nucleotide sequence is located in genomic DNA or extrachromosomal DNA.
65. The method of claim 64, wherein the target nucleotide sequence is an oncogene.
66. The method of claim 64, wherein the target nucleotide sequence is a tumor suppressor gene.
67. The method of claim 56, wherein the target nucleotide sequence is associated with a blood disorder.
68. The method of claim 67, wherein the blood disorder is sickle cell disease or betathalassemia.
69. The method of claim 68, wherein the target nucleotide sequence is BCL11 A.
70. The method of claim 56, wherein the target nucleotide sequence is associate with an inflammatory disease.
71. The method of claim 70, wherein the inflammatory disease is transthyretin amyloidosis.
72. The method of claim 71, wherein the target nucleotide sequence encodes TTR protein.
73. The method of claim 56, wherein the target nucleotide sequence is associated with cardiovascular disease.
74. The method of claim 73, wherein the target nucleotide sequence is PCSK9, ANGPTL3, or Lp(a).
75. The method of claim 56, wherein the target nucleotide sequence is viral DNA.
76. The method of claim 75, wherein the viral DNA is papovavirus, human papillomavirus (HPV), hepadnavirus, Hepatitis B Virus (HBV), herpesvirus, varicella zoster virus (VZV), Epstein-Barr virus (EBV), adenovirus, poxvirus, or parvovirus DNA.
77. The method of claim 56, wherein the modification is to a cell ex vivo.
78. The method of claim 77, wherein the cell is a T cell, a CAR T cell, a natural killer (NK) cell, a stem cell.
79. The method of claim 78, wherein the stem cell is a hematopoietic stem cell or a pancreatic derived stem cell.
80. The method of claim 56, wherein the Tas polypeptide comprises an amino acid sequence having at least 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of SEQ ID NO: 6 or 10.
81. The method of claim 56, wherein the Tas protein comprises a Nucleolar Protein (Nop) domain.
82. The method of claim 81, wherein the Nucleolar Protein (Nop) domain comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of SEQ ID NO: 3.
83. The method of claim 81, wherein the Tas protein comprises an HNH nuclease domain or a RuvC nuclease domain.
84. The method of claim 83, wherein the HNH nuclease domain comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of SEQ ID NO: 4.
85. The method of claim 83, wherein the RuvC nuclease domain comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of SEQ ID NO: 8.
86. An engineered Tandem Interspaced Guide RNA (TIGR) recombinase system comprising a Tandem Interspaced Guide RNA (TIGR)-associated (Tas) polypeptide;a recombinase capable of associating with the Tas polypeptide and directing site-specific insertion of a donor insertion sequence from a donor construct at a target insertion site; and a tandem-interspersed guide molecule (TIGR guide) capable of forming a complex with the Tas polypeptide and directing the Tas polypeptide and recombinase to the target insertion site.
87. The system of claim 86, wherein the Tas polypeptide comprises or consists of a polypeptide that is about 90% to 100% identical to any one of SEQ ID NO: 23910-23938.
88. The system of claim 86, wherein the recombinase polypeptide comprises or consists of a sequence that is 90-100% identical to any one of SEQ ID NO: 23939-23967.
89. The system of any one of claims 86 to 88, wherein the recombinase non-covalently complexes with the Tas polypeptide.
90. The system of any one of claims 86 to 88, wherein the recombinase is fused to the Tas polypeptide or covalently attached via a linker.
91. The system of any one of claims 86 to 90, wherein the donor construct comprises recombinase recognition site(s) and a donor insertion sequence.
92. The system of claim 91, wherein the size of the donor insertion sequence is 1 to 20 kb.
93. The system of any one of claims 86 to 92, wherein binding to the target sequence is in a PAM- or TAM-independent manner.
94. A polynucleotide encoding the system of any one of claims 86-93.
95. A vector or vector system comprising a polynucleotide of any one of claim 94.
96. A delivery vehicle configured to deliver one or more components of the system of claims 86 to 93, or a polynucleotide encoding one or more componenet of the system, or the vector or vecto system of claim 95.
97. A cell comprising (a) the system of any one of claims 86-93, the polynucleotide of claim 94, or the vector of claim 95, or the delivery vechicle of claim 96, or a any combination thereof.
98. A method of modifying a target polynucleotide comprising:contacting a target polynucleotide with (a) the system of any one of claims 86-93; (b) a polynucleotide of claim 94; (c) a vector or vector system of claim 95; (d) a delivery vehicle of claim 97; or any combination thereof, whereby the, system mediates site-directed insertion of a donor insert polynucleotide at a target insertion site.
99. An engineered or non-naturally occurring TIGR (Tandem Interspaced Guide RNA) methyltransferase system comprisingi. a Tandem Interspaced Guide RNA (TIGR)-associated (Tas) polypeptide, ii. a methyltransferase (MTase) capable of associating with the Tas polypeptide, and iii. a tandem interspersed guide molecule (TIGR guide) capable of forming a complex with the Tas polypeptide and directing sequence-specific base pairing with both sense and antisense strands of a double-stranded target oligonucleotide sequence.
100. The system of claim Error! Reference source not found., wherein the linker polypeptide is 1-100 or more amino acids in length.
101. The system of claim 99, wherein the methyltransferase (MTase) is covalently bound to the N terminus of the Tas polypeptide.
102. The system of claim 99, wherein the methyltransferase (MTase) is covalently bound to the C terminus of the Tas polypeptide.
103. The system of any one of claims99 to 102, wherein the methyltransferase (MTase) is a DNA methyltransferase.
104. The system of claim 99, wherein the DNA methyltransferase is an adenosine methyltransferase.
105. The system of claim 99, wherein the DNA methyltransferase is a cytosine methyltransferase.
106. The system of claim 105, wherein the cytosine methyltransferase comprises a Dnmtl (DNA (cytosine-5-)-methyltransferase 1), a Dnmt3a (DNA (cytosine-5-)-methyltransferase 3 alpha) and a Dnmt3b (DNA (cytosine-5-)-methyltransferase 3 beta).
107. The system of claim 106, wherein the methyltransferase (MTase) comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence chosen from 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102 and 104.
108. A Tas-associated methyltransferase (Tas-TMase).
109. A Tas-associated methyltransferase (Tas-TMase) comprising a Tas polypeptide having an amino acid sequence with at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with an amino acid sequence of 23997, 23998, 23999, 24000, 24001, 24002, 24003, 24004, 24005, 24006, 24007, 24008, 24009, 24010, and 24011.
110. The Tas-associated methyltransferase (Tas-TMase) of claim 108, further comprising a DNA methyltransferase.
111. The Tas-associated methyltransferase (Tas-TMase) of claim 108, wherein the DNA methyltransferase is an adenosine methyltransferase or a cytosine methyltransferase.
112. The Tas-associated methyltransferase (Tas-TMase) of claim 108, wherein the cytosine methyltransferase comprises aDnmtl (DNA(cytosine-5-)-methyltransferase 1), aDnmt3a (DNA (cytosine-5-)-methyltransferase 3 alpha) and a Dnmt3b (DNA (cytosine-5-)- methyltransferase 3 beta).
113. The Tas-associated methyltransferase (Tas-TMase) of claim 108, wherein the Tas polypeptide is noncovalently bound to the methyltransferase (MTase).
114. The Tas-associated methyltransferase (Tas-TMase) of claim 108, wherein the Tas polypeptide is covalently bound to the methyltransferase (MTase).
115. A method for methylating a base of a double-stranded DNA target nucleotide sequence at a locus of interest, comprising delivering to said locus an engineered composition comprising:a) a Tandem Interspaced Guide RNA (TIGR)-associated (Tas) polypeptide or a nucleotide sequence encoding the Tas polypeptide;b) a methyltransferase bound to the Tas polypeptide, andc) a split spacer guide RNA (guide molecule), or a nucleotide sequence encoding the guide molecule,i. wherein the split spacer guide RNA comprises a programmable nucleotide sequence that can base pair with both sense and antisense strands of the doublestranded DNA target nucleotide sequence, andii. wherein formation of a complex between the Tas polypeptide, the methyltransferase and the split spacer guide RNA (tigRNA) facilitates binding of the complex to the target nucleotide sequence and triggers methylation of a base of the target nucleotide sequence at the locus of interest.