Transposases and uses thereof
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2026-04-08
AI Technical Summary
Current transposases, such as PiggyBac, exhibit limited activity in human cells in culture and the portability of transposons remains unclear, with limited applications in genome editing due to restricted activity across various cell types.
Development of novel transposases from diverse species, including Poeciliopsis turrubarensis, Anthonomus grandis, and Xenopus tropicalis, which are optimized for enhanced transposition activity across multiple cell types, potentially combined with RNA-guided nucleases for targeted gene integration.
These novel transposases demonstrate improved activity and specificity in integrating transgenes into genomes across various cell types, enhancing the efficacy of genome editing tools for therapeutic and research applications.
Smart Images

Figure IMGF000028_0001 
Figure IMGF000029_0001 
Figure IMGF000031_0001
Abstract
Description
NOVEL TRANSPOSASES AND USES THEREOFFIELD OF INVENTION
[0001] The present invention relates to novel transposases and their use for gene editing.BACKGROUND OF INVENTION
[0002] PiggyBac (PB) derived transposase can integrate large cargos across a variety of cellular backgrounds enabling re-writing genetic information for therapeutic use, biomedical research, etc. Re-factoring self-mobilizing genomic elements lead to development of genome engineering tools. Further identification of PB transposases in several species suggested widespread distribution of this transposon system.
[0003] Identification of several piggy bac families further confirmed these observations (Choudhary, M.N.K., et al., Nat. Commun., 2023), together with cases of acquisition of transposases for cellular functions. The inventors explored the existing diversity of PiggyBac family of transposases by phylogenetic mapping and functional classification of PB features.
[0004] How portable are transposons remains an elusive fact, with limited activity reported in human cells in culture.
[0005] The extent to which language mode derived fitness predictions can be useful in genome mined protein reconstruction, is still to be explored. The inventors further demonstrate that protein language models can repair Transposon ORFs maximizing transposition activity across a variety of cell types.
[0006] The present invention provides novel transposases maximizing transposition activity across a variety of cell types.SUMMARY
[0007] This invention thus relates to a composition comprising at least one transposase ortholog.
[0008] In particular, the present invention relates to a composition comprising at least one transposase selected from the group consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof; an Anthonomus grandis DR1754440 transposase, or a variant thereof; an Anthonomus grandis DR1754053 transposase, or a variant thereof; a Anthonomus grandis DR1756049 transposase, or a variant thereof; a Xenopus tropicalis transposase, or a variant thereof; a Japanese Medaka transposase, or a variant thereof; a Leptobrachium leishanense transposase, or a variant thereof; a Scalopus aquaticus ScaAqu-5.3491 transposase, or a variant thereof; an Atlantic Salmon transposase, or a variant thereof; a Heliconius butterfly transposase, or a variant thereof; a Takifugu flavidus transposase, or a variant thereof; a Vaquita transposase, or a variant thereof; a Oryzias latipes transposase, or a variant thereof; a Salvelinus profundus transposase, or a variant thereof; a Cuniculus paca transposase, or a variant thereof; a Noctilio leporinus transposase, or a variant thereof; a Pipistrellus pipistrellus PipPip-6.1914 transposase, or a variant thereof; a Carlito syrichta transposase, or a variant thereof; a Myotis lucifugus PiggyBat transposase, or a variant thereof; a PiggyBac transposase, or a variant thereof; a_Solenopsis invicta DR3053925 transposase, or a variant thereof; a Simochromis diagramma genomic transposase, or a variant thereof; a Nematolebias whitei chromosome 17 transposase, or a variant thereof; and a Coremacera marginata DR1481656 transposase, or a variant thereof, or a nucleic acid encoding thereof.
[0009] In particular, the present invention relates to a composition comprising at least one transposase selected from the group consisting of:Poeciliopsis turrubarensis transposase, or a variant thereof;Anthonomus grandis DR1754440 transposase, or a variant thereof;Anthonomus grandis DR1754053 transposase, or a variant thereof; andAnthonomus grandis DR1756049 transposase, or a variant thereof; or a nucleic acid encoding thereof.
[0010] In some embodiments, the at least one transposase is selected from the group consisting of:Poeciliopsis turrubarensis transposase, or a variant thereof; andAnthonomus grandis DR1754440 transposase, or a variant thereof.
[0011] In some embodiments, the at least one transposase is a Poeciliopsis turrubarensis transposase, or a variant thereof.
[0012] In some embodiments, the Poeciliopsis turrubarensis transposase has an amino acid sequence having at least 75% sequence identity with SEQ ID NO: 1; for example having at least 75%, 80%, 85%, 90% 95% or 100% sequence identity with said sequence.
[0013] In some embodiments, the Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 1.
[0014] In some embodiments, the Poeciliopsis turrubarensis transposase variant has an amino acid sequence as set forth in any one of SEQ ID NO: 2 to SEQ ID NO: 13 or SEQ ID NO: 112.
[0015] In some embodiments, the Poeciliopsis turrubarensis transposase variant comprises at least one amino acid substitution on one or more of the amino acids atpositions 353, 356, and 432 corresponding to the amino acid numbering of SEQ ID NO: 1.
[0016] In some embodiments, the Poeciliopsis turrubarensis transposase variant has an amino acid sequence as set forth in SEQ ID NO: 112.
[0017] In some embodiments, the at least one transposase is an Anthonomus grandis DR1754440 transposase, or a variant thereof.
[0018] In some embodiments, the Anthonomus grandis DR1754440 transposase has an amino acid sequence having at least 75% sequence identity with SEQ ID NO: 30; for example having at least 75%, 80%, 85%, 90% 95% or 100% sequence identity with said sequence.
[0019] In some embodiments, the Anthonomus grandis DR1754440 transposase has an amino acid sequence as set forth in SEQ ID NO: 30.
[0020] In some embodiments, the Anthonomus grandis DR1754440 transposase variant comprises at least one amino acid substitution on one or more of the amino acids at positions 388, 389, and 393 corresponding to the amino acid numbering of SEQ ID NO: 30.
[0021] In some embodiments, the Anthonomus grandis DR1754440 transposase variant has an amino acid sequence as set forth in SEQ ID NO: 113.
[0022] In some embodiments, the at least one transposase is an Anthonomus grandis DR1756049 transposase, or a variant thereof.
[0023] In some embodiments, the Anthonomus grandis DR1756049 transposase has an amino acid sequence having at least 75% sequence identity with SEQ ID NO: 35 or SEQ ID NO: 108.
[0024] In some embodiments, the Anthonomus grandis DR1756049 transposase has an amino acid sequence as set forth in SEQ ID NO: 35 or SEQ ID NO: 108.
[0025] In some embodiments, the Anthonomus grandis DR1756049 transposase variant comprises at least one amino acid substitution on one or more of the amino acids at positions 348, 352, and 429 corresponding to the amino acid numbering of SEQ ID NO: 35.
[0026] In some embodiments, the Anthonomus grandis DR1756049 transposase variant has an amino acid sequence as set forth in SEQ ID NO: 111.
[0027] In some embodiments, the at least one transposase is an Anthonomus grandis DR1754053 transposase, or a variant thereof.
[0028] In some embodiments, the Anthonomus grandis DR1754053 transposase has an amino acid sequence having at least 75% sequence identity with SEQ ID NO: 31; for example having at least 75%, 80%, 85%, 90% 95% or 100% sequence identity with said sequence.
[0029] In some embodiments, the Anthonomus grandis DR1754053 transposase has an amino acid sequence as set forth in SEQ ID NO: 31.
[0030] In some embodiments, the transposase recognizes at least one ITR sequence, preferably at least two ITR sequences, selected from the group consisting of SEQ ID NO: 37 to SEQ ID NO: 84.
[0031] In some embodiments, the Poeciliopsis turrubarensis transposase, or variant thereof, recognizes a left ITR sequence having at least 75% sequence identity with SEQ ID NO: 37, and a right ITR sequence having at least 75% sequence identity with SEQ ID NO: 38.
[0032] In some embodiments, the Anthonomus grandis DR1754440 transposase, or variant thereof, recognizes a left ITR sequence having at least 75% sequence identity with SEQ ID NO: 71, and a right ITR sequence having at least 75% sequence identity with SEQ ID NO: 72.
[0033] In some embodiments, the Anthonomus grandis DR1756049 transposase, or variant thereof, recognizes a left ITR sequence having at least 75% sequence identity withSEQ ID NO: 81, and a right ITR sequence having at least 75% sequence identity with SEQ ID NO: 82.
[0034] In some embodiments, the Anthonomus grandis DR1754053 transposase, or variant thereof, recognizes a left ITR sequence having at least 75% sequence identity with SEQ ID NO: 73, and a right ITR sequence having at least 75% sequence identity with SEQ ID NO: 74.
[0035] In some embodiments, the composition further comprises a RNA-guided nuclease or nickase.
[0036] In some embodiments, the transposase and the RNA-guided nuclease or nickase are fused together by a covalent linkage or a non-covalent linkage.
[0037] In some embodiments, the covalent linkage comprises a linker.
[0038] In some embodiments, the transposase and the RNA-guided nuclease or nickase are not fused together.
[0039] In some embodiments, the composition further comprises at least one nucleic acid molecule comprising at least one transgene of interest to integrate into the genome of one or more cell.
[0040] The present invention further relates to an in vitro method for the integration of at least one transgene of interest into the genome of one or more cell, comprising contacting said one or more cell with the composition according to the invention, and at least one nucleic acid molecule comprising said at least one transgene of interest.
[0041] In some embodiments, the method is for a targeted integration of the at least one transgene of interest into the genome of one or more cell, wherein the composition further comprises a RNA-guided nuclease or nickase enabling said targeted integration.
[0042] The present invention further relates to a pharmaceutical composition comprising the composition according to the invention.
[0043] The present invention further relates to the composition according to the invention, or the pharmaceutical composition according to the invention, for use for treating a genetic disease in a subject in need thereof.DEFINITIONS
[0044] In the present text, the following terms have the following meanings:
[0045] The term “transposase” refers to an enzyme that binds to the end of a transposon and catalyzes its movement to another part of the genome by a cut-and-paste mechanism or a replicative transposition mechanism.
[0046] The terms “ortholog” and “orthologs” refer to genes that evolved in different species from a common ancestral gene by speciation. In general, orthologs retain the same function during evolution. By extension, “ortholog” and “orthologs” also refer to proteins, particularly to transposases, encoded by these genes, that evolved in different species from a common ancestral transposase. Therefore, transposase orthologs retain slightly the same function during evolution, of binding to the end of a transposon and catalyzing its movement to another part of the genome by a cut-and-paste mechanism or by a replicative transposition mechanism.
[0047] The terms “nucleic acid sequence” and “nucleotide sequence” may be used interchangeably to refer to any molecule composed of, or comprising, monomeric nucleotides. A nucleic acid may be an oligonucleotide or a polynucleotide. A nucleotide sequence may be a DNA, RNA, or a mix thereof. A nucleotide sequence may be chemically-modified or artificial.
[0048] The term “transgene” refers to an exogenous nucleic acid sequence, in particular an exogenous DNA or cDNA encoding a gene product. The gene product may be an RNA, peptide or protein. In addition to the coding region for the gene product (CDS), the transgene may include or be associated with one or more operational sequences to facilitate or enhance expression, such as a promoter, enhancer(s), response element(s), reporter element(s), insulator element(s), polyadenylation signal(s) and / or otherfunctional elements. Embodiments of the disclosure may utilize any known suitable promoter, enhancer(s), response element(s), reporter element(s), insulator element(s), polyadenylation signal(s) and / or other functional elements, unless specified otherwise. Suitable elements and sequences will be well known to those skilled in the art.
[0049] The terms "amino acid sequence ", "polypeptide ", "peptide” and "protein"' are used interchangeably to refer to a polymer of amino acid residues. Unless specified, a polymer of amino acid residues can be of any length. The terms also apply to amino acid polymers in which one or more amino acids are chemical analogues or modified derivatives of corresponding naturally-occurring amino acids.
[0050] The term “binding protein” refers to a protein that is able to bind non-covalently to another molecule. A binding protein can bind to, for example, a DNA molecule (a DNA-binding protein), an RNA molecule (an RNA-binding protein) and / or a protein molecule (a protein-binding protein). In the case of a protein-binding protein, it can bind to one or more molecules of the same protein to form homodimers, homotrimers, etc.; and / or it can bind to one or more molecules of a different protein or proteins. A binding protein can have more than one type of binding activity. For example, zinc finger proteins have DNA-binding, RNA-binding and protein-binding activity.
[0051] The terms “Cas9” or “Cas9 nuclease” refer to an RNA-guided nuclease comprising a Cas9 protein, or a fragment thereof (e.g., a protein comprising an active or inactive DNA cleavage domain of Cas9, and / or the gRNA binding domain of Cas9). A Cas9 nuclease is also referred to sometimes as a casnl nuclease or a CRISPR (Clustered Regularly Interspaced Short Palindromic Repeat)-associated nuclease. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to antecedent mobile elements, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of pre-crRNA requires a transencoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc) and a Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-aided processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circulardsDNA target complementary to the spacer. The target strand not complementary to crRNA is first cut endonucleolytically, then trimmed 3 ‘-5’ exonucleolytically. In nature, DNA-binding and cleavage typically requires protein and both RNAs. However, single guide RNAs (“sgRNA” or simply “gNRA”) can be engineered so as to incorporate aspects of both the crRNA and tracrRNA into a single RNA species. Cas9 recognizes a short motif in the CRISPR repeat sequences (the PAM or protospacer adjacent motif) to help distinguish self vs. non-self. Cas9 nuclease sequences and structures are well known to those of skill in the art. Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and 5. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski et al., 2013. RNA Biol. 10(5):726-37), the entire content of which is incorporated herein by reference. In some embodiments, a Cas9 nuclease has an inactive (e.g., an inactivated) DNA cleavage domain. A nuclease-inactivated Cas9 protein can interchangeably be referred to as a “dCas9” protein (for nuclease-“dead” Cas9). Methods for generating a Cas9 protein (or a fragment thereof) having an inactive DNA cleavage domain are known in the art (see, e.g., Jinek et al., 2012. Science. 337(6096):816-821; Qi et al., 2013. Cell. 152(5): 1173-83, the entire content of each being incorporated herein by reference).
[0052] The term “fusion” refers to a molecule in which two or more subunit molecules are linked. In some embodiments, the link between the two is covalent; alternatively, the link between the two can be non-covalent and rely, e.g., on intermolecular interactions. The subunit molecules can be the same chemical type of molecule, or can be different chemical types of molecules.
[0053] The term “fusion protein” refers to a hybrid polypeptide which comprises protein domains from at least two different proteins. For example, one protein domain may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy-terminal (C-terminal) protein, thus forming an “amino-terminal fusion protein” or a “carboxy-terminal fusion protein”, respectively. In preferred embodiments, a fusion protein is a single chain polypeptide which may be fully encoded by a nucleic acidsequence, and includes at least two protein domains directly covalently linked by peptidic bound or optionally covalently linked via a peptidic linker.
[0054] The terms "gene" or "genome" as used herein, includes a DNA region encoding a gene product, as well as all DNA regions which regulate the production of the gene product, whether or not such regulatory sequences are adjacent to coding and / or transcribed sequences. Accordingly, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment sites and locus control regions.
[0055] The term "eukaryotic cells” include, but are not limited to, fungal cells (such as yeast), plant cells, animal cells, mammalian cells and human cells (e.g., T-cells).
[0056] The term "linked" as used herein, refers to the juxtaposition of two or more components (such as sequence elements), in which the components are arranged such that both components function normally and allow the possibility that at least one of the components can mediate a function that is exerted upon at least one of the other components.
[0057] The term “specificity” refers to the ability to selectively bind a sequence which shares a degree of sequence identity to a selected sequence.
[0058] The term “at least 75% of sequence identity” with a reference sequence, in particular a polypeptide sequence, is meant to encompass having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity; for example, any sub-range comprising or consisting of 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, with the reference sequence.
[0059] The terms “insertion” and “integration” refer to the addition of a nucleic acid sequence into a second nucleic acid sequence or into a genome or part thereof. The terms “specific”, “site-specific”, “targeted” and “on-targeted” in relation to insertion or integration, are used herein interchangeably to refer to the insertion of a nucleic acid intoa specific site of a second nucleic acid or into a specific site of a genome or part thereof. Conversely, the terms “random”, “non-targeted” and “off-targeted” refer to non-specific and unintended insertion of a nucleic acid into an unwanted site. The terms “total” or “overall” refer to the total number of insertions.
[0060] The term “linker” refers to a chemical group or a molecule linking two adjacent molecules or moieties.
[0061] The terms "vector" and "plasmid" as used herein, refer to any polynucleotide that can carry, e.g., a second polynucleotide of interest, and e.g., which can transfer gene sequences to target cells. Thus, the term includes cloning, and expression vehicles, as well as integrating vectors. Particularly, the term "expression vector," as used herein, refers to any polynucleotide capable of directing the expression of a nucleic acid. In some aspects, the terms "vector" and "plasmid" are used interchangeably with the term "nucleic acid construct."
[0062] As used herein, the “percent identity” between two sequences is a function of the number of identical positions shared by the sequences (i. e., % identity = number of identical positions / total number of positions x 100), taking into account the number of gaps, and the length of each gap, which need to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm, as described below. The percent identity between two amino acid sequences can be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl. Biosci., 4:11- 17, 1988) which has been incorporated into the ALIGN program (version 2.0), using a PAM 120 weight residue table, a gap length penalty of 12 and a gap penalty of 4. Alternatively, the percent identity between two amino acid sequences can be determined using the Needleman and Wunsch (J. Mol, Biol. 48:444-453, 1970) algorithm which has been incorporated into the GAP program in the GCG software package (available at http: / / www.gcg.com), using either a Blossom 62 matrix or a PAM250 matrix, and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a length weight of 1, 2, 3, 4, 5, or 6. The percent identity between two nucleotide amino acid sequences may also be determined using for example algorithms such as the BLASTN program for nucleic acid sequences using asdefaults a word length (W) of 11, an expectation (E) of 10, M=5, N=4, and a comparison of both strands.
[0063] The term “subject” as used herein, refers to an individual organism, for example, an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, a goat, a cattle, a cat, or a dog. In some embodiments, the subject is a vertebrate, an amphibian, a reptile, a fish, an insect, a fly, or a nematode. In some embodiments, the subject is a research animal.
[0064] The terms “treatment,” “treat,” and “treating,” refer to a clinical intervention aimed to reverse, alleviate, delay the onset of, or inhibit the progress of a disease or disorder, or one or more symptoms thereof, as described herein. As used herein, the terms “treatment,” “treat,” and “treating” refer to a clinical intervention aimed to reverse, alleviate, delay the onset of, or inhibit the progress of a disease or disorder, or one or more symptoms thereof, as described herein. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent, reduce the likelihood of developing, or delay onset of a symptom or inhibit onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, for example, to prevent or delay their recurrence.DETAILED DESCRIPTION
[0065] Novel enzymes having transposase activity are presented herein, and their subsequent use for gene editing, in particular targeted gene insertion.
[0066] This invention relates to a composition comprising at least one protein having a transposase activity, or a nucleic acid encoding thereof. As used herein, a protein having transposase activity will be referred to as “transposase”; the term “transposase” thus encompasses transposases orthologs.
[0067] Thus, the present invention relates to a composition comprising at least one transposase, or a nucleic acid encoding thereof.
[0068] In some embodiments, the transposase is selected from the group comprising or consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof; a Xenopus tropicalis transposase herein interchangeably referred to as “tropical clawed frog” transposase, or a variant thereof; a Japanese Medaka transposase, or a variant thereof; a Leptobrachium leishanense transposase, herein interchangeably referred to as “Leishan spiny toad” transposase, or a variant thereof; a Scalopus aquaticus transposase, herein interchangeably referred to as “ScaAqu- 5.3491 transposase”, or a variant thereof; an Atlantic Salmon transposase, or a variant thereof; a Heliconius butterfly transposase, or a variant thereof; a Takifugu flavidus transposase herein interchangeably referred to as “Yellowbelly pufferfish” transposase, or a variant thereof; a Vaquita transposase, or a variant thereof; a Oryzias latipes transposase, herein interchangeably referred to as “Japanese Rice Fish” transposase, or a variant thereof; a Salvelinus profundus transposase, or a variant thereof;a Cuniculus paca transposase, or a variant thereof; a Noctilio leporinus transposase, or a variant thereof; a Pipistrellus pipistrellus transposase, herein interchangeably referred to as “PipPip-6.1914 transposase”, or a variant thereof; a Carlito syrichta transposase, herein interchangeably referred to as “Philippine tarsier” transposase, or a variant thereof; a Myotis lucifugus transposase, herein interchangeably referred to as “PiggyBat” transposase, or a variant thereof; a PiggyBac transposase, or a variant thereof; an Anthonomus grandis transposase, herein interchangeably referred to as “DR1754440” transposase, or a variant thereof; an Anthonomus grandis transposase, herein interchangeably referred to as “DR1754053” transposase, or a variant thereof; _Solenopsis invicta transposase, herein interchangeably referred to as “DR3053925” transposase, or a variant thereof; a Simochromis diagramma transposase, herein interchangeably referred to as “Simochromis diagramma genomic” transposase, or a variant thereof; a Nematolebias whitei transposase, herein interchangeably referred to as “Nematolebias whitei chromosome 17”, or a variant thereof; a _Anthonomus grandis transposase, herein interchangeably referred to as “DR1756049”, or a variant thereof; and_ a Coremacera marginata transposase, herein interchangeably referred to as “DR1481656”, or a variant thereof.
[0069] In some embodiments, the transposase is selected from the group consisting of:a Poeciliopsis turrubarensis transposase, or a variant thereof; a Xenopus tropicalis transposase herein interchangeably referred to as “tropical clawed frog” transposase, or a variant thereof; a Japanese Medaka transposase, or a variant thereof; a Leptobrachium leishanense transposase, herein interchangeably referred to as “Leishan spiny toad” transposase, or a variant thereof; a Scalopus aquaticus transposase, herein interchangeably referred to as “ScaAqu- 5.3491 transposase”, or a variant thereof; an Atlantic Salmon transposase, or a variant thereof; a Heliconius butterfly transposase, or a variant thereof; a Takifugu flavidus transposase herein interchangeably referred to as “Yellowbelly pufferfish” transposase, or a variant thereof; a Vaquita transposase, or a variant thereof; a Oryzias latipes transposase, herein interchangeably referred to as “Japanese Rice Fish” transposase, or a variant thereof; a Salvelinus profundus transposase, or a variant thereof; a Cuniculus paca transposase, or a variant thereof; a Noctilio leporinus transposase, or a variant thereof; a Pipistrellus pipistrellus transposase, herein interchangeably referred to as “PipPip-6.1914 transposase”, or a variant thereof; a Carlito syrichta transposase, herein interchangeably referred to as “Philippine tarsier” transposase, or a variant thereof;a Myotis lucifugus transposase, herein interchangeably referred to as “PiggyBaf transposase, or a variant thereof; a PiggyBac transposase, or a variant thereof; an Anthonomus grandis transposase, herein interchangeably referred to as “DR1754440” transposase, or a variant thereof; an Anthonomus grandis transposase, herein interchangeably referred to as “DR1754053” transposase, or a variant thereof; _Solenopsis invicta transposase, herein interchangeably referred to as “DR3053925” transposase, or a variant thereof; a Simochromis diagramma transposase, herein interchangeably referred to as “Simochromis diagramma genomic” transposase, or a variant thereof; a Nematolebias whitei transposase, herein interchangeably referred to as “Nematolebias whitei chromosome 17”, or a variant thereof; a _Anthonomus grandis transposase, herein interchangeably referred to as “DR1756049”, or a variant thereof; and_ a Coremacera marginata transposase, herein interchangeably referred to as “DR1481656”, or a variant thereof.
[0070] In some embodiments, the transposase is selected from the group consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 1, preferably a variant thereof having the amino acid sequence of any one of SEQ ID NO: 2 to SEQ ID NO: 13 or SEQ ID NO: 112, more preferably a variant thereof having the amino acid sequence of SEQ ID NO: 112;an Anthonomus grandis DR1754440 transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 30, preferably a variant thereof having the amino acid sequence of SEQ ID NO: 113; an Anthonomus grandis DR1756049 transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 35 or SEQ ID NO: 108, preferably a variant thereof having the amino acid sequence of SEQ ID NO: 111; an Anthonomus grandis transposase DR1754053, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 31; a Xenopus tropicalis transposase or a variant thereof having at least 75% sequence identity with SEQ ID NO: 14; a Japanese Medaka transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 15; a Leptobrachium leishanense transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 16; a Scalopus aquaticus transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 17; an Atlantic Salmon transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 18; a Heliconius butterfly transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 19; a Takifugu flavidus transposase or a variant thereof having at least 75% sequence identity with SEQ ID NO: 20; a Vaquita transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 21;a Oryzias latipes transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 22; a Salvelinus profundus transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 23; a Cuniculus paca transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 24; a Noctilio leporinus transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 25; a Pipistrellus pipistrellus transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 26; a Carlito syrichta transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 27; a Myotis lucifugus transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 28; a PiggyBac transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 29; _Solenopsis invicta transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 32; a Simochromis diagramma transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 33; a Nematolebias whitei transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 34; and a Coremacera marginata transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 36.
[0071] In some embodiments, the transposase, or a variant thereof is selected from the group comprising or consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof; a Xenopus tropicalis transposase herein interchangeably referred to as “tropical clawed frog” transposase, or a variant thereof; a Japanese Medaka transposase, or a variant thereof; a Leptobrachium leishanense transposase, herein interchangeably referred to as “Leishan spiny toad” transposase, or a variant thereof; a Scalopus aquaticus transposase, herein interchangeably referred to as “ScaAqu- 5.3491 transposase”, or a variant thereof; an Atlantic Salmon transposase, or a variant thereof; a Heliconius butterfly transposase, or a variant thereof; a Takifugu flavidus transposase herein interchangeably referred to as Yellowbelly pufferfish transposase, or a variant thereof; a Vaquita transposase, or a variant thereof; a Oryzias latipes transposase, herein interchangeably referred to as “Japanese Rice Fish” transposase, or a variant thereof; a Salvenius profundus transposase, or a variant thereof; a Cuniculus paca transposase, or a variant thereof; a Noctilio leporinus transposase, or a variant thereof; a Pipistrellus pipistrellus transposase, herein interchangeably referred to as “PipPip-6.1914 transposase”, or a variant thereof;a Carlito syrichta, herein interchangeably referred to as “Philippine tarsier’ transposase, or a variant thereof; an Anthonomus grandis transposase, herein interchangeably referred to as “DR1754440” transposase, or a variant thereof; an Anthonomus grandis transposase, herein interchangeably referred to as “DR1754053” transposase, or a variant thereof; _Solenopsis invicta transposase, herein interchangeably referred to as “DR3053925” transposase, or a variant thereof; a Simochromis diagramma transposase, herein interchangeably referred to as “Simochromis diagramma genomic” transposase, or a variant thereof; a Nematolebias whitei transposase, herein interchangeably referred to as “Nematolebias whitei chromosome 17”, or a variant thereof; a Anthonomus grandis transposase, herein interchangeably referred to as “DR1756049”, or a variant thereof; and a Coremacera marginata transposase, herein interchangeably referred to as “DR1481656”, or a variant thereof; or a nucleic acid encoding thereof.
[0072] In some embodiments, the transposase is selected from the group consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 1, preferably a variant thereof having the amino acid sequence of any one of SEQ ID NO: 2 to SEQ ID NO: 13 or SEQ ID NO: 112, more preferably a variant thereof having the amino acid sequence of SEQ ID NO: 112;an Anthonomus grandis DR1754440 transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 30, preferably a variant thereof having the amino acid sequence of SEQ ID NO: 113; an Anthonomus grandis DR1756049 transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 35 or SEQ ID NO: 108, preferably a variant thereof having the amino acid sequence of SEQ ID NO: 111; an Anthonomus grandis transposase DR1754053, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 31; a Xenopus tropicalis transposase or a variant thereof having at least 75% sequence identity with SEQ ID NO: 14; a Japanese Medaka transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 15; a Leptobrachium leishanense transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 16; a Scalopus aquaticus transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 17; an Atlantic Salmon transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 18; a Heliconius butterfly transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 19; a Takifugu flavidus transposase or a variant thereof having at least 75% sequence identity with SEQ ID NO: 20; a Vaquita transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 21;a Oryzias latipes transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 22; a Salvelinus profundus transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 23; a Cuniculus paca transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 24; a Noctilio leporinus transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 25; a Pipistrellus pipistrellus transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 26; a Carlito syrichta transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 27; a Myotis lucifugus transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 28; _Solenopsis invicta transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 32; a Simochromis diagramma transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 33; a Nematolebias whitei transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 34; and a Coremacera marginata transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 36.
[0073] In some embodiments, the transposase, or a variant thereof is selected from the group consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof;a Xenopus tropicalis transposase herein interchangeably referred to as “tropical clawed frog” transposase, or a variant thereof; a Japanese Medaka transposase, or a variant thereof; a Leptobrachium leishanense transposase, herein interchangeably referred to as “Leishan spiny toad” transposase, or a variant thereof; a Scalopus aquaticus transposase, herein interchangeably referred to as “ScaAqu- 5.3491 transposase”, or a variant thereof; an Atlantic Salmon transposase, or a variant thereof; a Heliconius butterfly transposase, or a variant thereof; a Takifugu flavidus transposase herein interchangeably referred to as Yellowbelly pufferfish transposase, or a variant thereof; a Vaquita transposase, or a variant thereof; a Oryzias latipes transposase, herein interchangeably referred to as “Japanese Rice Fish” transposase, or a variant thereof; a Salvenius profundus transposase, or a variant thereof; a Cuniculus paca transposase, or a variant thereof; a Noctilio leporinus transposase, or a variant thereof; a Pipistrellus pipistrellus transposase, herein interchangeably referred to as “PipPip-6.1914 transposase”, or a variant thereof; a Carlito syrichta, herein interchangeably referred to as “Philippine tarsier” transposase, or a variant thereof; an Anthonomus grandis transposase, herein interchangeably referred to as “DR1754440” transposase, or a variant thereof;an Anthonomus grandis transposase, herein interchangeably referred to as “DR1754053” transposase, or a variant thereof; _Solenopsis invicta transposase, herein interchangeably referred to as “DR3053925” transposase, or a variant thereof; a Simochromis diagramma transposase, herein interchangeably referred to as “Simochromis diagramma genomic” transposase, or a variant thereof; a Nematolebias whitei transposase, herein interchangeably referred to as “Nematolebias whitei chromosome 17”, or a variant thereof; a Anthonomus grandis transposase, herein interchangeably referred to as “DR1756049”, or a variant thereof; and a Coremacera marginata transposase, herein interchangeably referred to as “DR1481656”, or a variant thereof; or a nucleic acid encoding thereof.
[0074] In some embodiments, the transposase, or a variant thereof is selected from the group comprising or consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof; a Xenopus tropicalis transposase herein interchangeably referred to as “tropical clawed frog” transposase, or a variant thereof; a Japanese Medaka transposase, or a variant thereof; a Leptobrachium leishanense transposase, herein interchangeably referred to as “Leishan spiny toad” transposase, or a variant thereof; a Scalopus aquations transposase, herein interchangeably referred to as “ScaAqu- 5.3491 transposase”, or a variant thereof;an Anthonomus grandis transposase, herein interchangeably referred to as “DR1754440” transposase, or a variant thereof.
[0075] In some embodiments, the transposase, or a variant thereof is selected from the group consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof; a Xenopus tropicalis transposase herein interchangeably referred to as “tropical clawed frog” transposase, or a variant thereof; a Japanese Medaka transposase, or a variant thereof; a Leptobrachium leishanense transposase, herein interchangeably referred to as “Leishan spiny toad” transposase, or a variant thereof; a Scalopus aquaticus transposase, herein interchangeably referred to as “ScaAqu- 5.3491 transposase”, or a variant thereof; an Anthonomus grandis transposase, herein interchangeably referred to as “DR1754440” transposase, or a variant thereof.
[0076] In some embodiments, the transposase, or a variant thereof is selected from the group consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof; and an Anthonomus grandis transposase, or a variant thereof.
[0077] In some embodiments, the transposase, or a variant thereof is selected from the group consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof having an amino acid sequence having at least 75% sequence identity with the transposase; and an Anthonomus grandis transposase, or a variant thereof having an amino acid sequence having at least 75% sequence identity with the transposase.
[0078] In some embodiments, the transposase, or a variant thereof is selected from the group consisting of:Poeciliopsis turrubarensis transposase, or a variant thereof;Anthonomus grandis DR1754440 transposase, or a variant thereof;Anthonomus grandis DR1754053 transposase, or a variant thereof; andAnthonomus grandis DR1756049 transposase, or a variant thereof;
[0079] In some embodiments, the transposase is selected from the group consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 1, preferably a variant thereof having the amino acid sequence of any one of SEQ ID NO: 2 to SEQ ID NO: 13 or SEQ ID NO: 112, more preferably a variant thereof having the amino acid sequence of SEQ ID NO: 112; an Anthonomus grandis DR1754440 transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 30, preferably a variant thereof having the amino acid sequence of SEQ ID NO: 113; an Anthonomus grandis DR1756049 transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 35 or SEQ ID NO: 108, preferably a variant thereof having the amino acid sequence of SEQ ID NO: 111; and an Anthonomus grandis transposase DR1754053, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 31.
[0080] In some embodiments, the transposase, or a variant thereof is selected from the group consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof; an Anthonomus grandis transposase, herein interchangeably referred to as “DR1754440” transposase, or a variant thereof.
[0081] In some embodiments, the transposase is selected from the group consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 1, preferably a variant thereof having the amino acid sequence of any one of SEQ ID NO: 2 to SEQ ID NO: 13 or SEQ ID NO: 112, more preferably a variant thereof having the amino acid sequence of SEQ ID NO: 112; and an Anthonomus grandis DR1754440 transposase, or a variant thereof having at least 75% sequence identity with SEQ ID NO: 30, preferably a variant thereof having the amino acid sequence of SEQ ID NO: 113.
[0082] In certain embodiments, the transposase is a recombinant transposase. In certain embodiments, the transposase is a non-naturally occurring transposase.
[0083] In some embodiments, the transposase is a Poeciliopsis turrubarensis transposase, or a variant thereof.
[0084] In some embodiments, the Poeciliopsis turrubarensis transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 1.
[0085] In some embodiments, the Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 1.
[0086] In some embodiments, the Poeciliopsis turrubarensis transposase is a variant of Poeciliopsis turrubarensis transposase. In some embodiments, the variant of Poeciliopsis turrubarensis transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 1.
[0087] As used herein, “amino acid mutation” means substitution, deletion, insertion, and translocation, preferably substitution.
[0088] In some embodiments, the variant of Poeciliopsis turrubarensis transposase comprises at least one amino acid mutation, preferably at least one amino acid substitution, on one or more of the amino acids at positions 18, 22, 200, 238, 336, 342, 353, 356, 388, 413, 432, 547 and 549, corresponding to the amino acid numbering of SEQ ID NO: 1. In some embodiments, the variant of Poeciliopsis turrubarensis transposase comprises at least one amino acid mutation, preferably at least one amino acid substitution, on one or more of the amino acids at positions 18, 22, 200, 238, 336, 342, 388, 413, 547 and 549, corresponding to the amino acid numbering of SEQ ID NO: 1. In some embodiments, the variant of Poeciliopsis turrubarensis transposase comprises at least one amino acid mutation, preferably at least one amino acid substitution, on one or more of the amino acids at positions 353, 356, and 432, corresponding to the amino acid numbering of SEQ ID NO: 1.
[0089] In some embodiments, the variant of Poeciliopsis turrubarensis transposase comprises at least one amino acid substitution selected from the group comprising or consisting of W18S, V22S, T200R, I238A, I238R, R336A, Q342L, C388I, C388V, M413K, R353A, K356A, D432N, D547K, S549R, corresponding to the amino acid numbering of SEQ ID NO: 1. In some embodiments, the variant of Poeciliopsis turrubarensis transposase comprises at least one amino acid substitution selected from the group comprising or consisting of W18S, V22S, T200R, I238A, I238R, R336A,Q342L, C388I, C388V, M413K, D547K, S549R, corresponding to the amino acid numbering of SEQ ID NO: 1. In some embodiments, the variant of Poeciliopsis turrubarensis transposase comprises at least one amino acid substitution selected from the group comprising or consisting of R353A, K356A, and D432N, corresponding to the amino acid numbering of SEQ ID NO: 1. In some embodiments, the variant of Poeciliopsis turrubarensis transposase comprises at least one amino acid substitution selected from the group consisting of W18S, V22S, T200R, I238A, I238R, R336A, Q342L, C388I, C388V, M413K, R353A, K356A, D432N, D547K, S549R, corresponding to the amino acid numbering of SEQ ID NO: 1. In some embodiments, the variant of Poeciliopsis turrubarensis transposase comprises at least one amino acid substitution selected from the group consisting of W18S, V22S, T200R, I238A, I238R, R336A, Q342L, C388I, C388V, M413K, D547K, S549R, corresponding to the amino acid numbering of SEQ ID NO: 1. In some embodiments, the variant of Poeciliopsis turrubarensis transposase comprises at least one amino acid substitution selected from the group consisting of R353A, K356A, and D432N, corresponding to the amino acid numbering of SEQ ID NO: 1.
[0090] In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in any one of SEQ ID NO: 2 to SEQ ID NO: 13 or SEQ ID NO: 112. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in any one of SEQ ID NO: 2 to SEQ ID NO: 13. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 2. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 3. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 4. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 5. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 6. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 7. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQID NO: 8. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 9. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 10. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 11. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 12. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 13. In some embodiments, the variant of Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 112.
[0091] In certain embodiments, the Poeciliopsis turrubarensis transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 1.In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 1.
[0092] In some embodiments, the transposase is a Xenopus tropicalis transposase, herein interchangeably referred to as “tropical clawed frog” transposase, or a variant thereof.
[0093] In some embodiments, said tropical clawed frog transposase from Xenopus tropicalis has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 14.
[0094] In some embodiments, the Xenopus tropicalis transposase herein interchangeably referred to as “tropical clawed frog” transposase has an amino acid sequence as set forth in SEQ ID NO: 14.
[0095] In certain embodiments, the Xenopus tropicalis transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 14. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 14.
[0096] In some embodiments, the transposase is a Japanese Medaka transposase, herein interchangeably referred to as ' Sini erca chualsi" transposase or “DR0651552” transposase, or a variant thereof.
[0097] In some embodiments, the Japanese Medaka transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 15.
[0098] In some embodiments, the Japanese Medaka transposase has an amino acid sequence as set forth in SEQ ID NO: 15.
[0099] In certain embodiments, the Japanese Medaka transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 15. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 15.
[0100] In some embodiments, the transposase is a Leptobrachium leishanense transposase, herein interchangeably referred to as “Leishan spiny toad” transposase, or a variant thereof.
[0101] In some embodiments, said Leishan spiny toad transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 16.
[0102] In some embodiments, said Leishan spiny toad transposase has an amino acid sequence as set forth in SEQ ID NO: 16.
[0103] In certain embodiments, the Leishan spiny toad transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 16. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 16.
[0104] In some embodiments, the transposase is a Scalopus aquaticus transposase (or ScaAqu-5.3491 transposase), or a variant thereof.
[0105] In some embodiments, the Scalopus aquaticus transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 17.
[0106] In some embodiments, the Scalopus aquaticus transposase has an amino acid sequence as set forth in SEQ ID NO: 17.
[0107] In certain embodiments, the Scalopus aquaticus transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 17. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 17.
[0108] In some embodiments, the transposase is an Atlantic Salmon transposase, or a variant thereof.
[0109] In some embodiments, the Atlantic Salmon transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 18.
[0110] In some embodiments, the Atlantic Salmon transposase has an amino acid sequence as set forth in SEQ ID NO: 18.
[0111] In certain embodiments, the Atlantic Salmon transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 18. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 18.
[0112] In some embodiments, the transposase is a Heliconius butterfly transposase, or a variant thereof.
[0113] In some embodiments, the Heliconius butterfly transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 19.
[0114] In some embodiments, the Heliconius butterfly transposase has an amino acid sequence as set forth in SEQ ID NO: 19.
[0115] In certain embodiments, the Heliconius butterfly transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 19. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 19.
[0116] In some embodiments, the transposase is a Takifugu flavidus transposase herein interchangeably referred to as “Yellowbelly pufferfish” transposase, or a variant thereof
[0117] In some embodiments, said Yellowbelly pufferfish transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 20.
[0118] In some embodiments, said Y ellowbelly pufferfish transposase has an amino acid sequence as set forth in SEQ ID NO: 20.
[0119] In certain embodiments, the Yellowbelly pufferfish transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 20. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 20.
[0120] In some embodiments, the transposase is a Vaquita transposase, or a variant thereof.
[0121] In some embodiments, the Vaquita transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 21.
[0122] In some embodiments, the Vaquita transposase has an amino acid sequence as set forth in SEQ ID NO: 21.
[0123] In certain embodiments, the Vaquita transposase transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 21. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 21.
[0124] In some embodiments, the transposase is a Oryzias latipes transposase, (also referred to as “Japanese Rice Fish” transposase), or a variant thereof.
[0125] In some embodiments, said Japanese Rice Fish transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 22.
[0126] In some embodiments, said Japanese Rice Fish transposase has an amino acid sequence as set forth in SEQ ID NO: 22.
[0127] In certain embodiments, the Japanese Rice Fish transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 22. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 22.
[0128] In some embodiments, the transposase is a Salvelinus profundus transposase, or a variant thereof.
[0129] In some embodiments, the Salvelinus profundus transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 23.
[0130] In some embodiments, the Salvelinus profundus transposase has an amino acid sequence as set forth in SEQ ID NO: 23.
[0131] In certain embodiments, the Salvelinus profundus transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 23. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 23.
[0132] In some embodiments, the transposase is a Cuniculus paca transposase, or a variant thereof.
[0133] In some embodiments, the Cuniculus paca transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 24.
[0134] In some embodiments, the Cuniculus paca transposase has an amino acid sequence as set forth in SEQ ID NO: 24.
[0135] In certain embodiments, the Cuniculus paca transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 24. In someembodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 24.
[0136] In some embodiments, the transposase is a Noctilio leporinus transposase, or a variant thereof.
[0137] In some embodiments, the Noctilio leporinus transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 25.
[0138] In some embodiments, the Noctilio leporinus transposase has an amino acid sequence as set forth in SEQ ID NO: 25.
[0139] In certain embodiments, the Noctilio leporinus transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 25. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 25.
[0140] In some embodiments, the transposase is a Pipistrellus pipistrellus transposase (or PipPip-6.1914 transposase), or a variant thereof.
[0141] In some embodiments, the Pipistrellus pipistrellus transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 26.
[0142] In some embodiments, the Pipistrellus pipistrellus transposase has an amino acid sequence as set forth in SEQ ID NO: 26.
[0143] In certain embodiments, the Pipistrellus pipistrellus transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 26. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 26.
[0144] In some embodiments, the transposase is a Philippine tarsier transposase, or a variant thereof.
[0145] In some embodiments, the Philippine tarsier transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 27.
[0146] In some embodiments, the Philippine tarsier transposase has an amino acid sequence as set forth in SEQ ID NO: 27.
[0147] In certain embodiments, the Philippine tarsier transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 27. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 27.
[0148] In some embodiments, the transposase is a Myotis lucifugus transposase, herein interchangeably referred to as “PiggyBat” transposase, or a variant thereof.
[0149] In some embodiments, said “PiggyBat” transposase from Myotis lucifugus has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 28.
[0150] In some embodiments, said “PiggyBat” transposase from Myotis lucifugus has an amino acid sequence as set forth in SEQ ID NO: 28.
[0151] In certain embodiments, the PiggyBat transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 28. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 28.
[0152] In some embodiments, the transposase is a PiggyBac transposase, or a variant thereof.
[0153] In some embodiments, the PiggyBac transposase has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 29.
[0154] In some embodiments, said PiggyBac transposase has an amino acid sequence as set forth in SEQ ID NO: 29.
[0155] In certain embodiments, the PiggyBac transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 29. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 29.
[0156] In some embodiments, the transposase is a Anthonomus grandis transposase, herein interchangeably referred to as “DR1754440” transposase or “Antgra4440”, or a variant thereof. In some embodiments, the transposase is Anthonomus grandis DR1754440 transposase.
[0157] In some embodiments, said DR1754440 transposase from Anthonomus grandis has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 30.
[0158] In some embodiments, said DR1754440 transposase from Anthonomus grandis has an amino acid sequence as set forth in SEQ ID NO: 30.
[0159] In certain embodiments, the Anthonomus grandis DR1754440 transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 30. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 30.
[0160] In some embodiments, the Anthonomus grandis DR1754440 transposase is a variant of Anthonomus grandis DR1754440 transposase. In some embodiments, the variant of Anthonomus grandis DR1754440 transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 30.
[0161] In some embodiments, the variant of Anthonomus grandis DR1754440 transposase comprises at least one amino acid mutation, preferably at least one amino acid substitution, on one or more of the amino acids at positions 388, 389, and 393, corresponding to the amino acid numbering of SEQ ID NO: 30.
[0162] In some embodiments, the variant of Anthonomus grandis DR1754440 transposase comprises at least one amino acid substitution selected from the group comprising or consisting of R388A, K389A, and K393A, corresponding to the amino acid numbering of SEQ ID NO: 30. In some embodiments, the variant of Anthonomus grandis DR1754440 transposase comprises at least one amino acid substitution selected from the group consisting of R388A, K389A, and K393A, corresponding to the amino acid numbering of SEQ ID NO: 30.
[0163] In some embodiments, the variant of Anthonomus grandis DR1754440 transposase has an amino acid sequence as set forth in SEQ ID NO: 113.
[0164] In some embodiments, the transposase is a Anthonomus grandis transposase, herein interchangeably referred to as “DR1754053” transposase or “Antgra4053”, or a variant thereof. In some embodiments, the transposase is Anthonomus grandis DR1754053 transposase.
[0165] In some embodiments, said DR1754053 transposase from Anthonomus grandis has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 31.
[0166] In some embodiments, said DR1754053 transposase from Anthonomus grandis has an amino acid sequence as set forth in SEQ ID NO: 31.
[0167] In certain embodiments, the Anthonomus grandis DR1754053 transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 31. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 31.
[0168] In some embodiments, the transposase is a Solenopsis invicta transposase, herein interchangeably referred to as “DR3053925” transposase, or a variant thereof.
[0169] In some embodiments, said DR3053925 transposase from Solenopsis invicta has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 32.
[0170] In some embodiments, said DR3053925 transposase from Solenopsis invicta has an amino acid sequence as set forth in SEQ ID NO: 32.
[0171] In certain embodiments, the Solenopsis invicta transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 32. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 32.
[0172] In some embodiments, the transposase is a Simochromis diagramma transposase, herein interchangeably referred to as “Simochromis diagramma genomic” transposase, or a variant thereof.
[0173] In some embodiments, said Simochromis diagramma genomic transposase from Simochromis diagramma has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 33.
[0174] In some embodiments, said Simochromis diagramma genomic transposase from Simochromis diagramma has an amino acid sequence as set forth in SEQ ID NO: 33.
[0175] In certain embodiments, the Simochromis diagramma transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO:33. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 33.
[0176] In some embodiments, the transposase is a Nematolebias whitei transposase, herein interchangeably referred to as “Nematolebias whitei chromosome 17” transposase, or a variant thereof.
[0177] In some embodiments, said “Nematolebias whitei chromosome 17” transposase from Nematolebias whitei has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO:34.
[0178] In some embodiments, said “Nematolebias whitei chromosome 17” transposase from Nematolebias whitei has an amino acid sequence as set forth in SEQ ID NO: 34.
[0179] In certain embodiments, the Nematolebias whitei transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 34. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 34.
[0180] In some embodiments, the transposase is a Anthonomus grandis transposase, herein interchangeably referred to as “DR1756049” transposase or “Antgra6049”, or a variant thereof. In some embodiments, the transposase is Anthonomus grandis DR1756049 transposase.
[0181] In some embodiments, said “DR1756049” transposase tmwi Anthonomus grandis has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 108 or SEQ ID NO: 35.
[0182] In some embodiments, said “DR1756049” transposase from Anthonomus grandis has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 108.
[0183] In some embodiments, said “DR1756049” transposase from Anthonomus grandis has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 35.
[0184] In some embodiments, said “DR1756049” transposase from Anthonomus grandis has an amino acid sequence as set forth in SEQ ID NO: 35 or SEQ ID NO: 108. In some embodiments, said “DR1756049” transposase from Anthonomus grandis has an amino acid sequence as set forth in SEQ ID NO: 108. In some embodiments, said “DR1756049” transposase from Anthonomus grandis has an amino acid sequence as set forth in SEQ ID NO: 35.
[0185] In certain embodiments, the Anthonomus grandis DR1756049 transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO:35 or 108. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 35 or 108.
[0186] In some embodiments, the Anthonomus grandis DR1756049 transposase is a variant of Anthonomus grandis DR1756049 transposase. In some embodiments, the variant of Anthonomus grandis DR1756049 transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 35.
[0187] In some embodiments, the variant of Anthonomus grandis DR1756049 transposase comprises at least one amino acid mutation, preferably at least one amino acid substitution, on one or more of the amino acids at positions 348, 352, and 429, corresponding to the amino acid numbering of SEQ ID NO: 35.
[0188] In some embodiments, the variant of Anthonomus grandis DR1756049 transposase comprises at least one amino acid substitution selected from the group consisting of R348A, K352A, and D429N, corresponding to the amino acid numbering of SEQ ID NO: 35.
[0189] In some embodiments, the variant of Anthonomus grandis DR1756049 transposase has an amino acid sequence as set forth in SEQ ID NO: 111.
[0190] In some embodiments, the Anthonomus grandis transposase has an amino acid sequence selected from the group comprising or consisting of SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 108 and SEQ ID NO: 35. In some embodiments, the Anthonomus grandis transposase has an amino acid sequence selected from the group comprising or consisting of SEQ ID NO: 30, SEQ ID NO: 31 and SEQ ID NO: 108. In some embodiments, the Anthonomus grandis transposase has an amino acid sequence selected from the group comprising or consisting of SEQ ID NO: 30, SEQ ID NO: 31 and SEQ ID NO: 35. In some embodiments, the Anthonomus grandis transposase has an amino acid sequence selected from the group consisting of SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 108 and SEQ ID NO: 35. In some embodiments, the Anthonomus grandis transposase has an amino acid sequence selected from the group consisting of SEQ ID NO: 30, SEQ ID NO: 31 and SEQ ID NO: 108. In some embodiments, the Anthonomus grandis transposase has an amino acid sequence selected from the group consisting of SEQ ID NO: 30, SEQ ID NO: 31 and SEQ ID NO: 35.
[0191] In some embodiments, the Anthonomus grandis transposase variant has the amino acid sequence of SEQ ID NO: 111 or SEQ ID NO: 113.
[0192] In some embodiments, the transposase is a Coremacera marginata transposase, herein interchangeably referred to as “DR1481656”, or a variant thereof.
[0193] In some embodiments, the transposase is a Coremacera marginata transposase, herein interchangeably referred to as “DR1481656” transposase, or a variant thereof.
[0194] In some embodiments, said “DR1481656” transposase from Coremacera marginata has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 36.
[0195] In some embodiments, said “DR1481656” transposase from Coremacera marginata has an amino acid sequence as set forth in SEQ ID NO: 36.
[0196] In certain embodiments, the Coremacera marginata transposase is a recombinant transposase. In some embodiments, the recombinant transposase comprises at least one amino acid mutation compared to the amino acid sequence of SEQ ID NO: 36. In some embodiments, the recombinant transposase has less than 100% sequence identity with SEQ ID NO: 36.
[0197] In some embodiments, the transposase is a mutant transposase. In some embodiments, the mutant transposase comprises at least one amino acid mutation compared to the amino acid sequence of any one of SEQ ID NO: 1 or SEQ ID NO: 14 to SEQ ID NO: 36. In some embodiments, the mutant transposase comprises at least one amino acid mutation compared to the amino acid sequence of any one of SEQ ID NO: 14 to SEQ ID NO: 36.
[0198] In some embodiments, the at least one mutation is selected from the group comprising or consisting of substitution, deletion, insertion, and translocation.
[0199] In some embodiments, the mutant transposase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more mutation(s) compared to the amino acid sequence of any one of SEQ ID NO: 1 or SEQ ID NO: 14 to SEQ ID NO: 36. In some embodiments, the mutant transposase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more mutation(s) compared to the amino acid sequence of any one of SEQ ID NO: 14 to SEQ ID NO: 36.
[0200] In some embodiments, the mutant transposase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid substitution(s) compared to the amino acid sequence of any one of SEQ ID NO: 1 or SEQ ID NO: 14 to SEQ ID NO: 36. In some embodiments, the mutant transposase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more substitution(s) compared to the amino acid sequence of any one of SEQ ID NO: 14 to SEQ ID NO: 36.
[0201] In some embodiments, the mutant transposase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid deletion(s) compared to the amino acid sequence of any one of SEQ ID NO: 1 or SEQ ID NO: 14 to SEQ ID NO: 36. In some embodiments, the mutant transposase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more deletion(s) compared to the amino acid sequence of any one of SEQ ID NO: 14 to SEQ ID NO: 36.
[0202] In some embodiments, the mutant transposase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid insertion(s) compared to the amino acid sequence of any one of SEQ ID NO: 1 or SEQ ID NO: 14 to SEQ ID NO: 36. In some embodiments, the mutant transposase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more insertion(s) compared to the amino acid sequence of any one of SEQ ID NO: 14 to SEQ ID NO: 36.
[0203] In some embodiments, the mutant transposase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid translocation(s) compared to the amino acid sequence of any one of SEQ ID NO: 1 or SEQ ID NO: 14 to SEQ ID NO: 36. In some embodiments, the mutant transposase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more translocation(s) compared to the amino acid sequence of any one of SEQ ID NO: 14 to SEQ ID NO: 36.
[0204] Mutations of the hyperactive PiggyBac transposase (hyPB) have been described in W02020250181 and WO2022129438; residues R372, K375, and D450 of hyPB are critical for increasing on-target integration of this transposase when associated with a RNA-guided nuclease. Thus, in certain embodiments, the mutant transposase comprises at least one, preferably at least 2, more preferably 3 amino acid substitutions at positions corresponding to positions 372, 375, and / or 450 of hyPB. Within the scope of the present invention, such mutant may be referred to as “triple mutants” or “X3 mutants”.
[0205] In some embodiments, the transposase is a Poeciliopsis turrubarensis transposase of SEQ ID NO: 1 or Anthonomus grandis transposase of SEQ ID NO: 30,SEQ ID NO: 31, SEQ ID NO: 35 or SEQ ID NO: 108. In some embodiments, the transposase is a Poeciliopsis turrubarensis transposase variant of SEQ ID NO: 112 or a Anthonomus grandis transposase variant of SEQ ID NO: 111 or SEQ ID NO: 113.
[0206] In some embodiments, the transposase is a Poeciliopsis turrubarensis transposase of SEQ ID NO: 1 or a Poeciliopsis turrubarensis transposase variant of SEQ ID NO: 112.
[0207] In some embodiments, the transposase is a Anthonomus grandis transposase of SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 35 or SEQ ID NO: 108, or a Anthonomus grandis transposase variant of SEQ ID NO: 111 or SEQ ID NO: 113.
[0208] Transposases have the ability to recognize and bind specific nucleic acid sequences named Inverted Terminal Repeats (ITRs). The ITRs typically flank both sites of a nucleic acid sequence that is “cut-and-paste” by the transposase.
[0209] Thus, in some embodiments, the transposase recognizes and / or binds to at least one ITR sequence, preferably at least two ITR sequences. In some embodiments, the transposase recognizes and / or binds to a left ITR and a right ITR. As used herein, “left ITR” refers to the ITR sequence flanking the 5'-P extremity of the nucleic acid sequence flanked by the ITRs; and “right ITR” refers to the ITR sequence flanking the 3'-OH extremity of the nucleic acid sequence flanked by the ITRs.
[0210] In some embodiments, the transposase recognizes at least one ITR sequence, preferably at least two ITR sequences, selected from the group comprising or consisting of SEQ ID NO: 37 to SEQ ID NO: 84. In some embodiments, the transposase recognizes at least one ITR sequence, preferably at least two ITR sequences, selected from the group consisting of SEQ ID NO: 37 to SEQ ID NO: 84.
[0211] In some embodiments, the Poeciliopsis turrubarensis transposase recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 37, and a right ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 38.
[0212] In some embodiments, the Xenopus tropicalis transposase (also referred to as “tropical clawed frog” transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 39, and a right ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 40.
[0213] In some embodiments, the Japanese Medaka transposase (also referred to as Siniperca chuatsi transposase, or DR0651552 transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 41, and a right ITR sequence having at least75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 42.
[0214] In some embodiments, the Leptobrachium leishanense transposase (also referred to as “Leishan spiny toad” transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 43, and a right ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 44.
[0215] In some embodiments, the Scalopus aquaticus transposase recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 45, and a right ITR sequence having at least75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 46.
[0216] In some embodiments, the Atlantic Salmon transposase recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 47, and a right ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 48.
[0217] In some embodiments, the Heliconius butterfly transposase recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 49, and a right ITR sequence having at least75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 50.
[0218] In some embodiments, the Yellowbelly pufferfish transposase recognizes a left UR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 51, and a right ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 52.
[0219] In some embodiments, the Vaquita transposase recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 53, and a right ITR sequence having at least 75%,80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity withSEQ ID NO: 54.
[0220] In some embodiments, the Oryzias latipes transposase (also referred to as “Japanese Rice Fish” transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 55, and a right ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 56.
[0221] In some embodiments, the Salvenius profundus transposase recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 57, and a right ITR sequence having at least75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 58.
[0222] In some embodiments, the Cuniculus paca transposase recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 59, and a right ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 60.
[0223] In some embodiments, the Noctilio leporinus transposase recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 61, and a right ITR sequence having at least75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 62.
[0224] In some embodiments, the Pipistrellus pipistrellus transposase recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 63, and a right ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 64.
[0225] In some embodiments, the Carlito syrichta transposase (also referred to as “Philippine tarsier” transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity withSEQ ID NO: 65, and a right ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 66.
[0226] In some embodiments, the Myotis lucifugus transposase (also referred to as “Piggybat” transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 67, and a right ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 68.
[0227] In some embodiments, the transposase referred to as “PiggyBac” transposase recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 69, and a right ITRsequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 70.
[0228] In some embodiments, the Anthonomus grandis transposase (referred to as “DR1754440” transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 71, and aright ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 72.
[0229] In some embodiments, the Anthonomus grandis transposase (referred to as “DR1754053” transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 73, and aright ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 74.
[0230] In some embodiments, the Solenopsis invicta transposase (referred to as “DR3053925” transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 75, and aright ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 76.
[0231] In some embodiments, the Simochromis diagramma transposase (referred to as “Simochromis diagramma genomic” transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 77, and a right ITR sequence having at least 75%, 80%, 85%,90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO:78.
[0232] In some embodiments, the Nematolebias whitei transposase (referred to as “Nematolebias whitei chromosome 17” transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 79, and a right ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 80.
[0233] In some embodiments, the Anthonomus grandis transposase (referred to as “DR1756049” transposase) recognizes a left ITR sequence having at least 75%, 80%,85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 81, and aright ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 82.
[0234] In some embodiments, the Coremacera marginata transposase (referred to as “DR1481656” transposase) recognizes a left ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 83, and aright ITR sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 84.
[0235] In some embodiments, the transposase is an Anthonomus grandis transposase, or a Poeciliopsis turrubarensis transposase. In some embodiments, the transposase is anAnthonomus grandis transposase, a Poeciliopsis turrubarensis transposase, or a variant thereof.
[0236] In some embodiments, the transposase is selected from the group comprising or consisting of Anthonomus grandis transposase DR1756049, interchangeably referred to as Antgra6049, Anthonomus grandis transposase DR1754440, interchangeably referred to as Antgra4440, Anthonomus grandis transposase DR1754053, interchangeably referred to as Antgra4053, a Poeciliopsis turrubarensis transposase, or variants thereof. In some embodiments, the transposase is selected from the group consisting of Anthonomus grandis DR1756049 transposase, Anthonomus grandis DR1754440 transposase, Anthonomus grandis DR1754053 transposase, and Poeciliopsis turrubarensis transposase, or variants thereof.
[0237] In some embodiments, the transposase is selected from the group comprising or consisting of DR1756049 transposase (Antgra6049) having at least 75% sequence identity with SEQ ID NO: 108, DR1754440 transposase (Antgra4440) having at least 75% sequence identity with SEQ ID NO: 30, DR1754053 transposase (Antgra4053) having at least 75% sequence identity with SEQ ID NO: 31, or a Poeciliopsis turrubarensis transposase having at least 75% sequence identity with any one of SEQ ID NO: 1 to SEQ ID NO: 13.
[0238] In some embodiments, the transposase is selected from the group comprising or consisting of DR1756049 transposase (Antgra6049) having at least 80% sequence identity with SEQ ID NO: 108, DR1754440 transposase (Antgra4440) having at least 80% sequence identity with SEQ ID NO: 30, DR1754053 transposase (Antgra4053) having at least 80% sequence identity with SEQ ID NO: 31, or a Poeciliopsis turrubarensis transposase having at least 80% sequence identity with any one of SEQ ID NO: 1 to SEQ ID NO: 13.
[0239] In some embodiments, the transposase is selected from the group comprising or consisting of DR1756049 transposase (Antgra6049) having at least 85% sequence identity with SEQ ID NO: 108, DR1754440 transposase (Antgra4440) having at least 85% sequence identity with SEQ ID NO: 30, DR1754053 transposase (Antgra4053)having at least 85% sequence identity with SEQ ID NO: 31, or a Poeciliopsis turrubarensis transposase having at least 85% sequence identity with any one of SEQ ID NO: I to SEQ ID NO: 13.
[0240] In some embodiments, the transposase is selected from the group comprising or consisting of DR1756049 transposase (Antgra6049) having at least 90% sequence identity with SEQ ID NO: 108, DR1754440 transposase (Antgra4440) having at least 90% sequence identity with SEQ ID NO: 30, DR1754053 transposase (Antgra4053) having at least 90% sequence identity with SEQ ID NO: 31, or a Poeciliopsis turrubarensis transposase having at least 90% sequence identity with any one of SEQ ID NO: I to SEQ ID NO: 13.
[0241] In some embodiments, the transposase is selected from the group comprising or consisting of DR1756049 transposase (Antgra6049) having at least 95% sequence identity with SEQ ID NO: 108, DR1754440 transposase (Antgra4440) having at least 95% sequence identity with SEQ ID NO: 30, DR1754053 transposase (Antgra4053) having at least 95% sequence identity with SEQ ID NO: 31, or a Poeciliopsis turrubarensis transposase having at least 95% sequence identity with any one of SEQ ID NO: I to SEQ ID NO: 13.
[0242] In some embodiments, the transposase is selected from the group comprising or consisting of DR1756049 transposase (Antgra6049) having 100% sequence identity with SEQ ID NO: 108, DR1754440 transposase (Antgra4440) having 100% sequence identity with SEQ ID NO: 30, DR1754053 transposase (Antgra4053) having 100% sequence identity with SEQ ID NO: 31, or a Poeciliopsis turrubarensis transposase having 100% sequence identity with any one of SEQ ID NO: 1 to SEQ ID NO: 13.
[0243] In some embodiments, the composition of the invention comprises:- an Anthonomus grandis DR1756049 transposase (Antgra6049) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 108, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR)having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 81 and a right ITR (z.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 82.
[0244] In some embodiments, the composition of the invention consists of:- an Anthonomus grandis DR1756049 transposase (Antgra6049) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 108 or SEQ ID NO: 35, preferably an Anthonomus grandis DR1756049 transposase variant of amino acid sequence SEQ ID NO: 111, and- at least one nucleic acid molecule comprising at least one transgene of interest, wherein the at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 81 and a right ITR (z.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 82.
[0245] In some embodiments, the composition of the invention comprises:- an Anthonomus grandis DR1754440 transposase (Antgra4440) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 30, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 71 and a right ITR (z.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 72.
[0246] In some embodiments, the composition of the invention consists of:- an Anthonomus grandis DR1754440 transposase (Antgra4440) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 30, preferably an Anthonomus grandis DR1754440 transposase variant of amino acid sequence SEQ ID NO: 113, and- at least one nucleic acid molecule comprising at least one transgene of interest, wherein the at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 71 and a right ITR (i.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 72.
[0247] In some embodiments, the composition of the invention comprises:- an Anthonomus grandis DR1754053 transposase (Antgra4053) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 31, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR (i.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 73 and a right ITR i.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 74.
[0248] In some embodiments, the composition of the invention consists of:- an Anthonomus grandis DR1754053 transposase (Antgra4053) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 31, and- at least one nucleic acid molecule comprising at least one transgene of interest, wherein the at least one transgene of interest is flanked with a left ITR (i.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 73 and a right ITR (i.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 74.
[0249] In some embodiments, the composition of the invention comprises:- a Poeciliopsis turrubarensis transposase having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with any one of SEQ ID NO: 1 to SEQ ID NO: 13, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR (i.e., 5’-P ITR)having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 37 and a right ITR (z.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 38.
[0250] In some embodiments, the composition of the invention consists of:- a Poeciliopsis turrubarensis transposase having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with any one of SEQ ID NO: 1 to SEQ ID NO: 13 or SEQ ID NO: 112, preferably SEQ ID NO: 112, and- at least one nucleic acid molecule comprising at least one transgene of interest, wherein the at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 37 and a right ITR (i.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 38.
[0251] In some embodiments, the composition of the invention comprises:- a Poeciliopsis turrubarensis transposase having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with any one of SEQ ID NO: 2 to SEQ ID NO: 13, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR (i.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 37 and a right ITR i.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 38.
[0252] In some embodiments, the composition of the invention consists of:- a Poeciliopsis turrubarensis transposase having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with any one of SEQ ID NO: 2 to SEQ ID NO: 13 or SEQ ID NO: 112, preferably SEQ ID NO: 112, and- at least one nucleic acid molecule comprising at least one transgene of interest, wherein the at least one transgene of interest is flanked with a left ITR (i.e., 5’-P ITR) having atleast 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 37 and a right ITR (i.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 38.
[0253] In some embodiments, the composition of the invention further comprises a sitespecific DNA binding protein. In some embodiments, the site-specific DNA binding protein is a RNA-guided nuclease or nickase.
[0254] In some embodiments, the composition of the invention further comprises a RNA-guided nuclease or nickase.
[0255] In some embodiments, the RNA-guided nuclease or nickase comprises an active DNA cleavage domain and a guide RNA binding domain.
[0256] In some embodiments, the RNA-guided nuclease or nickase is a Cas protein.
[0257] In some embodiments, the Cas protein is selected from the group comprising or consisting of Cas9, Cas 12a (Cpfl), Cas 12b, Casl2f, and CasX. In some embodiments, the Cas protein is selected from the group consisting of Cas9, Casl2a (Cpfl), Casl2b, Casl2f, and CasX. It shall be understood that variants and functional fragments thereof are also encompassed, such as nickase Cas (nCas) or dead Cas (dCas) variants.
[0258] In some embodiments, the composition of the invention further comprises at least one RNA-guided nuclease or nickase selected from the group comprising or consisting of:Cas9 protein from Streptococcus pyogenes (SpCas9);Cas9 protein from Staphylococcus aureus (SaCas9);Cas9 protein from Campylobacter jejuni (CjCas9);Cas9 protein from Corynebacterium ulcerans;Cas9 protein from Corynebacterium diphtheria;Cas9 protein from Spiroplasma syrphidicola;Cas9 protein from Prevotella intermedia;Cas9 protein from Spiroplasma taiwanense;Cas9 protein from Streptococcus iniae;Cas9 protein from Belliella baltica;Cas9 protein from Psychroflexus torquisi;Cas9 protein from Streptococcus thermophilus;Cas9 protein from Listeria innocua;Cas9 protein from Neisseria meningitidis;Cas9 nickase from Streptococcus pyogenes Cas9 (nCas9);Cas9 nickase from Staphylococcus aureus Cas9 (SanCas9);Dead_Cas9;- Casl2a (Cpfl);- UnlCasl2fl- CasX;Dra2_TnpB;“Ancestral” Cas, refered to as LFCA; and“Ancestral” Cas refered to as LBCA.
[0259] Examples of Cas9 nucleases include, without limitation, Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus haemolyticus Cas9 (ShCas9), and Campylobacter jejuni Cas9 (CjCas9).
[0260] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein.
[0261] In some embodiments, the Cas9 protein may be a “Cas9 variant”. A “Cas9 variant”, as used herein, is a protein sharing homology to a Cas9 protein as described herein, and includes fragments thereof.
[0262] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Streptococcus pyogenes (SpCas9), or a variant thereof.
[0263] In some embodiments, the Cas9 protein from Streptococcus pyogenes (SpCas9), or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 85.
[0264] In some embodiments, the Cas9 protein from Streptococcus pyogenes (SpCas9) has an amino acid sequence as set forth in SEQ ID NO: 85.
[0265] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Staphylococcus aureus (SaCas9), or a variant thereof.
[0266] In some embodiments, the Cas9 protein from Staphylococcus aureus (SaCas9), or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 86.
[0267] In some embodiments, the Cas9 protein from Staphylococcus aureus (SaCas9) has an amino acid sequence as set forth in SEQ ID NO: 86.
[0268] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Campylobacter jejuni (CjCas9), or a variant thereof.
[0269] In some embodiments, the Cas9 protein from Campylobacter jejuni (CjCas9), or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 87.
[0270] In some embodiments, the Cas9 protein from Campylobacter jejuni (CjCas9) has an amino acid sequence as set forth in SEQ ID NO: 87.
[0271] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Corynebacterium ulcerans, or a variant thereof.
[0272] In some embodiments, the Cas9 protein from Corynebacterium ulcerans, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 88.
[0273] In some embodiments, the Cas9 protein from Corynebacterium ulcerans has an amino acid sequence as set forth in SEQ ID NO: 88.
[0274] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Corynebacterium diphtheria, or a variant thereof.
[0275] In some embodiments, the Cas9 protein from Corynebacterium diphtheria, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 89.
[0276] In some embodiments, the Cas9 protein from Corynebacterium diphtheria has an amino acid sequence as set forth in SEQ ID NO: 89.
[0277] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Spiroplasma syrphidicola, or a variant thereof.
[0278] In some embodiments, the Cas9 protein from Spiroplasma syrphidicola, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 90.
[0279] In some embodiments, the Cas9 protein from Spiroplasma syrphidicola has an amino acid sequence as set forth in SEQ ID NO: 90.
[0280] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Prevotella intermedia, or a variant thereof.
[0281] In some embodiments, the Cas9 protein from Prevotella intermedia, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 91.
[0282] In some embodiments, the Cas9 protein from Prevotella intermedia has an amino acid sequence as set forth in SEQ ID NO: 91.
[0283] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Spiroplasma taiwanense, or a variant thereof.
[0284] In some embodiments, the Cas9 protein from Spiroplasma taiwanense, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 92.
[0285] In some embodiments, the Cas9 protein from Spiroplasma taiwanense has an amino acid sequence as set forth in SEQ ID NO: 92.
[0286] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Streptococcus iniae, or a variant thereof.
[0287] In some embodiments, the Cas9 protein from Streptococcus iniae, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 93.
[0288] In some embodiments, the Cas9 protein from Streptococcus iniae has an amino acid sequence as set forth in SEQ ID NO: 93.
[0289] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Belliella baltica, or a variant thereof.
[0290] In some embodiments, the Cas9 protein from Belliella baltica, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 94.
[0291] In some embodiments, the Cas9 protein from Belliella baltica has an amino acid sequence as set forth in SEQ ID NO: 94.
[0292] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Psychroflexus lorquis or a variant thereof.
[0293] In some embodiments, the Cas9 protein from Psychroflexus lorquisL or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 95.
[0294] In some embodiments, the Cas9 protein from Psychroflexus torquisi has an amino acid sequence as set forth in SEQ ID NO: 95.
[0295] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Streptococcus thermophilus, or a variant thereof.
[0296] In some embodiments, the Cas9 protein from Streptococcus thermophilus, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 96.
[0297] In some embodiments, the Cas9 protein from Streptococcus thermophilus has an amino acid sequence as set forth in SEQ ID NO: 96.
[0298] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Listeria innocua, or a variant thereof.
[0299] In some embodiments, the Cas9 protein from Listeria innocua, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 97.
[0300] In some embodiments, the Cas9 protein from Listeria innocua has an amino acid sequence as set forth in SEQ ID NO: 97.
[0301] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 protein from Neisseria meningitidis, or a variant thereof.
[0302] In some embodiments, the Cas9 protein from Neisseria meningitidis, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 98.
[0303] In some embodiments, the Cas9 protein from Neisseria meningitidis has an amino acid sequence as set forth in SEQ ID NO: 98.
[0304] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 nickase from Streptococcus pyogenes (nCas9), or a variant thereof.
[0305] In some embodiments, the Cas9 nickase from Streptococcus pyogenes (nCas9), or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 99.
[0306] In some embodiments, the Cas9 nickase from Streptococcus pyogenes (nCas9) has an amino acid sequence as set forth in SEQ ID NO: 99.
[0307] In some embodiments, the RNA-guided nuclease or nickase is a dead Cas9 (dCas9), or a variant thereof.
[0308] In some embodiments, the dead Cas9, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 100.
[0309] In some embodiments, the dead Cas9 has an amino acid sequence as set forth in SEQ ID NO: 100.
[0310] In some embodiments, the RNA-guided nuclease or nickase is a Casl2a (Cpfl), or a variant thereof.
[0311] In some embodiments, the Casl2a (Cpfl), or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 101.
[0312] In some embodiments, the Cas 12a (Cpf 1) has an amino acid sequence as set forth in SEQ ID NO: 101.
[0313] In some embodiments, the RNA-guided nuclease or nickase is a UnlCasl2fl, or a variant thereof.
[0314] In some embodiments, the UnlCasl2fl, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 102.
[0315] In some embodiments, the UnlCasl2fl has an amino acid sequence as set forth in SEQ ID NO: 102.
[0316] In some embodiments, the RNA-guided nuclease or nickase is a CasX, or a variant thereof.
[0317] In some embodiments, the CasX, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 103.
[0318] In some embodiments, the CasX has an amino acid sequence as set forth in SEQ ID NO: 103.
[0319] In some embodiments, the RNA-guided nuclease or nickase is Dra2_TnpB, or a variant thereof.
[0320] In some embodiments, the Dra2_TnpB, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 104.
[0321] In some embodiments, the Dra2_TnpB has an amino acid sequence as set forth in SEQ ID NO: 104.
[0322] In some embodiments, the RNA-guided nuclease or nickase is an “Ancestral” Cas, refered to as LFCA, or a variant thereof.
[0323] In some embodiments, the LFCA, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 105.
[0324] In some embodiments, the LFCA has an amino acid sequence as set forth in SEQ ID NO: 105.
[0325] In some embodiments, the RNA-guided nuclease or nickase is an “Ancestral” Cas, referred to as LBCA, or a variant thereof.
[0326] In some embodiments, the LBCA, or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 106.
[0327] In some embodiments, the LBCA has an amino acid sequence as set forth in SEQ ID NO: 106.
[0328] In some embodiments, the RNA-guided nuclease or nickase is a Cas9 nickase from Staphylococcus aureus (SanCas9) or a variant thereof.
[0329] In some embodiments, the Cas9 nickase from Staphylococcus aureus (SanCas9), or the variant thereof has an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity with SEQ ID NO: 107.
[0330] In some embodiments, the Cas9 nickase from Staphylococcus aureus (SanCas9) has an amino acid sequence as set forth in SEQ ID NO: 107.
[0331] In some embodiments, the RNA-guided nuclease or nickase has an amino acid sequence selected from the group comprising or consisting of any one of SEQ ID NO: 85 to SEQ ID NO: 107. In some embodiments, the RNA-guided nuclease or nickase has an amino acid sequence selected from the group consisting of any one of SEQ ID NO: 85 to SEQ ID NO: 107.
[0332] The transposase and the RNA-guided nuclease or nickase may be administered separately (z.e., decoupled), or, alternatively, fused or otherwise linked together.
[0333] In one embodiment, the transposase and the RNA-guided nuclease are not fused nor linked together. In some embodiments, the transposase and the RNA-guided nuclease are decoupled or split. In some embodiments, the transposase and the RNA-guided nuclease are to be administered separately (e.g., when contacting cells with the composition of the invention). Illustratively, the transposase and the RNA-guided nuclease may be encoded by distinct vectors.
[0334] In another embodiment, the transposase and the RNA-guided nuclease are associated together, fused, or otherwise linked. Methods to associate two proteins, in particular methods to design fusion proteins, are well known in the art.
[0335] In some embodiments, the transposase and the RNA-guided nuclease or nickase are fused together in a fusion protein. In some embodiments, the RNA-guided nuclease or nickase is fused in C-terminus or N-terminus of the transposase. In certain embodiments, the transposase and the RNA-guided nuclease or nickase are covalently or non-covalently linked. In one embodiment, the transposase and the RNA-guided nuclease or nickase are covalently linked. In another embodiment, the transposase and the RNA- guided nuclease or nickase are non-covalently linked.
[0336] In some embodiments, the fusion protein further comprises at least one linker.
[0337] In some embodiments, the composition further comprises a gRNA. In some embodiments, the gRNA is recognized by the RNA-guided nuclease or nickase. Typically, the gRNA is complementary and / or specific of at least one locus in the genome of a cell, thereby forcing the localization of the RNA-guided nuclease or nickase to this specific locus.
[0338] In some embodiments, the transposase is fused to an aptamer binding protein, and the gRNA of the RNA-guided nuclease or nickase comprises at least one aptamer sequence. The aptamer binding protein may be fused in C-terminus or N-terminus of the transposase, optionally through a linker. The at least one aptamer sequence may be DNAor RNA, preferably RNA. The gRNA may comprise more than one aptamer sequence, i.e., 2, 3, 4, 5, 6, 7, 8, 9, or more.
[0339] In some embodiments, the aptamer binding protein is MS2 bacteriophage coat protein (MCP) and the at least one aptamer is a MS2 RNA tetraloop binding sequence. In some embodiments, MCP has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity with SEQ ID NO: 109 (encoded, e.g., by the nucleic acid sequence with SEQ ID NO: 110).
[0340] Examples of covalent and non-covalent association of transposases with RNA- guided nuclease or nickase and gRNA are disclosed in W02020250181 and WO2022129438, the entire content of which is incorporated herein by reference. It will be understood that the MCP in the transposase MCP-fusion protein binds non-covalently to the at least one MS2 RNA tetraloop binding sequence comprised in the gRNA itself non-covalently bound to a RNA-guided nuclease or nickase; in particular, the binding of the fusion protein to the RNA-guided nuclease or nickase / gRNA complex directs the activity of the transposase of the invention towards the site specifically recognized by the RNA-guided nuclease or nickase / gRNA complex.
[0341] The present invention further relates to the use of the composition according to the invention for gene editing.
[0342] In some embodiments, gene editing includes gene disruption and gene integration, preferably gene integration. Gene integration may be referred to as gene insertion.
[0343] In one embodiment, gene editing includes targeted gene disruption and targeted gene integration, preferably targeted gene integration. In another embodiment, gene editing includes non-targeted gene disruption and non-targeted gene integration, preferably non-targeted gene integration.
[0344] It will be apparent to the person skilled in the art which embodiment of the present invention enables targeted gene integration or disruption, or non-targeted gene integration or disruption. Typically, covalent or non-covalent association of thetransposase with the RNA-guided nuclease or nickase and the gRNA enables targeted gene integration or disruption, as the RNA-guided nuclease or nickase and the gRNA increases the targeting of the transposase activity to the locus recognized by the gRNA. Alternatively, non-targeted gene integration or disruption does not require the presence of a RNA-guided nuclease or nickase nor gRNA.
[0345] The present invention further relates to an in vitro method for the modifying the genome of one or more cell, comprising contacting said one or more cell with the composition according to the invention. As used herein, “one or more cell” may refer to a population of cells.
[0346] The present invention further relates to an in vitro method for the integration of at least one transgene of interest into the genome of one or more cell, comprising contacting said one or more cell with the composition according to the invention, and at least one nucleic acid molecule comprising said at least one transgene of interest. As used herein, “one or more cell” may refer to a population of cells.
[0347] In a preferred embodiment, the integration is targeted integration, z. e. , integration of the at least one transgene of interest into a specific locus of the genome. In this embodiment, the composition according to the invention preferably comprises a RNA- guided nuclease or nickase and a gRNA. In another embodiment, the integration is nontargeted or non-specific integration.
[0348] In some embodiments, the at least one nucleic acid molecule comprising the at least one transgene of interest is a DNA or RNA molecule. In a preferred embodiment, the at least one nucleic acid molecule comprising the at least one transgene of interest is a DNA molecule.
[0349] In some embodiments, the method is for the integration of large nucleic acid sequence, preferably the integration of large transgenes, more preferably the targeted integration of large transgenes. Hence, in some embodiments, the at least one transgene of interest has a size of at least 5 kb, at least 10 kb, at least 15 kb, at least 20 kb, at least 25 kb, or more. In some embodiments, the at least one nucleic acid molecule comprisingthe at least one transgene of interest has a size of at least 5 kb, at least 10 kb, at least 15 kb, at least 20 kb, at least 25 kb, or more.
[0350] In some embodiments, the at least one nucleic acid molecule comprising the at least one transgene of interest is a transposon (or transposable element), i.e., in some embodiment, the at least one nucleic acid molecule comprising the at least one transgene of interest comprises at least one ITR sequence, preferably at least two ITR sequences, more preferably two ITR sequences.
[0351] In some embodiments, the at least one ITR sequence, preferably at least two ITR sequences, is selected from the group of sequences having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with the group comprising or consisting of SEQ ID NO: 37 to SEQ ID NO: 84. In some embodiments, the at least one ITR sequence, preferably at least two ITR sequences, is selected from the group comprising or consisting of SEQ ID NO: 37 to SEQ ID NO: 84. In some embodiments, the at least one ITR sequence, preferably at least two ITR sequences, is selected from the group of sequences having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with the group consisting of SEQ ID NO: 37 to SEQ ID NO: 84. In some embodiments, the at least one ITR sequence, preferably at least two ITR sequences, is selected from the group consisting of SEQ ID NO: 37 to SEQ ID NO: 84.
[0352] In a preferred embodiment, the at least one ITR sequence is adjacent to the nucleic acid sequence of the at least one gene of interest.
[0353] In a more preferred embodiment, the at least two ITR sequences flank the nucleic acid sequence of the at least one gene of interest.
[0354] In a more preferred embodiment, the two ITR sequences flank the nucleic acid sequence of the at least one gene of interest. Typically, the at least one nucleic acid molecule comprising the at least one transgene of interest comprises a left ITR and a right ITR.
[0355] As used herein, “left ITR” refers to the ITR sequence flanking the 5'-P extremity of the nucleic acid sequence of the at least one gene of interest; and “right ITR” refers tothe ITR sequence flanking the 3'-OH extremity of the nucleic acid sequence of the at least one gene of interest.
[0356] In some embodiments, the left ITR is selected from the group comprising or consisting of SEQ ID NO: 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, and 83. In some embodiments, the right ITR is selected from the group comprising or consisting of SEQ ID NO: 38, 40, 42, 44, 46, 48, 50, 52, 54, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, and 84. In some embodiments, the left ITR is selected from the group consisting of SEQ ID NO: 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, and 83. In some embodiments, the right ITR is selected from the group consisting of SEQ ID NO: 38, 40, 42, 44, 46, 48, 50, 52, 54, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, and 84.
[0357] In a preferred embodiment, the ITR sequences are recognized by the transposase as described herein.
[0358] In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 37 and the right ITR (3’-OH ITR) of SEQ ID NO: 38. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 39 and the right ITR (3’-OH ITR) of SEQ ID NO: 40. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 41 and the right ITR (3’-OH ITR) of SEQ ID NO: 42. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 43 and the right ITR (3’- OH ITR) of SEQ ID NO: 44. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 45 and the right ITR (3’-OH ITR) of SEQ ID NO: 46. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’- P ITR) of sequence SEQ ID NO: 47 and the right ITR (3’-OH ITR) of SEQ ID NO: 48. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 49 and the right ITR (3’- OH ITR) of SEQ ID NO: 50. In some embodiments, the nucleic acid sequence of the atI ll least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 51 and the right ITR (3’-OH ITR) of SEQ ID NO: 52. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’- P ITR) of sequence SEQ ID NO: 53 and the right ITR (3’-OH ITR) of SEQ ID NO: 54. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 55 and the right ITR (3’- OH ITR) of SEQ ID NO: 56. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 57 and the right ITR (3’-OH ITR) of SEQ ID NO: 58. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’- P ITR) of sequence SEQ ID NO: 59 and the right ITR (3’-OH ITR) of SEQ ID NO: 60. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 61 and the right ITR (3’- OH ITR) of SEQ ID NO: 62. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 63 and the right ITR (3’-OH ITR) of SEQ ID NO: 64. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’- P ITR) of sequence SEQ ID NO: 65 and the right ITR (3’-OH ITR) of SEQ ID NO: 66. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 67 and the right ITR (3’- OH ITR) of SEQ ID NO: 68. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 69 and the right ITR (3’-OH ITR) of SEQ ID NO: 70. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’- P ITR) of sequence SEQ ID NO: 71 and the right ITR (3’-OH ITR) of SEQ ID NO: 72. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 73 and the right ITR (3’- OH ITR) of SEQ ID NO: 74. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 75 and the right ITR (3’-OH ITR) of SEQ ID NO: 76. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’- P ITR) of sequence SEQ ID NO: 77 and the right ITR (3’-OH ITR) of SEQ ID NO: 78.In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 79 and the right ITR (3’- OH ITR) of SEQ ID NO: 80. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’-P ITR) of sequence SEQ ID NO: 81 and the right ITR (3’-OH ITR) of SEQ ID NO: 82. In some embodiments, the nucleic acid sequence of the at least one gene of interest is flanked by the left ITR (or 5’- P ITR) of sequence SEQ ID NO: 83 and the right ITR (3’-OH ITR) of SEQ ID NO: 84.
[0359] The present invention further relates to a method, preferably an in vitro method, for the integration, preferably the targeted integration, of at least one transgene of interest into the genome of one or more cell, or a population of cells, comprising contacting said one or more cell, or population of cells, with a composition comprising a transposase as described herein, and at least one nucleic acid molecule comprising the at least one transgene of interest, wherein said at least one transgene of interest is flanked with ITR recognized by the transposase.
[0360] The present invention further relates to a method, preferably an in vitro method, for the integration, preferably the targeted integration, of at least one transgene of interest into the genome of one or more cell, or a population of cells, comprising contacting said one or more cell, or population of cells, with:- a composition comprising an Anthonomus grandis DR1756049 transposase (Antgra6049) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 108, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 81 and a right ITR (z.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 82.
[0361] The present invention further relates to a method, preferably an in vitro method, for the integration, preferably the targeted integration, of at least one transgene of interestinto the genome of one or more cell, or a population of cells, comprising contacting said one or more cell, or population of cells, with:- a composition comprising an Anthonomus grandis DR1756049 transposase (Antgra6049) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 108 or SEQ ID NO: 35, or having an amino acid sequence as set forth in SEQ ID NO: 111, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 81 and a right ITR (z.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 82.
[0362] The present invention further relates to a method, preferably an in vitro method, for the integration, preferably the targeted integration, of at least one transgene of interest into the genome of one or more cell, or a population of cells, comprising contacting said one or more cell, or population of cells, with:- a composition comprising an Anthonomus grandis DR1754440 transposase (Antgra4440) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 30, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 71 and a right ITR (z.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 72.
[0363] The present invention further relates to a method, preferably an in vitro method, for the integration, preferably the targeted integration, of at least one transgene of interest into the genome of one or more cell, or a population of cells, comprising contacting said one or more cell, or population of cells, with:- a composition comprising an Anthonomus grandis DR1754440 transposase (Antgra4440) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 30, or having an amino acid sequence as set forth in SEQ ID NO: 113, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 71 and a right ITR (z.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 72.
[0364] The present invention further relates to a method, preferably an in vitro method, for the integration, preferably the targeted integration, of at least one transgene of interest into the genome of one or more cell, or a population of cells, comprising contacting said one or more cell, or population of cells, with:- a composition comprising an Anthonomus grandis DR1754053 transposase (Antgra4053) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 108, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 73 and a right ITR (z.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 74.
[0365] The present invention further relates to a method, preferably an in vitro method, for the integration, preferably the targeted integration, of at least one transgene of interest into the genome of one or more cell, or a population of cells, comprising contacting said one or more cell, or population of cells, with:- a composition comprising a Poeciliopsis turrubarensis transposase having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with any one of SEQ ID NO: 1 to SEQ ID NO: 13, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 37 and a right ITR (z.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 38.
[0366] The present invention further relates to a method, preferably an in vitro method, for the integration, preferably the targeted integration, of at least one transgene of interest into the genome of one or more cell, or a population of cells, comprising contacting said one or more cell, or population of cells, with:- a composition comprising a Poeciliopsis turrubarensis transposase having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with any one of SEQ ID NO: 1 to SEQ ID NO: 13 or SEQ ID NO: 112, preferably having an amino acid sequence as set forth in SEQ ID NO: 112, and- at least one nucleic acid molecule comprising at least one transgene of interest, wherein the at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 37 and a right ITR (i.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 38.
[0367] The present invention further relates to a method, preferably an in vitro method, for the integration, preferably the targeted integration, of at least one transgene of interest into the genome of one or more cell, or a population of cells, comprising contacting said one or more cell, or population of cells, with:- a composition comprising a Poeciliopsis turrubarensis transposase having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with any one of SEQ ID NO: 2 to SEQ ID NO: 13, and- at least one nucleic acid molecule comprising said at least one transgene of interest, wherein said at least one transgene of interest is flanked with a left ITR i.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO:37 and a right ITR (z.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 38.
[0368] The present invention further relates to a method, preferably an in vitro method, for the integration, preferably the targeted integration, of at least one transgene of interest into the genome of one or more cell, or a population of cells, comprising contacting said one or more cell, or population of cells, with:- a composition comprising a Poeciliopsis turrubarensis transposase having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with any one of SEQ ID NO: 2 to SEQ ID NO: 13 or SEQ ID NO: 112, preferably having an amino acid sequence as set forth in SEQ ID NO: 112, and- at least one nucleic acid molecule comprising at least one transgene of interest, wherein the at least one transgene of interest is flanked with a left ITR (z.e., 5’-P ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 37 and a right ITR (i.e., 3’-OH ITR) having at least 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO: 38.
[0369] The present invention further relates to a cell comprising at least one transposase, or nucleic acid encoding thereof, as defined herein.
[0370] In some embodiments, the cell further comprises a RNA-guided nuclease or nickase or a nucleic acid encoding thereof.
[0371] In some embodiments, the cell further comprises at least one nucleic acid molecule comprising the at least one transgene of interest, as defined herein.
[0372] The present invention further relates to a vector encoding at least one transposase, or nucleic acid encoding thereof, as defined herein. Suitable vector are known in the art.
[0373] The present invention further relates to an in vivo method for the integration of at least one transgene of interest into the genome of one or more cell, comprising contacting said one or more cell with the composition according to the invention, and atleast one nucleic acid molecule comprising said at least one transgene of interest. Any feature of the in vitro method as described herein may be applied to the in vivo method.
[0374] The present invention further relates to a pharmaceutical composition comprising the transposase according to the invention.
[0375] In particular, the present invention further relates to a pharmaceutical composition comprising the transposase according to the invention, and a pharmaceutically acceptable excipient.
[0376] In some embodiments, the pharmaceutically acceptable excipient is selected in a group comprising or consisting of a solvent, a diluent, a carrier, a vehicle, a dispersion medium, a coating, an antibacterial agent, an antifungal agent, an isotonic agent, an absorption delaying agent and any combinations thereof. The carrier, diluent, solvent or vehicle must be “acceptable” in the sense of being compatible with the transposase, and not be deleterious upon being administered to an individual. Typically, the vehicle does not produce an adverse, allergic or other untoward reaction when administered to an individual, preferably a human individual. For the particular purpose of human administration, the pharmaceutical compositions should meet sterility, pyrogenicity, general safety and purity standards as required by regulatory offices, such as, for example, the Food and Drugs Administration (FDA) Office or the European Medicines Agency (EMA). Suitable excipients include, without limitation, mannitol, dextrose, lactose, starch, magnesium stearate, sodium saccharine, cellulose, magnesium carbonate, and the like. Acceptable carriers, solvents, diluents and vehicles for therapeutic use are well known in the pharmaceutical art. The choice of a suitable pharmaceutical carrier, solvent, excipient or vehicle can be made with regard to the intended route of administration and standard pharmaceutical practice. The pharmaceutical compositions may comprise as, or in addition to, the carrier, vehicle, solvent or diluent any suitable binder, lubricant, suspending agent, coating agent, or solubilizing agent. Preservatives, stabilizers, dyes and even flavoring agents may be provided in the pharmaceutical composition.
[0377] The present invention further relates to a medicament comprising the composition or pharmaceutical composition according to the invention. The presentinvention further relates to the composition according to the invention, or the pharmaceutical composition according to the invention, for use as a medicament.
[0378] The present invention further relates to the composition, pharmaceutical composition or medicament according to the invention, for use for treating a genetic disease in a subject in need thereof. In some embodiments, the composition, pharmaceutical composition, or medicament for use for treating a genetic disease further comprises at least one nucleic acid molecule comprising at least one transgene, wherein the expression of the transgene compensates the genetic defect of the genetic disease. Illustratively, the at least one transgene may encode a protein and thereby compensate the lack or absence of expression of said protein in the genetic disease; or, alternatively, the at least one transgene may encode a siRNA or shRNA or miRNA inhibiting the expression of an overexpressed gene.
[0379] The present invention further relates to a method for treating a subject in need thereof, comprising administering to said subject the composition, pharmaceutical composition or medicament according to the invention.
[0380] The present invention further relates to a method for treating a genetic disease in a subject in need thereof, comprising administering to said subject a therapeutically effective dose of the composition, pharmaceutical composition or medicament according to the invention. In some embodiments, the method further comprises co-administering to the subject at least one nucleic acid molecule comprising at least one transgene, wherein the expression of the transgene compensates the genetic defect of the genetic disease.
[0381] The composition, pharmaceutical composition, or medicament according to the invention may be administered once, twice, or any number of times.
[0382] The present invention further relates to the composition or pharmaceutical composition according to the invention, for use for the manufacture of a medicament for treating a genetic disease in a subject in need thereof.BRIEF DESCRIPTION OF THE DRAWINGS
[0383] Figure 1 is a scheme showing the phylogenetic analysis of PiggyBac derived transposases across eukaryotic genomes.
[0384] Figure 2 is a histogram showing the transposition activity by PiggyBac derived transposases measured as Stable RFP Cargo integration after 3 weeks.
[0385] Figure 3 is a histogram representation showing the transposition activity by additional PiggyBac derived transposases measured as Stable RFP Cargo integration after 3 weeks.
[0386] Figure 4 is a schematic representation of the programable insertion reporter cell line.
[0387] Figure 5 is a histogram representation showing the programable transposition activity by PiggyBac derived transposases variants measured as GFP reporter reconstitution upon cargo insertion.
[0388] Figure 6 is a scheme representing the phylogenetic analysis of piggyBac derived transposases across eukaryotic genomes.
[0389] Figure 7A-7B is a set of histograms assessment of selected transposase orthologs for random integration of an RFP cargo in HEK293T cells, 22 days post plasmid transfection. (A) activity and (B) activity normalized to day 22.
[0390] Figure 8A-8C is a set of histograms showing random integration of an RFP cargo in HEK293T cells, 22 days post plasmid transfection (Fig. 8A) or 2 weeks post plasmid transfection (Fig. 8B), with other transposase orthologs. Fig. 8C is a histogram showing random integration of an RFP cargo in HEK293T cells, 16 days post plasmid transfection, with transposase orthologs and hyperactive PiggyBac transposase (hyPB). Comparison is made in conditions with the transposon alone or with Cas9 or with the transposon and the transposase (PB)
[0391] Figure 9 is a histogram showing targeted integration of an RFP cargo in HEK293T cells with the transposase orthologs (3X, coupled to Cas9) and hyperactive PiggyBac transposase (hyPB).
[0392] Figure 10 is a photograph showing targeted integration compatibility across PiggyBac orthologs.
[0393] Figure 11A-11B is a set of graphs and schemes showing on-target insertion obtained with the transposase PoeTurrub (from Poeciliopsis Turrubarensis). (Fig. 11A) NGS reads:PoeTurrub AAVS1::::3’ITR. (Fig. 11B) Profiles of the top alleles for PoeTurrub 3X NGS reads: The red line delineates the boundary between the AAVS1 genomic region on the left and the 3’ ITR on the right.
[0394] Figure 12A-12B is a set of graphs and schemes showing on-target insertion obtained with the transposase Antgra6049 (from Anthonomus grandis). (Fig. 12A) NGS reads: Antgra6049 AAVS1::::3TTR (Fig. 12B) Profiles of the top alleles for Antgra6049 3X NGS reads: The red line delineates the boundary between the AAVS 1 genomic region on the left and the 3’ ITR on the right.
[0395] Figure 13A-13B is a set of graphs and schemes showing on-target insertion obtained with the transposase Antgra4440 (from Anthonomus grandis). (Fig. 13A) NGS reads: Antgra4440 AAVS1::::3TTR. (Fig. 13B) Profiles of the top alleles for Antgra4440 3X NGS reads: The red line delineates the boundary between the AAVS 1 genomic region on the left and the 3’ ITR on the right.
[0396] Figure 14A-14B is a set of graphs and schemes showing on-target insertion obtained with the transposase Antgra4053 (from Anthonomus grandis). (Fig. 14A) NGS reads: Antgra4053 AAVS1::::3TTR. (Fig. 14B) Profiles of the top alleles for Antgra4053 3X NGS reads: The red line delineates the boundary between the AAVS 1 genomic region on the left and the 3’ ITR on the right.
[0397] Figure 15 is a scheme showing the percentage of RFP positive cells, i.e., the overall integration obtained with various transposase orthologs tested.
[0398] Figure 16A-16B is a set of histograms showing ortholog activity in primary cells from a first donor (Fig. 16A) and a second donor (Fig. 16B).
[0399] Figure 17A-17B is a set of histograms showing targeted integration with ortholog transposases. The best transposases identified in the bioprospecting screen were evaluated for programmable integration by co-transfection of Cas9, AAVS1 targeting gRNA and transposase-transposon plasmid pairs. Fig. 17A shows the percentage of RFP positive cells; Fig. 17B shows junction qPCR A.U.
[0400] Figure 18 a scheme showing the sequence alignment of the original PiggyBac transposase and the hyperactive PiggyBac transposase with the transposase orthologs.EXAMPLES
[0401] The present invention is further illustrated by the following examples.Example 1Materials and Methods
[0402] Identification of PiggyBac ORFs:
[0403] In order to identify putative active PiggyBac transposon ORFs, a hmm model was built using hmmbuild using all the active identified PiggyBac sequences reported in the literature. All sequences identified as PiggyBac transposases in Dfam database were fetched and searched for putative functional ORFs using the build hmm model and phramer. ORFs longer than 400aa containing all functional PB domains were used for subsequent ITR annotation. In order to further expand the number of putative PiggyBac transposases, all eukaryotic genomes were searched for PB derived ORFs using the same phramer based strategy, and elements above lOelO e value where and longer than 400aa containing all pb functional domains where kept. The phylogenetic analysis of PiggyBac derived transposases across eukaryotic genomes is shown on Figure 1.
[0404] In order to identify ITR sequences from the putative PiggyBac Transposases, 2000 bp in both upstream and downstream direction from the PB ORF where fetched andsearched for inverted repeats using PALINDROME (Hwee Kim, et al., Bioinformatics, 2016). Transposases with ITRs longer than lObp and an intact catalytic DDE motif were considered as putative active transposases. Clustering of the sequences was performed, and representatives from each cluster were selected for experimental validation.
[0405] Assembly of transposase expressing plasmids:
[0406] Transposase ORF aminoacidic sequences were codon optimized for Homo Sapiens and ordered as synthesized as gene fragments to TWIST biosciences. Gene fragments were cloned in to CMV based expression vector by Golden Gate assembly using Esp3I. Transposon (cargo vector) plasmid sequences were defined as the first 150bp from the transposon ends from both 5’ and 3’ ITR sequences and synthesized as gene fragments by TWIST biosciences with added overhangs for golden gate assembly. EFla RFP polyA expression cassette was included between ITRs.
[0407] Transposition activity assays:
[0408] 120k HEK293T cells were seeded in p24 wells the day prior to transfection. Transposase and transposon DNA was transfected in to HEK293T using. 0.4 ug of transposase expressing vector and 1.5ug of transposon plasmid and 7.5ul of PEI (see Figure 4). ITR vector RFP expression was measured two days after transfection and 20 days after transfection. RFP signal at day 20 was taken as the integration efficiency for each of the tested systems, as it indicated stable transgene integration.
[0409] Programable transposition activity assays
[0410] 120k HEK293T Hershey reporter cells were seeded in p24 wells the day prior to transfection. Programmable nuclease, gRNA, variants of the transposase and transposon DNA was transfected in to reporter HEK293T using. 0.4 ug each of nuclease, gRNA, and transposase expressing vector and 1.5ug of transposon plasmid and 8ul of PEI. Reporter GFP expression was measured 5 days after transfection, and signal was taken as the programmable integration efficiency for each of the tested systems, as it indicated site specific insertion reconstituting GFP expression.Results
[0411] Results of RFP integration assays are shown on Figures 2, 3 and 5.Example 2Materials and Methods
[0412] Identification of PiggyBac ORFs; plasmid assembly; transposition activity assay
[0413] See Example 1.
[0414] Molecular biology and Cloning
[0415] To test the sequences discovered in PB mining (structural search), the orthogonal piggyBac sequences were cloned in a pcDNA 3.1 (A8-H7 of pcDNA™3.1(+)-Esp3I) vector and a Cas9-GG (pSico-Cas9-GG) backbone to express transposase. The paired Yi GFP or RFP ITRs were cloned in a7-Al-pucl9-Esp3I backbone with Vi GFP from t5-I3 EFla Vi emGFP and full RFP from t6-F6 EFla-RFP-pA respectively. The resulting plasmids were sequence verified using sanger sequencing.
[0416] Transfection
[0417] Transfection of the plasmids containing PB orthologs ORFs with their corresponding ITRs containing Full RFP was done in Hek293 T cells with a ratio of 1:5 of transposases to transposons. The initial FACS reading was taken 48 hours post transfection on Fortes sa or Aurora and the cells were maintained for 2 to 3 weeks for the final FACS reading by passing every other day until the episomal control reading was close to zero. Overall efficiency of the PB orthologs for their ability to integrate randomly was calculated by normalizing the final percentage of RFP cells by their initial reading at 48 hours. Additionally, RFP positive cells were sorted and expanded from which genomic DNA were collected and processed for NGS miSeq analysis.
[0418] Targeted Transposition activity assays
[0419] Triple mutant residue selection was performed by alignment of the functionalpiggyBacs to the Trichoplusia Ni PiggyBac sequence. Triple mutant encoding plasmids (PBx3) were co-transfected with Cas9 and gRNA, transposon plasmids in to 0.5M HEK293T cells seeded in a p6 plate. Cells were analyzed for RFP expression two days after transfection. Two rounds of enrichment via RFP sorting were performed, one week after transfection and again two weeks after transfection. Genomic DNA was extracted using quiagen columns 4 days after second sorting. 3’ Junction PCR was purified in the cases where the PCR product was sanger sequenced.
[0420] Illumina sequencing for targeted insertion analysis
[0421] Genomic DNA was extracted from enriched cellular samples. Subsequently, junction PCR was conducted employing primers containing P5 and P7 adapters. The selected primers for the junction PCR were designed to anneal at the 3’ end of the transposon (forward) and the genomic locus of AAVS1. Illumina reads were processed utilizing CRISPR-A to acquire both the quantity and profile of indels.Results
[0422] PiggyBac transposases orthologs sequences from Dfam database are filtered to exclude non-functional sequences. Remaining sequences were prioritized by visual inspection of alignment and ITRs, inclusion of diverse sequences, and synthesis, which led to final selection of 13 transposase sequences (see Figure 6).
[0423] These 13 selected PiggyBac transposase orthologs sequences from bioprospecting analysis were assessed for random integration of an RFP cargo in HEK293T cells, 22 days post plasmid transfection. Results are shown in Figure 7A-7B.
[0424] 12 subsequent bioprospected PiggyBac transposase orthologs sequences were identified from a second identification round. The activities of these transposases were tested as in round 1: random integration of an RFP cargo in HEK293T cells, 22 days post plasmid transfection. Results are shown on Figure 8A.
[0425] Another selection was performed and 10 PiggyBac transposase orthologs that showed transpositional acitivity in Hek 293T cells when co-transfected with transposons flanked by their corresponding ITRs were identified. The orthologs presented in Table 1were identified as positive hits.
[0426] Table 1
[0427] The PiggyBac orthologs demonstrated successful transposition of the respective transposons containing RFP, as evidenced by the presence of RFP signal two weeks post transfection (Figure 8B). In contrast, the corresponding episomal control signals were negligible, indicating the effective integration of the transposons into the host genome. The degree of integrability varied across orthologs, with integration efficiencies ranging from 3 percent to as high as 30 percent. The data shown here is representative of multiple biological replicates.
[0428] Further, it can be seen on Figure 8C the RFP cargo random integration activity in HEK293T cells 16 days after transfection. The results show that several PiggyBactransposase orthologs sequences have integration activity comparable to hyperactive PiggyBac transposase (hyPB). In particular, the transposases from Poeciliopsis Turrubarensis and Anthonomus grandis show equal or superior integration activity when compared to hyPB.
[0429] The three best performing transposases were evaluated for programmable integration. Targeted integration is shown on Figure 9.
[0430] To assess the ability of these orthologs for targeted integration, cells were transfected with Cas9, the orthologs and their variants along with their corresponding transposons containing Full RFP. Cells were sorted to enrich for RFP- positive cells, followed by expansion and maintenance for two weeks. Genomic DNA was extracted from these cells and used as templates for PCR, specifically using primers that annealed to the transposon and the genomic locus of AAVS1. The results from the PCR gel demonstrated bands of the correct size (as illustrated in Figure 10), which were subsequently excised, gel purified, and sent for Sanger sequencing. The sequencing results confirmed the presence of positive junctions between the corresponding transposons and AAVS1, validating the precise and targeted integration of RFP. Qualitatively, it was observed that the 3X mutant exhibited stronger bands in the gel compared to their WT equivalents, suggesting enhanced integration efficiency. The Sanger sequencing result showed positive junction between the ITRs and AAVS1 locus for the three selected PiggyBac orthologs.
[0431] The on-target insertion was characterized by a junction PCR followed by next generation sequencing to precisely capture and profile the payload-genome junctions at AAVS1. While precise on-target insertion of the pay load in all four examined PiggyBac orthologs can be detected (Figures 11A, 11B, 12A, 12B, 13A, 13B, 14A, and 14B), no insertion at nearby TTAA site was detected, demonstrating integration on DSB sites generated by Cas9. Furthermore, the lack of plasmid element in the insertion site further demonstrates the clean excision of the payload flanking TTAA by the PiggyBac orthologs.
[0432] Despite the absence of detection of indels in the vicinity of the integration site,notably all PiggyBac orthologs examined, except Antgra6049, exhibited WT allele (perfect aligned allele against the reference sequence) as the most abundant alleles in the pool (as illustrated in Figures 11B, 12B, 13B and 14B). The next-generation sequencing results for WT PiggyBac orthologs indicated low read counts at the junctions, potentially attributable to the presence of a proportionately higher number of randomly integrated payloads, which may compete for primer annealing during the PCR process.
[0433] These observations collectively provide strong scientific evidence for the successful and targeted integration of payloads by the PiggyBac orthologs. These findings hold substantial promise for a wide range of applications in the fields of genetic engineering, biotechnology, and biomedical research. These applications include:Gene Therapy: The precise integration of therapeutic genes into the host genome is a critical component of gene therapy. These observations suggest that the PiggyBac orthologs, especially the 3X mutant, could serve as valuable tools for developing more effective and reliable gene therapy techniques. This has the potential to revolutionize the treatment of genetic disorders and other diseases at the molecular level.Stem Cell Research: Stem cell therapy and regenerative medicine depend on the accurate integration of specific genes into stem cells. The observed efficiency in targeted integration opens new avenues for the genetic modification of stem cells, enabling their use in tissue repair, organ transplantation, and disease modeling.Functional Genomics: Understanding the function of specific genes often involves manipulating their expression in a controlled manner. The orthologs’ ability to integrate payloads with precision offers researchers a powerful tool for studying gene function in various organisms.Bioproduction and Biopharmaceuticals: In bioproduction processes, such as the production of therapeutic proteins, optimizing host cells for efficient protein expression is crucial. The precise integration of genes, as demonstrated by these orthologs, can improve the development of high-yield cell lines for biopharmaceutical production.Synthetic Biology: The precise and efficient integration of transposons is a cornerstone of synthetic biology. It enables the creation of custom genetic circuits and metabolic pathways for various applications, including biofuel production, environmental remediation, and bioplastic synthesis.
[0434] In summary, the observations of successful and targeted integration of transposons containing RFP by the PiggyBac orthologs, with the 3X mutant showing enhanced efficiency, open up a multitude of possibilities across the life sciences. These applications have the potential to drive innovation and advancements in fields critical to human health, agriculture, and the environment, making the orthologs and the methods described in this research highly promising and valuable assets for future scientific and commercial endeavors.Example 3Materials and Methods
[0435] Identification of piggybac ORFs
[0436] In order to identify putative active piggyBac transposon ORFs, a hmm model was built using all the active identified piggybac sequences reported in the literature. All sequences identified as piggyBac transposases in Dfam database were fetched and searched for putative functional orfs using the build hmm model and phramer. ORFs longer than 400 amino acids containing all functional PB domains were used for subsequent ITR annotation. In order to further expand the number of putative piggyBac transposases, all eukaryotic genomes were searched for PB derived ORFs using the same phramer based strategy, and elements above lOelO e value where and longer than 400aa containing all pb functional domains where kept.
[0437] In order to identify ITR sequences from the putative piggybac transposases, 2000bp in both upstream and downstream direction from the PB ORF where fetched and searched for inverted repeats using PAEINDROME. Transposases with ITRs longer than lObp and an intact catalytic DDE motif were considered as putative active transposases.Clustering of the sequences was performed, and representatives from each cluster were selected for experimental validation.
[0438] In October 2022, all transposons labeled as piggyBac (PB) in the Dfam database were downloaded. They were screened them by running them against the wild type PB seed, made from the PB transposase, in HMMER, and those that had a positive hit were recovered. The resulting transposons were filtered by selecting only those between 1400- 4000 bp in length (846 sequences). A second filter based on transposons having ITR’s longer than 10 bp and including the tetranucleotide TTAA or trinucleotide TAA was applied. This step reduced the dataset to 83 transposons. The ITRs were annotated using EMBOSS palindrome, by searching for palindromes formed from the DNA sequences flanking the transposase which had a TTAA or TAA motif. From the 83 transposons, 16 top transposons were selected through a manual curation based on them having a transposase with an open reading frame (ORF) between 400-700 amino acids, ITR’s 20- 500 bp away from the beginning or end of the transposase, ITR’s composed of at least 2 different palindromes, and the species they came from (higher priority for mammals).
[0439] Piggybac transposon identification
[0440] In order to identify active piggybac transposons an hmm model was built, using hmmbuild, adding active piggybac transposases reported in literature. An hmm search was performed with the program BATH (formerly known as frahmmer) on all Mammalian genomes in NCBI, and all entries in the Dfam database. Bath was used as it can detect and align translated homologies, even in the presence of frameshift errors. Only piggybac ORFs longer than 400aa containing all functional PB domains were used for ITR identification.
[0441] In order to identify ITR’s, 4000bp both upstream and downstream from the ORF were taken to identify palindromes with the program Palindrome from EMBOSS. Sequences with palindromes longer than lObp, in which the two most common nucleotides account for less than 80% of the palindrome, and palindromes formed between a sequence upstream and a sequence downstream were considered as putative active transposons. Transposons were then delimited by palindromes containing theTTAA motif. Clustering of the sequences was performed using MMseqs2 with an identity and coverage of 0.9, and representatives from each cluster were selected for experimental validation.
[0442] 20 final sequences were selected for experimental validation based on their ITR architecture similarity to the piggybac transposon. An additional 2 sequences were selected as representatives from the zoonomia project in DFAM. 4 additional sequences were selected from a blastn search in NCBI with the most active transposon (Poeciliopsis), with a sequence identity higher than 0.85 and were tested.
[0443] Plasmid assembly
[0444] Transposase ORF amino acid sequences were codon optimized for Homo Sapiens and ordered and synthesized as gene fragments to TWIST biosciences. Gene fragments were cloned into a CMV based expression vector by Golden Gate assembly using Esp3I restriction enzyme. Transposon (cargo vector) plasmid sequences were defined as the first 150bp from the transposon ends from both 5’ and 3’ ITR sequences and synthesized as gene fragments by TWIST biosciences with added overhangs for golden gate assembly. EFla RFP polyA expression cassette was included between the ITRs.
[0445] Transposition activity assays
[0446] 120k HEK293T cells were seeded in p24 wells the day prior to transfection. Transposase and transposon DNA was transfected into HEK293T using. 0.4 ug of transposase expressing vector and 1.5ug of transposon plasmid and 7.5ul of PEI. ITR vector RFP expression was measured two days after transfection and 20 days after transfection. RFP signal at day 20 was taken as the integration efficiency for each of the tested systems, as it indicated stable transgene integration.
[0447] Hek293T cells, obtained from Thermo Fisher Scientific, were cultured in Dulbecco’s modified eagle medium (DMEM) supplemented with high glucose (Gibco, Thermo Fisher), 10% Fetal Bovine Serum (FBS), 2 mM glutamine, and 100 U penicillin / 0.1 mg / mL streptomycin at 37°C in a 5% CO2 incubator. For transfectionexperiments, cells were treated with Polyethyleneimine (PEI, Thermo Fisher Scientific) at a 1:3 DNA-PEI ratio in OptiMem. Prior to transfection, cells were seeded to achieve 70% confluency the following day, typically using 120,000 cells per adherent p24 well plate. Plasmid DNA was mixed at a ratio of 1 transposase:3 gRNA:5 transposon, with 0.035 pmols of transposase used per p24 well plate.
[0448] The expression of the ITR vector RFP was assessed two days and twenty days post-transfection using cell cytometry with the Cytek Aurora™ CS System. The RFP signal at day twenty was considered indicative of stable transgene integration and was utilized to determine the integration efficiency for each of the tested systems.
[0449] Excision activity assays
[0450] 48h post-transfection, cells were harvested and a miniprep protocol was performed using the NZYMiniprep kit from NZYtech to extract the remaining plasmids. Primers flanking the ITRs in the transposon plasmids were utilized to detect transposon excision. Expecting a 2900bps band indicated not-excised transposon, whereas a 1200bps band indicated excision of both ITR and transposon.
[0451] Targeted Transposition activity assays
[0452] Triple mutant residue selection was performed by alignment of the functional piggyBacs to the Trichoplusia Ni PiggyBac sequence. Triple mutant encoding plasmids (PBx3) where co-transfected with Cas9 and gRNA, transposon plasmids into HEK293T cells seeded in a p6 plate. Cells were analyzed for RFP expression two days after transfection. Two rounds of enrichment via RFP sorting were performed, one week after transfection and again two weeks after transfection. Genomic DNA was extracted using quiagen columns 4 days after second sorting. 3’ Junction PCR was performed and gel purified in the cases the PCR product was sanger sequenced.
[0453] Triple mutant residue selection was performed by alignment of the functional piggyBacs to the Trichoplusia Ni PiggyBac sequence. Plasmids encoding the triple mutant variants (PBx3) were co-transfected with Cas9, gRNA and transposon plasmids in a 1: 1:3:5 molar ratio into 0.5M Hek23T cells seeded in a p6 plate the day beforetransfection. Cells were analyzed for RFP expression two days after transfection using cell cytometry with the Cytek Aurora™ CS System. Subsequently, two rounds of enrichment via RFP sorting were conducted with BD FACSAria (Biosciences), one week after transfection and again two weeks after transfection. Genomic DNA was extracted using Quiagen DNeasy Blood & Tissue Kit column’s four days after the second sorting. A 3’ Junction PCR was performed and gel purification using the QIAquick Gel Extraction Kit was performed when necessary before Sanger sequencing the amplified bands.Results
[0454] With the aim of exploring transposition activity across bioprospected PiggyBac elements, a subset of transposase were selected for evaluating transposition activity in mammalian cells. 22 bioprospected were selected for testing. Results are shown in Figure 15 and Figure 16A-16B.
[0455] Long term integration in mammalian was analyzed from which 6 sequences are considered active (5% of RFP positive cells after 2 weeks) and excision of the active ones was measured.
[0456] An analysis of activity vs transposase and transposon architecture was done. A fisher test between the selected variables and active sequences (sequences with 5% RFP) showed a clear correlation with sequences that completed all the qualitative checks and had at least 2 palindromes. Odds ratio 22.4 were obtained with P-value of 0.0036. There was a positive correlation with sequences that completed all qualitative checks but had at least 3 palindromes. Odds ratio 9.8 were obtained with P-value of 0.016.
[0457] An ANOVA was performed taking this two as classifiers and activity (%RFP cells) which showed in both cases a significant difference in means. The same was done for the individual qualitative variables but no association was shown for any single variable.
[0458] An in silico deep mutational scan with esm-lv was performed on Poeciliopsis turrubarensis . 12 mutations were chosen to be experimentally tested, 8 being unique mutations in a single position and 4 being 2 mutations per position. For this selection, 2mutations from each domain were selected. The results can be observed in Figure 5, where 3 out of the 12 tested mutants showed higher activity than transposase 7 wild type, 2 showed no considerable changes and the rest showed lower activity. The mutant with the highest increase in activity showed a log2 fold change of 0.195209, which is equivalent to approximately 10% increase in activity. While the mutant with the lowest activity showed a log2 fold change of about -0.757048 which is a decrease in activity of approximately 40%.
[0459] Compatibility of the identified orthologs with previously described FiCAT programmable integration system, as described in W02020250181 and WO2022129438, was tested. Results are shown in Figure 17A-17B.
[0460] The results show long term integration with Cas9 and gRNA with the transposase orthologs tested. Both WT sequence of orthologs with SEQ ID NO: 35, 1 and 30, as well as HyPB R372A / K375A / D450N triple mutant equivalents (herein referred to as X3 mutants), i.e., transposases of SEQ ID NO: 111, 112 and 113, were tested. Of note, R372A / K375A / D450N are the most important residues for increasing on-target in FiCAT in respect to the HyPB sequence; the sequence alignments of the orthologs with hyPB can be seen on Figure 18. X3 mutants seem to have higher on-target efficiencies, matching with observations in FiCAT with HyPB. This result was confirmed by junction PCR and was quantified with qPCR.
[0461] These results show new active transposase orthologs with transposition activity and that are compatible in FiCAT and enable targeted integration.CitationsChoudhary, M.N.K., el al., (2023), Widespread contribution of transposable elements to the rewiring of mammalian 3D genomes., Nat. Commun., 14, 634.Hwee Kim, Yo-Sub Han, April 2016, OMPPM: online multiple palindrome pattern matching, Bioinformatics, Volume 32, Issue 8, Pages 1151-1157.
Claims
CLAIMS1. A composition comprising at least one transposase selected from the group consisting of: a Poeciliopsis turrubarensis transposase, or a variant thereof; an Anthonomus grandis DR1754440 transposase, or a variant thereof; an Anthonomus grandis DR1754053 transposase, or a variant thereof; a Anthonomus grandis DR1756049 transposase, or a variant thereof; a Xenopus tropicalis transposase, or a variant thereof; a Japanese Medaka transposase, or a variant thereof; a Leptobrachium leishanense transposase, or a variant thereof; a Scalopus aquaticus ScaAqu-5.3491 transposase, or a variant thereof; an Atlantic Salmon transposase, or a variant thereof; a Heliconius butterfly transposase, or a variant thereof; a Takifugu flavidus transposase, or a variant thereof; a Vaquita transposase, or a variant thereof; a Oryzias latipes transposase, or a variant thereof; a Salvelinus profundus transposase, or a variant thereof; a Cuniculus paca transposase, or a variant thereof; a Noctilio leporinus transposase, or a variant thereof; a Pipistrellus pipistrellus PipPip-6.1914 transposase, or a variant thereof; a Carlito syrichta transposase, or a variant thereof;a Myotis lucifugus PiggyBat transposase, or a variant thereof; a PiggyBac transposase, or a variant thereof; _Solenopsis invicta DR3053925 transposase, or a variant thereof; a Simochromis diagramma genomic transposase, or a variant thereof; a Nematolebias whitei chromosome 17 transposase, or a variant thereof; and a Coremacera marginata DR1481656 transposase, or a variant thereof; or a nucleic acid encoding thereof.
2. The composition according to claim 1, wherein said at least one transposase is selected from the group consisting of:Poeciliopsis turrubarensis transposase, or a variant thereof;Anthonomus grandis DR1754440 transposase, or a variant thereof; an Anthonomus grandis DR1754053 transposase, or a variant thereof; and a Anthonomus grandis DR1756049 transposase, or a variant thereof.
3. The composition according to claim 1 or 2, wherein said at least one transposase is selected from the group consisting of:Poeciliopsis turrubarensis transposase, or a variant thereof; andAnthonomus grandis DR1754440 transposase, or a variant thereof.
4. The composition according to any one of claims 1 to 3, wherein said at least one transposase is a Poeciliopsis turrubarensis transposase, or a variant thereof.
5. The composition according to claim 4, wherein said Poeciliopsis turrubarensis transposase has an amino acid sequence having at least 75% sequence identity with SEQ ID NO: 1.
6. The composition according to claim 4, wherein said Poeciliopsis turrubarensis transposase has an amino acid sequence as set forth in SEQ ID NO: 1.
7. The composition according to claim 4, wherein said Poeciliopsis turrubarensis transposase variant has an amino acid sequence as set forth in any one of SEQ ID NO: 2 to SEQ ID NO: 13 or SEQ ID NO: 112.
8. The composition according to claim 4, wherein said Poeciliopsis turrubarensis transposase variant comprises at least one amino acid substitution on one or more of the amino acids at positions 353, 356, and 432 corresponding to the amino acid numbering of SEQ ID NO: 1.
9. The composition according to claim 4, wherein said Poeciliopsis turrubarensis transposase variant has an amino acid sequence as set forth in SEQ ID NO: 112.
10. The composition according to any one of claims 1 to 3, wherein said at least one transposase is an Anthonomus grandis DR1754440 transposase, or a variant thereof.
11. The composition according to claim 10, wherein said Anthonomus grandis DR1754440 transposase has an amino acid sequence having at least 75% sequence identity with SEQ ID NO: 30.
12. The composition according to claim 10, wherein said Anthonomus grandis DR1754440 transposase has an amino acid sequence as set forth in SEQ ID NO: 30.
13. The composition according to claim 10, wherein said Anthonomus grandis DR1754440 transposase variant comprises at least one amino acid substitution on one or more of the amino acids at positions 388, 389, and 393 corresponding to the amino acid numbering of SEQ ID NO: 30.
14. The composition according to claim 10, wherein said Anthonomus grandis DR1754440 transposase variant has an amino acid sequence as set forth in SEQ ID NO: 113.
15. The composition according to claim 1 or 2, wherein said at least one transposase is an Anthonomus grandis DR1756049 transposase, or a variant thereof.
16. The composition according to claim 15, wherein said Anthonomus grandis DR1756049 transposase has an amino acid sequence having at least 75% sequence identity with SEQ ID NO: 35 or SEQ ID NO: 108.
17. The composition according to claim 15, wherein said Anthonomus grandis DR1756049 transposase has an amino acid sequence as set forth in SEQ ID NO: 35 or SEQ ID NO: 108.
18. The composition according to claim 15, wherein said Anthonomus grandis DR1756049 transposase variant comprises at least one amino acid substitution on one or more of the amino acids at positions 348, 352, and 429 corresponding to the amino acid numbering of SEQ ID NO: 35.
19. The composition according to claim 15, wherein said Anthonomus grandis DR1756049 transposase variant has an amino acid sequence as set forth in SEQ ID NO: 111.
20. The composition according to claim 1 or 2, wherein said at least one transposase is an Anthonomus grandis DR1754053 transposase, or a variant thereof.
21. The composition according to claim 20, wherein said Anthonomus grandis DR1754053 transposase has an amino acid sequence having at least 75% sequence identity with SEQ ID NO: 31.
22. The composition according to claim 20, wherein said Anthonomus grandis DR1754053 transposase has an amino acid sequence as set forth in SEQ ID NO: 31.
23. The composition according to any one of claims 1 to 22, wherein said transposase recognizes at least one ITR sequence, preferably at least two ITR sequences, selected from the group consisting of SEQ ID NO: 37 to SEQ ID NO: 84.
24. The composition according to any one of claims 4 to 9, wherein said Poeciliopsis turrubarensis transposase, or variant thereof, recognizes a left UR sequence having at least 75% sequence identity with SEQ ID NO: 37, and a right ITR sequence having at least 75% sequence identity with SEQ ID NO: 38.
25. The composition according to any one of claims 10 to 14, wherein said Anthonomus grandis DR1754440 transposase, or variant thereof, recognizes a left ITR sequence having at least 75% sequence identity with SEQ ID NO: 71, and a right ITR sequence having at least 75% sequence identity with SEQ ID NO: 72.
26. The composition according to any one of claims 15 to 19, wherein said Anthonomus grandis DR1756049 transposase, or variant thereof, recognizes a left ITR sequence having at least 75% sequence identity with SEQ ID NO: 81, and a right ITR sequence having at least 75% sequence identity with SEQ ID NO: 82.
27. The composition according to any one of claims 20 to 22, wherein said Anthonomus grandis DR1754053 transposase, or variant thereof, recognizes a left ITR sequence having at least 75% sequence identity with SEQ ID NO: 73, and a right ITR sequence having at least 75% sequence identity with SEQ ID NO: 74.
28. The composition according to any one of claims 1 to 27, further comprising a RNA- guided nuclease or nickase.
29. The composition according to claim 28, wherein the transposase and the RNA- guided nuclease or nickase are fused together by a covalent linkage or a non- co valent linkage.
30. The composition according to claim 29, wherein the covalent linkage comprises a linker.
31. The composition according to claim 28, wherein the transposase and the RNA- guided nuclease or nickase are not fused together.
32. The composition according to any one of claims 1 to 31, further comprising at least one nucleic acid molecule comprising at least one transgene of interest to integrate into the genome of one or more cell.
33. An in vitro method for the integration of at least one transgene of interest into the genome of one or more cell, comprising contacting said one or more cell with the composition according to any one of claims 1 to 32, and at least one nucleic acid molecule comprising said at least one transgene of interest.
34. The in vitro method according to claim 33, for a targeted integration of the at least one transgene of interest into the genome of one or more cell, wherein the composition further comprises a RNA-guided nuclease or nickase enabling said targeted integration.
35. The composition according to any one of claims 1 to 32, which is a pharmaceutical composition.
36. The composition according to any one of claims 1 to 32, or claim 35, for use in a method for treating a genetic disease in a subject in need thereof.