Polypeptides with translocation activity
Patent Information
- Application Number
- JP2024515313
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-09
- Filing Date
- 2022-09-08
- Publication Date
- 2025-08-29
AI Technical Summary
Existing transposon systems for genome manipulation, such as Sleeping Beauty (SB), exhibit near-random integration into the genome, posing genotoxic risks and potential oncogenic transformation, necessitating the development of transposases with specificity for targeted integration to enhance safety in gene therapy applications.
Engineering transposases with specific amino acid substitutions at positions 187, 247, and 248 of the Sleeping Beauty transposase to acquire targeted integration into palindromic AT repeat sequences, reducing exon integration events by at least 25% and enhancing DNA flexibility at non-nucleosomal sites.
The engineered transposases demonstrate reduced integration into exons and transcriptional control regions, improving safety and utility in gene therapy by minimizing genotoxic risks and enhancing targeted integration efficiency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to polypeptides with transposition activity, in particular engineered polypeptides with transposition activity, nucleic acids encoding such polypeptides, vectors containing said nucleic acids, cells containing said nucleic acids, methods for using said polypeptides to integrate exogenous nucleic acids into a genome, said polypeptides for use in medicine and / or in some gene therapy methods; and pharmaceutical compositions comprising said polypeptides, nucleic acids, vectors or cells. In particular, the present invention relates to said polypeptides that have been engineered to acquire specificity of integration into the genome. [Background technology]
[0002] Polypeptides with transposase activity are versatile tools for therapeutic genome engineering. Transposons or transposable elements are (short) nucleic acid sequences with upstream and downstream terminal repeats. Active transposons encode enzymes that facilitate the excision and insertion of nucleic acids into target DNA sequences.
[0003] The Sleeping Beauty (SB, Tc1 / Mariner superfamily transposon) transposon is one of the known tools used in therapeutic genome engineering. This exemplary tool has been previously studied, and in particular, variants that exhibit hyperactivity have been developed that enhance the efficient insertion of transposons of various sizes into the nucleic acid of a cell or the insertion of DNA into the genome of a cell, thus allowing for more efficient transcription / translation than known transposons.
[0004] SB is a synthetic transposon reconstructed based on the sequence of a transposition-inactive element isolated from a fish genome. SB is the most thoroughly studied vertebrate transposon to date, and directs a whole range of genetic engineering applications, including transgenic cell line production, induced pluripotent stem cell (iPSC) reprogramming, phenotype-driven insertional mutation screening in cancer biology, germline gene transfer in experimental animals, and somatic cell gene therapy both ex vivo and in vivo. On a genome-wide scale, SB transposons exhibit a near-random integration profile with a slight bias towards integration into genes and upstream regulatory sequences in cultured mammalian cell lines. On a local scale, SB preferentially inserts into TA dinucleotides (e.g., Mos1) and shows additional target site preferences based on physical properties of DNA, including flexibility, A affinity, and symmetric patterns of hydrogen-bonding positions in the major groove of tDNA. However, random integration into genomes poses certain genotoxic risks in human applications. There is therefore a need for a transposon system in which integration is guaranteed and whose insertion into the genome does not pose or at least reduces the risks associated with oncogenic transformation.
[0005] The inventors unexpectedly discovered specific transposase variants that show specificity of integration into genomes, especially palindromic AT repeat target sequences. The inventors showed enhanced DNA flexibility of non-nucleosomal DNA-rich target sites by the discovered variants. The variants were further found to detarget exons of genes in the human genome as well as transcriptional regulatory regions, thus expanding the safety and utility of the variants in gene therapy applications. Summary of the Invention
[0006] In a first aspect (aspect 1a), the present invention provides a method for producing a medicament for the treatment of a pulmonary arthritis, comprising: α1 helix-α2 helix-β1 sheet-β2 sheet-β3 sheet-β4 sheet-β5 sheet-α3 helix-β6 sheet-η1-α4 helix-η2-α5 helix-α6 helix-α7 helix-α8 helix (where: (i) the naturally occurring H in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, L, M, F, P, R, S, T, W, K, E or Y, preferably V, P or T; or The naturally occurring F in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, L, M, P, R, S, T, W, K, H, E or Y, preferably V, P or T, or the naturally occurring Y in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, L, M, F, P, R, S, T, W, H, E or K, preferably V, P or T; or The naturally occurring L in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, M, F, P, R, S, T, W, H, Y, E or K, preferably V, P or T, or The naturally occurring S in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, L, M, F, P, R, T, W, H, Y, E or K, preferably V, P or T; and / or (ii) at least one naturally occurring Q, P, S or T in the β6 sheet connecting the β6 sheet and the α4 helix is replaced with R, S, C or A, preferably R or S; and / or (iii) at least one naturally occurring K, I, S, T, V, P or C in the β6 sheet connecting the β6 sheet and the α4 helix is replaced with R, I, C or V, preferably R). Provided is a polypeptide having transposase activity comprising a variant of a naturally occurring transposase that comprises a secondary structure element, wherein the variant of the naturally occurring transposase comprises at least one of substitutions (i), (ii) or (iii).
[0007] Alternatively, in aspect 1b, the present invention provides a polypeptide having transposition activity comprising or consisting of a transposase of the Tc1 / mariner superfamily, in which the amino acid positions corresponding to amino acid positions 248, 247 and / or 187 of sleeping beauty (SB) transposase in SEQ ID NO:1 are substituted with different amino acids, and the transposase acquires specificity for integration into a genome.
[0008] Gain in specificity means that the number of exon integration events due to transposition in the polypeptide is reduced by at least 25% compared to the number of exon integration events observed using SB of SEQ ID NO:1.
[0009] Optionally, the transposase is a sleeping beauty transposase or a variant thereof having at least 70% sequence identity to SEQ ID NO:1.
[0010] The polypeptides disclosed herein can, in certain embodiments, comprise a substitution at an amino acid position corresponding to amino acid 248 of SB transposase, where preferably the substitution is selected from the group consisting of K248R, K248S, K248V, K248I and K248C.
[0011] In certain embodiments, the polypeptide optionally comprises a substitution at an amino acid position corresponding to amino acid 247 of SB transposase, where preferably the substitution is selected from the group consisting of P247R, P247C, P247A and P247S.
[0012] The polypeptide may, in one embodiment, comprise a substitution at an amino acid position corresponding to amino acid 187 of SB transposase, wherein preferably the substitution is selected from the group consisting of H187A, H187N, H187C, H187Q, H187G, H187I, H187L, H187M, H187S, H187V, H187W, H187K, H187R, H187E, H187P and H187T.
[0013] In some embodiments, the polypeptide has the following substitutions: (i) H187V, H187P, H187R, H187E, H187T, H187S, H187A, H187N, H187C, H187Q, H187G, H187I, H187L, H187M, H187W, H187K, preferably H187P, H187V or H187T, more preferably H187V; and / or (ii) P247R, P247S, P247C or P247A; preferably P247R or P247S; and / or (iii) K248R, K248S, K248V, K248I or K248C; preferably K248R; or a substitution at the corresponding position of its variant may include one or more of:
[0014] Alternatively, the present invention relates in embodiment 1c to a polypeptide comprising SEQ ID NO:1 or a variant thereof having translocation activity and at least 70% sequence identity with SEQ ID NO:1, wherein SEQ ID NO:1 or its variant comprises the following substitutions: (i) H187A, H187N, H187C, H187Q, H187G, H187I, H187L, H187M, H187F, H187S, H187V, H187W, H187K, H187Y, H187R, H187E, H187P or H187T, preferably H187P, H187V or H187T, more preferably H187V; and / or (ii) P247R, P247C, P247A or P247S, preferably P247R or P247S and / or (iii) K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; or a substitution at the corresponding position of its variant Provide something that includes one or more of the above.
[0015] The polypeptide may have secondary structure elements as described above. It may be a polypeptide of embodiment 1a. Preferably, it shows a gain in specificity of integration into the genome compared to SB100X.
[0016] In a second aspect, the present invention relates to a nucleic acid comprising a nucleic acid sequence encoding a polypeptide of the first aspect.
[0017] In a third aspect, the present invention relates to a vector comprising the nucleic acid molecule of the second aspect.
[0018] In a fourth aspect, the present invention relates to a cell comprising the nucleic acid of the second aspect or the vector of the third aspect.
[0019] In a fifth aspect, the present invention provides an in vitro method for integrating an exogenous nucleic acid into the genome of a cell, comprising: a) providing an isolated cell; b) providing a cell with a polypeptide of the first aspect; and c) providing a cell with a nucleic acid containing an exogenous nucleic acid or a vector containing the same; The present invention relates to a method comprising the steps of:
[0020] In a sixth aspect, the present invention relates to an in vivo method of integrating an exogenous nucleic acid into the genome of a cell of a subject, comprising the step of administering to a subject a polypeptide, nucleic acid or vector according to any of the above aspects of the invention and a nucleic acid comprising the exogenous nucleic acid or a vector comprising said nucleic acid.
[0021] In another aspect, the invention relates to a polypeptide, a nucleic acid, a vector or a cell as described for use in medicine, in particular in gene therapy.
[0022] In a further aspect, the invention relates to a pharmaceutical composition comprising the described polypeptide, nucleic acid, vector or cell and a pharma- ceutically acceptable carrier, adjuvant or vehicle.
[0023] List of Figures In the following, the content of the figures contained in this specification is explained, in which context reference is also made to the detailed description of the invention given above and / or below. [Brief description of the drawings]
[0024] [Figure 1] Amino acid sequence alignment of exemplary transposases of the Tc1 / Mariner superfamily transposase. All amino acids corresponding to amino acid positions 187, 247, and 248 of SB are highlighted by boxes. Thus, FIG. 1 allows one skilled in the art to determine which amino acid positions correspond to amino acid positions 187, 247, and 248, respectively, in all other aligned transposases and can be replaced with different amino acids in the same manner as illustrated for SB. FIG. 1 only describes the alignment of transposases of selected Tc1 / Mariner superfamily transposases. Additional transposases of this superfamily can be added to this alignment to allow identification of amino acids corresponding to amino acid positions 187, 247, and 248 of SB in these transposases. The transposases shown belong to a group of transposases known to have transposase activity. The amino acid sequence of SB (SEQ ID NO: 1) was used as the reference amino acid sequence in this alignment.
[0025] [Diagram 2]Alignment of the amino acid sequences of Sleeping Beauty transposase and closely related transposases. All amino acids corresponding to amino acids 187, 247, and 248 of SB are highlighted by boxes. Thus, FIG. 1 allows one skilled in the art to determine which amino acid positions in all other aligned transposases correspond to amino acids 187, 247, and 248, respectively, and can be replaced with different amino acids in the same manner as illustrated for SB. FIG. 2 depicts an alignment of only a select number of transposition-active members of the Tc1 / mariner transposase superfamily. Additional transposases can be added to this alignment to allow identification of amino acids corresponding to amino acids 187, 247, and 248 of SB in these transposases. The secondary structure α1 helix-α2 helix-β1 sheet-β2 sheet-β3 sheet-β4 sheet-β5 sheet-α3 helix-β6 sheet-η1-α4 helix-η2-α5 helix-α6 helix-α7 helix-α8 helix belonging to the corresponding amino acid stretch is shown on top of the sequence alignment. This further allows the skilled person to determine the amino acid positions in all other aligned transposases that correspond to the specific amino acids in η1 that connect the β6 sheet and the α4 helix or that connect the β3 sheet and the β4 sheet. In a similar manner, further transposases of this superfamily can be added to this alignment to allow the identification of the amino acids that correspond to the positions in η1 that connect the β6 sheet and the α4 helix or that connect the β3 sheet and the β4 sheet. The amino acid sequence of SB (SEQ ID NO: 1) was used as the reference amino acid sequence in this alignment.
[0026] [Diagram 3]Transposition activity of Sleeping Beauty transposase mutants. Relative transposition activity of H187 (A), P247 (B) and K248 (C). Plasmids expressing transposase mutants were transiently co-transfected with a transposon donor plasmid (pT2B / puro) into HeL cells. Cells were selected for puromycin resistance and stained with methylene blue to identify viable cell colonies. Colony numbers were normalized to the SB100X positive control, where transposition efficiency was set to 100%. Inactive SB transposase (D3) was included as a negative control. Data are shown as mean ± SD, n=3 biological replicates. Differences in transposition activity are significant as determined by Student's t-test for the indicated mutants.
[0027] [Figure 4] Integration sites of transposase mutants. Sequence logos show majority rule consensus sequences at genomic insertion loci in 60 bp windows centered around the targeted TA dinucleotide. A value of 2 (log24) on the y-axis indicates the highest possible frequency. Pie charts show the percentage of insertions occurring at the SB-specific ATATATAT consensus motif. N values represent the number of uniquely mappable SB transposon insertions.
[0028] [Diagram 5]Representation of insertion sites in genomic features and chromatin-defined functional segments. (A) Integration loci of MLV, HIV, SB100X and their mutant derivatives were counted in gene-associated segments of the human genome. Numbers indicate fold change in insertion frequency increase (red) or decrease (blue) compared to a theoretical random control (set to 1). The dendrogram on the left is based on the mean frequency values of the rows. Statistical significance (Fisher's exact test) measured for values of K248R, P247R and H187V mutants versus SB100X is indicated by stars; * p<0.05, ** p<0.001, ns: non-significant. (B) Representation of insertion sites in functional genomic segments defined by epigenetic signal patterns. Color code, dendrogram and statistical significance as indicated for panel (A). Enrichment of simple repeat integrations catalyzed by H187V and K248R is shown in panel (C) and preference for integration into the ATATATAT, H187V variant is shown as panel (D). The overall percentage of insertions in these genomic regions is shown in (E).
[0029] [Figure 6] Insertion frequency in genomic safe harbors. (A) Numbers indicate the percentage representation of insertions in the safe harbor subcategories listed in the table below. Darker colors indicate lower insertion frequency in a safe harbor category, as opposed to the ideal of 100%. **p<0.001, Fisher's test compared to SB100X. (B) Overall representation of insertions in genomic safe harbors.
[0030] [Figure 7] Nucleosome occupancy at transposon insertion sites. Relative frequencies of insertions by SB100X and K248R, P247R and H187V mutants into sites associated with nucleosomes as determined by MNase-Seq data in human HepG2 cells. Frequencies are shown in windows extending 500 bp upstream and 500 bp downstream from the transposon integration site.
[0031] [Figure 8]H187, P247 and K248 transposase mutants preferentially target TA-rich genomic regions. Panel (A) shows that H187V and K248R insertions generally lack histone modifications. Exons, especially coding exons, are relatively deficient in TA dinucleotides (Figure 8B). However, exons are preferred target sites for SB transposition.
[0032] [Figure 9] Representation of insertion sites of additional SB100x mutants in genomic features. Integration loci H187P, H187R, H187E, H187S, H187T, P247S, P247A, P247C, P247S, K248C, K248I and K248V of MLV, HIV, SB100X and its mutant derivatives were counted in gene-related segments of the human genome. Numbers indicate fold increase (red) or decrease (blue) in insertion frequency compared to a theoretical random control (set to 1). Dendrogram on the left is based on row mean frequency values. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0033] Detailed Description of the Invention Before describing the present invention in detail below, it is to be understood that this invention is not limited to the particular methodology, protocols, and reagents described herein, as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention, which should be limited only by the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0034] Preferably, the terms used herein are as described in "A multilingual glossary of biotechnological terms: (IUPAC Recommendations)", Leuenberger, HGW, Nagel, B. and Klbl, H. eds. (1995), Helvetica Chimica Acta, CH-4010 Basel, Switzerland) and in "Pharmaceutical Substances: Syntheses, Patents, Applications" by Axel Kleemann and Jurgen Engel, Thieme Medical Publishing, 1999; "Merck Index: An Encyclopedia of Chemicals, Drugs, and Biologicals", edited by Susan Budavari et al., CRC Press, 1996, and in the United States Pharmacopeial 25 / National Formulary-20, 2001, published by the United States Pharmacopeial Convention, Inc., Rockville Md.
[0035] The practice of the present invention employs, unless otherwise indicated, conventional methods of chemistry, biochemistry, cell biology, immunology and recombinant DNA techniques which are described in the art.
[0036] Throughout this specification and the following claims, unless the context requires otherwise, the terms "comprises" and variations thereof, such as "comprises" and "comprising" refer to the inclusion of a stated integer or step or group of integers or steps, but not to the exclusion of any other integer or step or group of integers or steps. In the following description, the various aspects of the invention are defined in more detail. Each aspect thus defined may be combined with any one or more of the other aspects, unless clearly stated to the contrary. In particular, any feature indicated as optional, preferred or advantageous may be combined with any other feature or features indicated as optional, preferred or advantageous.
[0037] In the following, the elements of the present invention are described. These elements are described with specific embodiments. However, it should be understood that they can be combined in any manner and in any number to create further embodiments. The various described examples and preferred embodiments should not be interpreted as limiting the present invention to only the embodiments explicitly described. This description should be understood to support and encompass embodiments that combine the embodiments explicitly described with any number of the disclosed and / or preferred elements. Furthermore, any permutation and combination of all elements described herein should be interpreted as disclosed by this specification, unless the context requires otherwise.
[0038] definition We provide below definitions of terms commonly used in this specification, which have their respective defined and preferred meanings in each instance where they are used, as well as elsewhere in this specification.
[0039] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0040] As used herein, the term "transposase" refers to an enzyme that is a component of a functional nucleic acid-protein complex capable of transposition and mediates transposition. The term "transposase" also refers to an integrase from a retrotransposon or of retroviral origin. As used herein, a "transposition reaction" refers to a reaction in which a transposon is inserted into a target nucleic acid. The participants in the transposition reaction are a transposon and a transposase or integrase enzyme.
[0041] The terms "naturally occurring transposase" or "wild-type transposase" are used interchangeably herein as the unmodified amino acid sequence of a transposase that has been isolated from a naturally occurring species. Such amino acid sequences are readily available from EBI or NCBI. Sleeping Beauty is intended to be encompassed by the term "naturally occurring transposase."
[0042] As used herein, the term "transposition activity" refers to the activity of a transposase, which can be assessed in a transposition reaction. A suitable experimental setup can be described in the experimental section herein or can use the classical two-component transposition assay described in Ivics, 1997 Cell 91: 501-510.
[0043] The term "transposase" refers to a stretch of amino acids that has the transposition activity of a naturally occurring full-length transposase, e.g., at least 1% of the activity, at least 10% of the activity, at least 20% of the activity, at least 50% of the activity, or optimally 100% or more of the activity. The transposase may lack 1-10 N- and / or C-terminal amino acid sequences of a naturally occurring full-length transposase. Preferably, the transposase may lack the methionine at position M1 if it is contained in a fusion protein that contains another protein at the N-terminus of the transposase.
[0044] As used herein, the term "substitution" refers to the replacement of one nucleotide or amino acid with one nucleotide of a nucleic acid sequence and one amino acid of an amino acid sequence, respectively.
[0045] The term "consensus sequence" as used herein refers to the calculated order of the most frequent residues of nucleotides or amino acids found at each position in a sequence alignment between two or more sequences. It represents the result of a multiple sequence alignment, where related sequences are compared to each other and similar sequence motifs are calculated. Conserved sequence motifs are described as consensus sequences and show identical amino acids, i.e., amino acids that are identical between the compared sequences, conserved amino acids, i.e., amino acids that vary between the compared amino acid sequences but all of which belong to a functional or structural group of amino acids, such as polar or neutral, and variable amino acids, i.e., amino acids that do not show a clear relationship between the compared sequences. Figure 1 describes the alignment of sequences of transposases, including some transposition-active members of the Tc1 / mariner transposase superfamily, and below the alignment, the derived consensus sequence. The alignment algorithm used to generate the alignments presented is CLUSTAL Omega using the following settings: dealigned input sequences: no, mbed-like clustering tree: yes, mbed-like clustering iterations: yes, combined iteration count: default, highest tree iterations: default, highest hmm iterations: default. Amino acid conservation is highlighted with MView. Figure 2 presents the same alignment as in Figure 1, but with secondary structure information obtained for the catalytic domain of SB transposase highlighted in ESPript3.0 (PDB entry code: 5CR4).
[0046] The term "corresponding position" as used herein refers to an amino acid that aligns in an amino acid sequence alignment with the amino acid sequence of a naturally occurring full-length control transposase, preferably the full-length amino acid sequence of SB of SEQ ID NO:1. The position of the control transposase is determined from the first N-terminal amino acid. Thus, position 187 of SB in the full-length amino acid sequence of SEQ ID NO:1 is "H". Thus, the amino acid position corresponding to position 187 of SB in other transposases is the amino acid that aligns with the "H" of SB at position 187 of SEQ ID NO:1 in the alignment, as illustrated in Figures 1 and 2. Similarly, the amino acid position corresponding to position 247 of SB is aligned with the "P" of SB at position 247 of SEQ ID NO:1, and the position corresponding to position 248 of SB is aligned with the "K" of SB at position 248 of SEQ ID NO:1.
[0047] In the present invention, the "primary structure" of a protein or polypeptide is the sequence of amino acids in a polypeptide chain. The "secondary structure" of a protein is the general three-dimensional shape of a local segment of a protein. However, it does not describe the specific atomic positions in three-dimensional space that are considered to be the tertiary structure. In proteins, the secondary structure is determined by the hydrogen bonding pattern between the main chain amide groups and the carboxyl groups. The "tertiary structure" of a protein is the three-dimensional structure of the protein as determined by atomic coordinates. The "quaternary structure" is the arrangement of multiple folded or coiled protein or polypeptide molecules in a multi-subunit complex.
[0048] As used herein, the term "folding" or "protein folding" refers to the process by which a protein assumes its three-dimensional shape or structure, i.e., whereby the protein is directed to form a specific three-dimensional shape through non-covalent and / or covalent interactions, such as, but not limited to, hydrogen bonding, metal coordination, hydrophobic forces, van der Waals forces, pi-pi interactions, electrostatic effects and / or via intramolecular Cys bonds. Thus, the term "folded protein" refers to the three-dimensional shape of some or all of a protein, such as its secondary, tertiary or quaternary structure.
[0049] The term "beta strand" as used herein refers to a 5-10 amino acid long section in a polypeptide chain where the N-Cα-CN torsion angle in the backbone is approximately 120°. Beta strands in a protein sequence can be predicted by retrieving annotations from pdb files after (multiple) sequence alignment (e.g. using Clustal omega). Prediction of beta strands can be performed using publicly available software tools such as JPred. The start and end positions of the seven beta strands can be confirmed by further structural alignment of the PDB files using pyMol.
[0050] The terms "β-sheet" or "beta-sheet", as used interchangeably herein, refer to two β-strands that form a β-sheet. β-sheets are a common motif of ordered secondary structure in proteins. As peptide chains are oriented by their N- and C-termini, β-strands can also be said to be oriented. This is usually indicated in protein topology diagrams by an arrow pointing to the C-terminus. Adjacent β-strands can form hydrogen bonds in antiparallel, parallel or mixed configurations. In the antiparallel configuration, successive β-strands alternate orientation such that the N-terminus of one strand is adjacent to the C-terminus of the next. This is the configuration that provides the strongest interstrand stability since it allows interstrand hydrogen bonds between carbonyls and amines to be planar, which is the preferred orientation. β-structures are characterized by long, extended polypeptide chains. The amino acid composition of β-strands tends to favor hydrophobic (water-fearing) amino acid residues. The side chains of these residues tend to be less soluble in water than the more hydrophilic (water-loving) residues. β-structures tend to be found within the core structure of proteins, where hydrogen bonds between strands are protected from competing water molecules.
[0051] The terms "α-helix" or "alpha-helix", as used interchangeably herein, refer to a common motif in protein secondary structure, a right-handed helical conformation in which all main-chain NH hydrogens are bonded to the main-chain C=O of the amino acid located four residues earlier along the protein sequence. Among the local structure types of proteins, the α-helix is the most extreme, the most predictable from sequence, and the most frequent.
[0052] The terms "intervening region" or "loop," as used interchangeably herein, refer to the amino acid residues between two beta sheets or between a beta sheet and an alpha helix. Generally, the "intervening region" or "loop" has little or no structure, which provides flexibility.
[0053] When referring to a substitution within an amino acid sequence, the following nomenclature "XNo.Z" is used, meaning that amino acid "X" at a particular position "No." number in the original amino acid sequence is substituted with amino acid "Z", where "XNo.Z / X'No.'Z' / X"No."Z"" means that the original amino acids "X", "X'" and "X"" may be selectively substituted at position "No." with several different amino acids represented as "Z", "Z'" and "Z"", respectively. The amino acid positions are indicated with respect to a particular length of amino acid sequence, preferably the wild-type, full-length sequence of a particular transposase.
[0054] The terms "vector" or "expression vector" are used interchangeably and refer to a polynucleotide or a mixture of polynucleotides and proteins that can be introduced into a cell, preferably a mammalian cell, or that can introduce the collection of nucleic acids of the invention or a nucleic acid that is part of the collection of nucleic acids of the invention. Examples of vectors include, but are not limited to, plasmids, cosmids, phages, viruses or artificial chromosomes. In particular, vectors can be used to transport a promoter and a collection of nucleic acids or a nucleic acid that is part of the collection of nucleic acids of the invention into a suitable host cell. An expression vector can contain a "replicon" polynucleotide sequence that facilitates the autonomous replication of the expression vector in the host cell. Once inside the host cell, the expression vector can replicate independently or simultaneously with the host chromosomal DNA, producing several copies of the vector and its inserted DNA. When a replication-deficient expression vector is used - often for safety reasons - the vector cannot replicate and can simply direct the expression of the nucleic acid. Depending on the type of expression vector, the expression vector can express the neoantigen encoded by the nucleic acid only transiently, i.e., lost from the cell, or can be stable in the cell. An expression vector typically contains an expression cassette, ie, the essential elements that allow transcription of a nucleic acid into an mRNA molecule.
[0055] The term "viral vector" as used herein refers to a single-stranded or double-stranded nucleic acid sequence capable of assembling into an infectious viral particle. This nucleic acid sequence may be a complete or partial viral genome. In the latter case, the viral genome preferably contains one or more heterologous genes. For some viral particles, very short sequences of the viral genome are necessary to allow assembly of infectious viral particles. For example, only short (about 200 bp long) repeat sequences located at the 5' and 3' of the heterologous nucleic acid of a certain length (typically 4.5-5.3 kB for adeno-associated viruses) allow assembly of infectious adeno-associated viral particles. The minimum nucleic acid sequence for assembly of certain viruses is well known. The larger the viral genome and the smaller the minimum viral sequence for infectious viral particles, the larger the heterologous gene that can be inserted into the viral vector.
[0056] "Sequence similarity" refers to the percentage of amino acids that represent the same or conservative amino acid substitutions. The term "sequence identity" between two amino acid sequences refers to the percentage of amino acids that are identical between the two sequences. Alignment to determine sequence similarity, preferably sequence identity, can be performed with known tools, preferably using best sequence alignment, for example, using CLC main Workbench (CLC bio) or Align, using standard settings, preferably EMBOSS::needle, Matrix:Blosum62, Gap Open 10.0, Gap Extend 0.5. The identity percentage is determined with reference to the full length sequence used for comparison, and not simply to the most similar sequence or sequence stretch. Thus, an amino acid that shares 100% sequence identity with 50 consecutive amino acids of a 100 amino acid long sequence used for comparison only has 50% sequence identity (under the assumption that there are no additional amino acids that share any identity outside the 50 consecutive amino acids).
[0057] The term "pharmaceutical acceptable" as used herein preferably refers to the non-toxic nature of a substance that does not interact with the action of the active agent of the pharmaceutical composition. In particular, "pharmaceutical acceptable" means listed in the United States Pharmacopoeia, the European Pharmacopoeia or other generally recognized pharmacopoeias approved by a regulatory agency of the federal or state government for use in animals, more particularly in humans.
[0058] The term "carrier" refers to a natural or synthetic organic or inorganic ingredient with which the active ingredient is combined for the purpose of facilitating, enhancing or enabling application. According to the present invention, the term "carrier" also includes one or more compatible solid or liquid fillers, diluents, additives or encapsulating substances suitable for administration to a subject. Possible carrier substances (e.g., diluents) are, for example, sterile water, Ringer's solution, lactated Ringer's solution, physiological saline, bacteriostatic saline (e.g., saline containing 0.9% benzyl alcohol), phosphate-buffered saline (PBS), Hank's solution, fixed oils, polyalkylene glycols, hydrogenated naphthalenes and biocompatible lactide polymers, lactide / glycolide copolymers or polyoxyethylene / polyoxy-propylene copolymers. In one embodiment, the carrier is PBS. The resulting solution or suspension is preferably isotonic with the blood of the recipient. Suitable carriers and their formulations are described in great detail in Remington's Pharmaceutical Sciences, 17th ed., 1985, Mack Publishing Co.
[0059] The term "cell" is used herein to refer to a eukaryotic or prokaryotic cell, such as a bacterial cell, a yeast cell or a mammalian cell, preferably a human, mouse, rat, rabbit, dog, monkey or cat cell.
[0060] Aspects and Preferred Embodiments of the Invention The various aspects of the invention are now defined in more detail. Each aspect thus defined may be combined with any other aspect or aspects, unless clearly indicated to the contrary. In particular, any feature indicated as being preferred or advantageous may be combined with one or more other features indicated as being preferred or advantageous.
[0061] In the research leading to the present invention, it was surprisingly found that variants of polypeptides with transposase activity show a significant gain of genomic integration properties.
[0062] The present invention provides a polypeptide having transposition activity comprising or consisting of a transposase of the Tc1 / mariner superfamily, in which the amino acid positions corresponding to amino acid positions 248, 247 and / or 187 of the sleeping beauty (SB) transposase of SEQ ID NO: 1 are substituted with different amino acids, wherein the transposase has acquired specificity for integration into the genome. The polypeptide can be a sleeping beauty transposase or a variant thereof having at least 70% sequence identity to SEQ ID NO: 1.
[0063] "Gain of specificity" preferably means that the number of exon integration events due to transposition in the polypeptide is reduced by at least 10%, preferably at least 25%, compared to the number of exon integration events observed when using the SB of SEQ ID NO:1.
[0064] The polypeptide can, in certain embodiments, comprise a substitution at an amino acid position corresponding to amino acid 248 of SB transposase, where preferably the substitution is selected from the group consisting of K248R, K248S, K248V, K248I and K248C.
[0065] The polypeptide can, in certain embodiments, comprise a substitution at an amino acid position corresponding to amino acid 247 of SB transposase, where preferably the substitution is selected from the group consisting of P247R, P247C, P247A and P247S.
[0066] The polypeptide may, in one embodiment, comprise a substitution at an amino acid position corresponding to amino acid 187 of SB transposase, wherein preferably the substitution is selected from the group consisting of H187A, H187N, H187C, H187Q, H187G, H187I, H187L, H187M, H187S, H187V, H187W, H187K, H187R, H187E, H187P and H187T.
[0067] Polypeptides also preferably contain combinations of the described substitutions at these positions. Other mutations, such as those that enhance translocation activity (eg, as described below), can also be included.
[0068] The polypeptides of the invention may contain one or more of the following substitutions: (i) H187V, H187P, H187R, H187E, H187T or H187S, H187A, H187N, H187C, H187Q, H187G, H187I, H187L, H187M, H187S, H187W, H187K, preferably H187P, H187V or H187T, more preferably H187V; and / or (ii) P247R, P247S, P247C or P247A, preferably P247R or P247S and / or (iii) K248R, K248S, K248V, K248I or K248C; preferably K248R; or a substitution at the corresponding position of a variant thereof.
[0069] In one embodiment, the present invention also provides a method for the preparation of a method for treating a cancer cell comprising the steps of: α1 helix-α2 helix-β1 sheet-β2 sheet-β3 sheet-β4 sheet-β5 sheet-α3 helix-β6 sheet-η1-α4 helix-η2-α5 helix-α6 helix-α7 helix-α8 helix (where: (i) the naturally occurring H in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, L, M, F, P, R, S, T, W, K, Y or E, preferably V, P or T; or The naturally occurring F in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, L, M, P, R, S, T, W, K, H, E or Y, preferably V, P or T, or the naturally occurring Y in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, L, M, F, P, R, S, T, W, H, E or K, preferably V, P or T; or The naturally occurring L in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, M, F, P, R, S, T, W, H, Y, E or K, preferably V, P or T, or The naturally occurring S in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, L, M, F, P, R, T, W, H, Y, E or K, preferably V, P or T; and / or (ii) at least one naturally occurring Q, P, S or T in the β6 sheet connecting the β6 sheet and the α4 helix is replaced with R, S, C or A, preferably R or S; and / or (iii) at least one naturally occurring K, I, S, T, V, P or C in the β6 sheet connecting the β6 sheet and the α4 helix is preferably replaced with R, I, V or C) Provided is a polypeptide having transposition activity comprising, consisting essentially of, or consisting of a variant of a naturally occurring transposase having secondary structure elements, wherein the variant of the naturally occurring transposase comprises at least one of substitutions (i), (ii) or (iii).
[0070] The present invention is based on the surprising discovery that certain amino acid substitutions in transposases enhance DNA flexibility at target sites rich in non-nucleosomal DNA, leading to transposases that acquire specificity for integration into genomes, particularly palindromic AT repeat target sequences, detargeting exons and transcriptional regulatory regions of genes in the human genome. All these properties contribute to enhanced safety and practicality of the polypeptides of the present invention in gene therapy applications. These properties are highly favorable for transposases, since they are less likely to inactivate or mutate expressed genes. The improved integration pattern can be evaluated by the experiments described in Examples 3 and 7 herein and shown in Figures 5 and 9, respectively. Thus, if a polypeptide comprises or consists of a variant of SB of SEQ ID NO: 1, it is preferred that the number of exon integration events is reduced by at least 10%, better at least 25%, more preferably at least 40% or more preferably at least 50% with the polypeptides of the present invention, when compared to the number of exon integration events observed when using SB of SEQ ID NO: 1.
[0071] Preferably, the variants of the present invention retain the transposition activity of the full-length wild-type transposase, i.e., have at least 1%, at least 10%, at least 20%, at least 30%, at least 50%, at least 75%, optionally 100% or more of the transposition activity of the full-length wild-type transposase on which the variant is based, preferably 75%, optionally 100% of the transposition activity of SB of SEQ ID NO:1.
[0072] Figures 1 and 2 depict alignments of amino acid sequences of transposases of various degrees of relatedness to SB. Although the amino acid sequences differ considerably between at least some of the transposases included in the alignment in Figure 1 and SB, all transposases share similar secondary structures. They contain secondary structural elements such as α-helices and β-sheets in a certain order and length. Figures 1 and 2 show the location of α1 helix-α2 helix-β1 sheet-β2 sheet-β3 sheet-β4 sheet-β5 sheet-α3 helix-β6 sheet-η1-α4 helix-η2-α5 helix-α6 helix-α7 helix-α8 helix in these amino acid sequences. With respect to SB, the elements α1 helix-α2 helix-β1 sheet-β2 sheet-β3 sheet-β4 sheet-β5 sheet-α3 helix-β6 sheet-η1-α4 helix-η2-α5 helix-α6 helix-α7 helix-α8 helix span the following amino acids of SEQ ID NO:1: α1 helix: AA128-138 of SEQ ID NO:1; α2 helix: AA143-147 of SEQ ID NO:1; β1 sheet: AA149-158 of SEQ ID NO:1; β2 sheet: AA169-171 of SEQ ID NO:1; β3 sheet: AA173-176 of SEQ ID NO:1; β4 sheet: AA191-199 of SEQ ID NO:1; β5 sheet: AA202-208 of SEQ ID NO:1; α3 helix: AA215-233 of SEQ ID NO:1; β6 sheet: AA240-241 of SEQ ID NO:1; η1: AA247-250 of SEQ ID NO:1; α4 helix: AA252-260 of SEQ ID NO:1; η2: AA273-275 of SEQ ID NO:1; α5 helix: AA278-291 of SEQ ID NO:1; α6 helix: AA297-309 of SEQ ID NO:1; α7 helix: AA313-321 of SEQ ID NO:1; α8 helix: AA323-332 of sequence number 1.
[0073] Thus, the α1 helix-α2 helix-β1 sheet-β2 sheet-β3 sheet-β4 sheet-β5 sheet-α3 helix-β6 sheet-η1-α4 helix-η2-α5 helix-α6 helix-α7 helix-α8 helix of other transposases begins and ends at amino acid positions corresponding to the amino acid positions shown above in SB. This allows one skilled in the art to determine the secondary structure element: α1 helix-α2 helix-β1 sheet-β2 sheet-β3 sheet-β4 sheet-β5 sheet-α3 helix-β6 sheet-η1-α4 helix-η2-α5 helix-α6 helix-α7 helix-α8 helix in a transposase and to determine accordingly the amino acids in the loop connecting the β3 and β4 sheets and in η1 connecting the β6 sheet and the α4 helix in a transposase.
[0074] Furthermore, one skilled in the art can analyze the three-dimensional structure and amino acid sequence in relation to its structure in other transposases and predict the alignment of amino acids based on the three-dimensional structure. Furthermore, computer programs can help determine the secondary structure of a protein or polypeptide of interest, i.e., any other transposase of interest, for example. One method is based on homology modeling. Usually, two polypeptides or proteins with more than 30% sequence identity or more than 40% similarity often have similar structural topologies. The development of protein structure databases (PDBs) has expanded the predictability of secondary structures, including the number of possible folds within a polypeptide or protein structure. Additional methods for predicting secondary structures of proteins and polypeptides are known in the art and include, for example, threading, profile analysis, and progressive docking. Thus, one of skill in the art can easily add additional transposases of this superfamily to the alignment, allowing identification of amino acids in these transposases that correspond to amino acids 187, 247, and 248 of SB and / or amino acids in the loop connecting the β3 and β4 sheets of a transposase or amino acids in the η1 connecting the β6 sheet and α4 helix of a transposase.
[0075] The loop connecting the β3 and β4 sheets of a transposase may contain in its naturally occurring sequence H, F, L, S or Y. In this case, it is preferred to replace these amino acids with other amino acids that may alter the protein-DNA interactions of the transposase.
[0076] If the naturally occurring amino acid in the loop connecting the β3 and β4 sheets is H, it is preferred that this amino acid is replaced by V, A, N, C, Q, G, I, L, M, F, P, R, S, T, W, K, Y or E, preferably V, I, L, M, P, S or T, more preferably V, I, L, T or P, most preferably V, P or T.
[0077] If the naturally occurring amino acid in the loop connecting the β3 and β4 sheets is F, it is preferred that this amino acid is replaced by V, A, N, C, Q, G, I, L, M, P, R, S, T, W, K or Y, preferably V, I, L, M, P, S or T, more preferably V, I, L, T or P, most preferably V, P or T.
[0078] If the naturally occurring amino acid in the loop connecting the β3 and β4 sheets is Y, then it is preferred that this amino acid is replaced by V, A, N, C, Q, G, I, L, M, F, P, R, S, T, W or K, preferably V, I, L, M, P, S or T, more preferably V, I, L, T or P, most preferably V, P or T.
[0079] If the naturally occurring amino acid in the loop connecting the β3 and β4 sheets is L, it is preferred that this amino acid is replaced by V, A, N, C, Q, G, I, M, F, P, R, S, T, W, H, Y or K, preferably V, P or T.
[0080] If the naturally occurring amino acid in the loop connecting the β3 and β4 sheets is S, it is preferred that this amino acid is replaced by V, A, N, C, Q, G, I, L, M, F, P, R, T, W, H, Y or K, preferably V, P or T.
[0081] The η1, which connects the β6 sheet and the α4 helix of some transposases, may contain the naturally occurring sequences K, I, S, T, V, P or C (position 248 of SB) or Q, P, S or T (position 247 of SB). In this case, it is preferred to replace these amino acids with other amino acids that may alter the protein-DNA interactions of the transposase.
[0082] If the naturally occurring amino acid in η1 connecting the β6 sheet and the α4 helix is at least one naturally occurring Q, P, S or T, it is preferred to replace this amino acid with R, C, S or A, preferably R or S.
[0083] If the naturally occurring amino acid in η1 connecting the β6 sheet and the α4 helix is at least one naturally occurring K, I, S, T, V, P or C, then it is preferred to replace this amino acid with R, V, I, C, A, P or Q, preferably I, C or R, most preferably R.
[0084] In a further preferred embodiment, both the naturally occurring Q, P, S or T and K, I, S, T, V, P or C in η1 connecting the β6 sheet and the α4 helix are substituted as described above, preferably both are substituted with R.
[0085] In further embodiments or other aspects of the invention, substitutions may also occur in variants of naturally occurring transposases. The term "variant of a naturally occurring transposase" refers to an amino acid sequence having at least 70% amino acid sequence identity with the amino acid sequence of a naturally occurring transposase, preferably one of the transposases of SEQ ID NO: 1 to 17. More preferably, a variant of a transposase further modified by one or more substitutions as described above has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity with the amino acid sequence of a naturally occurring transposase, preferably one of the transposases of SEQ ID NO: 1 to 17, most preferably the SB transposase of SEQ ID NO: 1. If the substitution in the present invention is in the context of a variant or naturally occurring transposase, the degree of identity is determined in the absence of the substitution in the present invention. Thus, a naturally occurring transposase variant has 70% amino acid sequence identity to SEQ ID NO:1 and may further have substitutions (i), (ii) and / or (iii) above.
[0086] When the number of exon integration events is compared to the number of exon integration events observed when using a transposase of SEQ ID NO: 1 to 17, respectively, it is preferably reduced by, for example, at least 25%, more preferably at least 40% or more preferably at least 50%, in a polypeptide of the present invention based on the amino acid sequence of SEQ ID NO: 1 to 17, if the polypeptide of the present invention comprises or consists of a variant of a transposase of SEQ ID NO: 1 to 17.
[0087] In a preferred embodiment, the polypeptide of the invention comprises, consists of or essentially consists of a naturally occurring full-length transposase, preferably a naturally occurring full-length transposase of SEQ ID NO: 1-17 or a variant thereof. In this preferred embodiment, the control amino acid sequence for the determination of the variant is the full-length sequence. Thus, in this preferred embodiment of the polypeptide of the invention, the variant comprises, consists of or essentially consists of an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity with the amino acid sequence of a naturally occurring transposase, preferably a transposase of SEQ ID NO: 1-17, most preferably transposase SB of SEQ ID NO: 1. Similarly, the amino acid sequence identity of a variant of a naturally occurring transposase is determined in the absence of the substitutions of the invention of (i), (ii) and / or (iii).
[0088] In each of the above examples, the variants of the present invention retain the transposition activity of naturally occurring or wild-type transposase. The amino acid changes may affect the producibility and / or stability of the resulting transposase or may affect other activities, such as transposition activity. Substitutions that enhance the transposition activity of transposases, particularly SB, are described further below and previously in WO2009 / 003671. It is particularly preferred that such substitutions are present in addition to the above (i), (ii) and / or (iii) substitutions.
[0089] In certain embodiments of the first aspect of the invention, the naturally occurring transposase is selected from the group consisting of Sleeping Beauty (SB) (SEQ ID NO:1), Tdr1 (SEQ ID NO:2), ZB (SEQ ID NO:3), FP (SEQ ID NO:4), Passport (SEQ ID NO:5), TCB2 (SEQ ID NO:6), S (SEQ ID NO:7), Quetzal (SEQ ID NO:8), Paris (SEQ ID NO:9), Tc1 (SEQ ID NO:10), Minos (SEQ ID NO:11), Uhu (SEQ ID NO:12), Bari (SEQ ID NO:13), Tc3 (SEQ ID NO:14), Impala (SEQ ID NO:15), Himar (SEQ ID NO:16) or Mos1 (SEQ ID NO:17), preferably the transposase is SB (SEQ ID NO:1), ZB (SEQ ID NO:3), FP (SEQ ID NO:4), Passport (SEQ ID NO:5) or Minos (SEQ ID NO:11). In a most preferred embodiment, the transposase is SB (SEQ ID NO:1).
[0090] In certain preferred embodiments or further aspects of the invention, the polypeptide of the invention comprises, consists of or consists essentially of a transposase of SEQ ID NO: 1 to 17 or a variant thereof comprising at least one of the substitutions (i), (ii) or (iii) above.
[0091] Transposases closely related to SB not only share secondary structure elements arranged in the order shown above, but also share significant amino acid similarity or identity in certain key regions involved in DNA-protein interactions. Thus, the substitutions of the present invention in this subgroup of transposases closely related to SB can be characterized by the consensus amino acid sequence around amino acid 187 of SB. This consensus sequence is X1X2X3X 12 X4X5GX6 (SEQ ID NO: 18), where the variables have the meanings shown below. Another consensus amino acid sequence around amino acid positions 247 and 248 of SB is X7X8DX9X 10 X 13 X 11 X 14H (SEQ ID NO:19), with the variables having the meanings defined below. Certain transposases closely related to SB may comprise an amino acid sequence that fulfills the consensus sequence of SEQ ID NO:18 or SEQ ID NO:19 or both. Preferably, transposases closely related to SB comprise an amino acid sequence that fulfills both consensus sequences. Such transposases are highly similar to SB in both the region around amino acid 187 and around amino acid positions 247 and 248 of SB. Thus, substitutions (i), (ii) and / or (iii) of the invention preferably include transposases that comprise two consecutive amino acid sequences that fulfill the consensus of SEQ ID NO:18 and SEQ ID NO:19, respectively. Thus, in one embodiment, the polypeptide of the first aspect of the invention or alternatively the other aspect of the invention is a polypeptide having transposition activity that comprises, consists essentially of or consists of a variant of a naturally occurring transposase, (i) a first amino acid stretch: X1X2X3X 12 X4X5GX6 (SEQ ID NO: 18) (where: X1 is T, F, P or R; X2 is I, T, N, R, K, Q or V; X3 is K, N, L, R or Q, preferably K; X4 is G, N, P, A or Q, preferably G; X5 is G, K, A or absent, preferably X5 is G; X6 is S or absent, and X 12 is F, H, Y, L, or S in the wild-type protein, and X 12 X in naturally occurring transposases 12 when unsubstituted or substituted with A, N, C, Q, G, I, L, M, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X 12 X in naturally occurring transposases 12is unsubstituted or substituted by A, N, C, Q, G, I, L, M, F, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X 12 X in naturally occurring transposases 12 is unsubstituted or substituted by A, N, C, Q, G, I, L, M, F, P, R, S, T, V, W, E or K, preferably V, P or T; X 12 X in naturally occurring transposases 12 when unsubstituted or substituted with A, N, C, Q, G, I, F, M, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X 12 X in naturally occurring transposases 12 and is unsubstituted or substituted with A, N, C, Q, G, I, L, F, M, P, R, T, V, W, K, E or Y, preferably V, P or T; and / or (ii) a second amino acid stretch: X7X8DX9X 10 X 13 X 11 X 14 H (SEQ ID NO: 19) (where: X7 is Q, L, M, H or Y, and X8 is M, Q or H; X9 is N, H or G; X 10 is A or D, X 11 is absent, V, T or K (preferably absent), and X 13 is Q, P, S or T in the naturally occurring transposase, unsubstituted or substituted with R, S, C or A, preferably R or S; and / or X 14is K, I, S, T, V, P or C in the naturally occurring transposase, unsubstituted or substituted with R, I, C, V, A, P or Q, preferably R; where X 12 , X 13 and X 14 At least one of the following is substituted: The present invention relates to a polypeptide comprising:
[0092] In a preferred embodiment, X1 is T, X2 is V or I, preferably V; X3 is K or Q, preferably K; X4 is G, The X5 is a G. X6 is S, and X 12 is F, H, Y, L, or S in the wild-type protein, and X 12 X in naturally occurring transposases 12 when unsubstituted or substituted with A, N, C, Q, G, I, L, M, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X 12 X in naturally occurring transposases 12 is unsubstituted or substituted by A, N, C, Q, G, I, L, M, F, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X 12 X in naturally occurring transposases 12 is unsubstituted or substituted by A, N, C, Q, G, I, L, M, F, P, R, S, T, V, W, E or K, preferably V, P or T; X 12 X in naturally occurring transposases 12 when unsubstituted or substituted with A, N, C, Q, G, I, F, M, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X12 X in naturally occurring transposases 12 is unsubstituted or substituted with A, N, C, Q, G, I, L, F, M, P, R, T, V, W, K, E or Y, preferably V, P or T; and / or X7 is Q, X8 is M, Q or H; X9 is N, H or G; X 10 is A or D, X 11 is non-existent, X 13 is P in the naturally occurring transposase and is unsubstituted or substituted with R, S, C or A, preferably R or S; and / or X 14 is K in the naturally occurring transposase, unsubstituted or substituted with R, I, C, V, A, P or Q, preferably R, and Where X 12 , X 13 and X 14 At least one of is substituted.
[0093] If there is only one substitution in a polypeptide of the invention, it is preferably not H187F, H187Y or K248A.
[0094] When any of these substitutions are present in a variant of a transposase, the sequence identity of the variant to the respective wild-type transposase sequence is similarly determined prior to introduction of substitutions (i), (ii) and / or (iii).
[0095] In certain embodiments, the polypeptides of the invention comprise SEQ ID NO:1 or a variant thereof having translocation activity and having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO:1, wherein SEQ ID NO:1 or its variant comprises one or more of the following substitutions: (i) H187A, H187N, H187C, H187Q, H187G, H187I, H187L, H187M, H187F, H187S, H187V, H187W, H187K, H187Y, H187R, H187E, H187P or H187T, preferably H187P, H187V or H187T, more preferably H187V; and / or (ii) P247R, P247C, P247A or P247S; preferably P247R or P247S and / or (iii) K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; or a substitution at the corresponding position of a variant thereof.
[0096] In a preferred embodiment, the polypeptide of the present invention comprises SEQ ID NO:1 or a variant thereof having transposition activity and having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with SEQ ID NO:1, wherein SEQ ID NO:1 or its variant comprises a substitution that enhances its transposition activity, wherein said substitution is preferably one or more of the above substitutions. The enhanced transposition activity can be an enhancement compared to the activity of wtSB or SB of SEQ ID NO:1. As mentioned above, various naturally occurring transposase variants have been described. For example, WO2009 / 003671 describes transposases, particularly SB variants, that are hyperactive, i.e. have enhanced transposition activity compared to wild-type transposase, particularly SB. It is preferred that the substitutions of the present invention, including a substitution or group of substitutions that enhance one or more properties of the transposase, particularly the transposition activity, are introduced into the transposase variant. Thus, in some embodiments, the polypeptides of the invention further comprise at least one of the following substitutions or groups of substitutions: (1)K14R, K13D, K13A, K30R, K33A, T83A, I100L, R115H, R143L, R147E, A205K / H207V / K208R / D210E, H207V / K208R / D210E, R214D / K215A / E216V / N217Q;M243Q, E267D, T314N and / or G317E;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (2)K14R / / R214D / K215A / E216V / N217Q;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (3)K33A / R115H / R214D / K215A / E216V / N217Q / M243H;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (4)K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (5)K13D / K33A / T83N / H207V / K208R / D210E / M243Q;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (6)K13A / K33A / R214D / K215A / E216V / N217Q;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (7)K33A / T83N / R214D / K215NE216V / N217Q / / G317E;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (8)K14R / T83A / M243Q;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (9)K14R / T83A / I100L / M243Q;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (10)K14R / T83A / R143L / M243Q;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (11)K14R / T83A / R147E / M243Q;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (12)K14R / T83A / M243Q / E267D;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (13)K14R / T83A / M243Q / T314N;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (14)K14R / K30R / I110L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (15)K14R / K30R / R143L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (16)K14R / K30R / R147E / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> (17)K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / E267D;<h2 style=";text-align:left;direction:ltr"> (18)K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H / T314N; (19)K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / G317E; (20)K14R / K33A / R115H / R214D / K215A / E216V / N217Q / M243H; (21)K14R / K30R / R147E / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / T314N; (22)K14R / K30R / R143U / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / E267D; (23)K14R / K30R / R143L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / T314N; (24)K14R / K30R / R143L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / G317E; (25)K14R / K33A / R115H / R143L / R214D / K215A / E216V / N217Q / M243H; (26)K14R / K33A / R115H / R147E / / R214D / K215A / E216V / N217Q / M243H; (27)K14R / K33A / R115H / R214D / K215A / E216V / N217Q / / M243H / E267D; (28)K14R / K33A / R115H / R214D / K215A / E216V / N217Q / / M243H / T314N; (29)K14R / K33A / R115H / R214D / K215A / E216V / N217Q / / M243H / G317E; (30)K14R / T83A / M243Q / G317E; or (31)K13A / K33A / T83N / R214D / K215A / E216V / N217Q; or a variant thereof at the corresponding position. Preferably, in addition to the substitutions of the invention, the substitutions or substitutions include at least one of the following: Tdr1 (SEQ ID NO:2), ZB (SEQ ID NO:3), FP (SEQ ID NO:4), Passport (SEQ ID NO:5), TCB2 (SEQ ID NO:6), S (SEQ ID NO:7), Quetzal (SEQ ID NO:8), Paris (SEQ ID NO:9), Tc1 (SEQ ID NO:10), Minos (SEQ ID NO:11), Uhu (SEQ ID NO:12), Bari (SEQ ID NO:13), Tc3 (SEQ ID NO:14), Impala (SEQ ID NO:15), Him a naturally occurring transposase selected from the group consisting of ZB (SEQ ID NO: 3), FP (SEQ ID NO: 4), Passport (SEQ ID NO: 5) or Minos (SEQ ID NO: 11)) at the corresponding position of the transposase, preferably (i) K248R, H187A, H187N, H187C, H187Q, H187G, H187I, H187L, H187M, H187F, H187S, H187V, H187W or H187Y, preferably H187I, H187V or H187L, more preferably H187V; and / or (ii) P247R; and / or (iii) K248A or K248S; preferably K248R; or Tdr1 (SEQ ID NO:2), ZB (SEQ ID NO:3), FP (SEQ ID NO:4), Passport (SEQ ID NO:5), TCB2 (SEQ ID NO:6), S (SEQ ID NO:7), Quetzal (SEQ ID NO:8), Pa Substitution at the corresponding position in ris (SEQ ID NO:9), Tc1 (SEQ ID NO:10), Minos (SEQ ID NO:11), Uhu (SEQ ID NO:12), Bari (SEQ ID NO:13), Tc3 (SEQ ID NO:14), Impala (SEQ ID NO:15), Himar (SEQ ID NO:16) or Mos1 (SEQ ID NO:17) (preferably the transposase is SB (SEQ ID NO:1), ZB (SEQ ID NO:3), FP (SEQ ID NO:4), Passport (SEQ ID NO:5) or Minos (SEQ ID NO:11)).
[0097] The polypeptide of the first aspect of the invention most preferably comprises or consists of an amino acid sequence based on SEQ ID NO:1, with the following amino acids substituted within the amino acid sequence of SEQ ID NO:1: (1) K14R and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C, preferably K248R; (2) K13D and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (3) K13A and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (4) K30R and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (5) K33A and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (6) T83A and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (7) I100L and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (8) R115H and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (9) R143L and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (10) R147E and H187I, H187V, H187E, H187T, H187P or H187L, more preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; which enhances the transposition activity (11) A205K / H207V / K208R / D210E and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; which enhances the transposition activity (12) H207V / K208R / D210E and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R, which enhances the transposition activity; (13) R214D / K215A / E216V / N217Q and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R or P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R, which enhances the transposition activity; (14) M243Q and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (15) E267D and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (16) T314N and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (17) G317E and H187I, H187V, H187E, H187T, H187P or H187L, more preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (18) A combination of two or more of the substitutions or groups of substitutions that enhance the translocation activity as shown under (1) to (18) is combined with the substitutions H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; Preferred combinations of the substitutions that enhance the transposition activity shown under (1) to (18) and the substitutions of the present invention are as follows: (19) K14R / / R214D / K215A / E216V / N217Q and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; which enhances the transposition activity (20) K33A / R115H / R214D / K215A / E216V / N217Q / M243H and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (21) K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S; preferably K248R; (22) K13D / K33A / T83N / H207V / K208R / D210E / M243Q and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; which enhances the transposition activity (23) K13A / K33A / R214D / K215A / E216V / N217Q and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S; preferably K248R; which enhances transposition activity (24) K33A / T83N / R214D / K215NE216V / N217Q / / G317E and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (25) K14R / T83A / M243Q and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R, which enhances the transposition activity; (26) K14R / T83A / I100L / M243Q and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; which enhances the transposition activity (27) K14R / T83A / R143L / M243Q and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R, which enhances the transposition activity; (28) K14R / T83A / R147E / M243Q and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R or P247S, P247, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; which enhances the transposition activity (29) K14R / T83A / M243Q / E267D and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R, which enhances the transposition activity; (30) K14R / T83A / M243Q / T314N and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; which enhances the transposition activity (31) K14R / K30R / I110L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (32) K14R / K30R / R143L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (33) K14R / K30R / R147E / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (34) K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / E267D and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (35) K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H / T314N and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (36) K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / G317E and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (37) K14R / K33A / R115H / R214D / K215A / E216V / N217Q / M243H and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (38) K14R / K30R / R147E / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / T314N and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (39) K14R / K30R / R143U / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / E267D and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (40) K14R / K30R / R143L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / T314N and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (41) K14R / K30R / R143L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / G317E and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (42) K14R / K33A / R115H / R143L / R214D / K215A / E216V / N217Q / M243H and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; which enhances the transposition activity (43) K14R / K33A / R115H / R147E / / R214D / K215A / E216V / N217Q / M243H and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; (44) K14R / K33A / R115H / R214D / K215A / E216V / N217Q / / M243H / E267D and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; which enhances the transposition activity (45) K14R / K33A / R115H / R214D / K215A / E216V / N217Q / / M243H / T314N and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; which enhances the transposition activity (46) K14R / K33A / R115H / R214D / K215A / E216V / N217Q / / M243H / G317E and H187I, H187V, H187E, H187T, H187P or H187L, more preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; which enhances the transposition activity (47) K14R / T83A / M243Q / G317E and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R, which enhances the transposition activity; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; or (48) K13A / K33A / T83N / R214D / K215A / E216V / N217Q and H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R, which enhances transposition activity.
[0098] In a preferred embodiment, the polypeptide of the first aspect comprises or consists of the amino acid sequence of SEQ ID NO: 1, wherein the following amino acids have been substituted to enhance translocation activity: K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N, and comprises the following substitutions of the invention H187I, H187V, H187E, H187T, H187P or H187L, preferably H187V and / or P247R, P247S, P247C or P247A, preferably P247R; and / or K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R. In particular, a preferred embodiment of the polypeptide of the first aspect of the invention has an amino acid sequence comprising or consisting of SEQ ID NO: 1, wherein the following amino acids have been substituted: substitutions: K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and H187V to enhance transposition activity; or K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and P247R to enhance transposition activity; or K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and K248R to enhance transposition activity; or K14 enhancements R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and P247R and K248R for enhancing transposition activity; preferably K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and H187V to enhance transposition activity; or K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and H187P to enhance transposition activity; or K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and H187T for enhancing transposition activity; or K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and H187S for enhancing transposition activity; or K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and H187E for enhancing transposition activity; or K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and P247R to enhance transposition activity; or K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and P247S for enhancing transposition activity; or K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and K248R to enhance transposition activity; K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and K248C to enhance transposition activity; K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and K248I to enhance transposition activity; K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N and K248V for enhancing transposition activity.
[0099] Preferably, the sequence identity with SEQ ID NO:1 is at least 70%, at least 80%, at least 90% or at least 95%, although of course other substitutions may be included.
[0100] Particularly preferred transposases of the present invention that combine enhanced transposition activity with the properties of the substitutions of the present invention are as follows (amino acid substitutions that enhance transposition activity are highlighted in bold + underline, amino acid substitutions of the present invention are highlighted in bold, underline and italics): H187V and K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N (SEQ ID NO: 20): [ka] P247R and K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N (SEQ ID NO:21): [ka] K248R and K14R, K33A, R115H, R214D / K215A / E216V / N217Q, M243H and T314N (SEQ ID NO: 22): [ka]
[0101] The polypeptides (transposase variants) of the present invention have several advantages over prior art approaches, most notably exhibiting superior genomic integration properties.
[0102] In a second aspect, the present invention relates to a nucleic acid comprising a nucleic acid sequence encoding a polypeptide of the first aspect.
[0103] In certain embodiments of the second aspect of the invention, the (encoding) nucleic acid sequence is operably linked to at least one transcriptional control unit.
[0104] In certain embodiments of the nucleic acid of the invention, the nucleic acid comprises at least one open reading frame. In other embodiments, the nucleic acid further comprises at least a regulatory region of the gene. Preferably, the regulatory region is a transcriptional regulatory region, more particularly, the regulatory region is selected from the group consisting of promoters, enhancers, silencers, locus control regions and boundary elements. The nucleic acid of the invention is typically a ribonucleic acid, including mRNA, DNA, cDNA, chromosomal DNA, extrachromosomal DNA, plasmid DNA, viral DNA or RNA, and also recombinant viral vectors. In certain embodiments of the nucleic acid of the invention, the nucleic acid is DNA or RNA, and in other preferred embodiments, the nucleic acid is part of a plasmid or recombinant viral vector. The nucleic acid of the invention is preferably selected from any nucleic acid sequence that codes for the amino acid sequence of the polypeptide of the invention. Thus, all nucleic acid variants code for the mutant polypeptide variants of the invention as described above, including nucleic acid variants with different nucleotide sequences due to the degeneracy of the genetic code. Nucleic acid variant nucleotide sequences that result in improved expression of the encoded fusion protein, especially in the host organism of choice, are preferred. Tables of appropriate adjustment of nucleic acid sequences for specific transcription / translation machinery of the host cell are known to those skilled in the art. In general, it is preferred to adapt the G / C content of the nucleotide sequence to the specific host cell conditions. For expression in human cells, it is preferred to increase the G / C content of the maximum G / C content (encoding each peptide variant of the invention) by at least 10%, more preferably at least 20%, 30%, 50%, 70%, more preferably 90%. The production and purification of such nucleic acids and / or derivatives are usually carried out by standard techniques.
[0105] These sequence variants preferably lead to a protein selected from the polypeptides of the invention having translocation activity of the invention or variants of said polypeptides comprising the amino acid sequences of SEQ ID NO: 1 to 17, preferably SEQ ID NO: 1, in which at least one amino acid has been substituted compared to the native nucleic acid sequence of the corresponding polypeptide having translocation activity. Thus, the nucleic acid sequences of the invention code for modified (non-natural) variants of said polypeptides. Furthermore, promoters or other expression control regions can be controlled to be linked to the nucleic acids encoding the polypeptides of the invention in order to control the polypeptide / protein expression in a quantitative or tissue-specific manner.
[0106] In a third aspect, the present invention relates to a vector comprising the nucleic acid molecule of the second aspect. As already mentioned above, the nucleic acid encoding the polypeptide of the present invention may be RNA or DNA. Similarly, the nucleic acid of the present invention encoding the polypeptide or transposon of the present invention may be a linear fragment or a circular, isolated fragment or may be incorporated into a vector, preferably as a plasmid or recombinant viral DNA.
[0107] In a fourth aspect, the present invention relates to a cell comprising the nucleic acid of the second aspect or the vector of the third aspect.
[0108] In some embodiments, the cell is from an animal, preferably from a vertebrate, preferably from a mammal, for example, from a human, preferably selected from the group consisting of fish, bird or mammal.The cell can be an immune cell, for example, a T cell, for example, a primary T cell.It can also be a tumor cell.
[0109] In a fifth aspect, the present invention provides an in vitro method for integrating an exogenous nucleic acid into the genome of a cell, comprising: a) providing an isolated cell; b) providing the cell with a polypeptide of the invention; and c) providing a cell with a nucleic acid containing an exogenous nucleic acid or a vector containing the same; The present invention relates to a method comprising the steps of:
[0110] In one embodiment of the in vitro method, the polypeptide is provided to a cell by introducing a nucleic acid or vector of the invention into the cell. In this case, the exogenous nucleic acid can be contained within the nucleic acid or can be a nucleic acid that comprises the exogenous nucleic acid, as desired. The vector containing the exogenous nucleic acid can alternatively be provided to the cell separately.
[0111] The polypeptide can also be provided to the cells in the form of a polypeptide, for example by adding the polypeptide to the cell culture medium.
[0112] In one embodiment, the nucleic acid of the invention, the vector of the invention and / or the exogenous nucleic acid is provided to the cell using a method selected from the group consisting of electroporation, microinjection, lipofection, etc. Electroporation has proven to be particularly advantageous for transduction of, for example, primary cells.
[0113] The cells are from an animal, preferably a vertebrate, preferably selected from the group consisting of a fish, a bird or a mammal, preferably a mammal, such as a human.
[0114] In a sixth aspect, the present invention relates to an in vivo method of integrating an exogenous nucleic acid into the genome of a cell of a subject, comprising the step of administering to a subject a polypeptide, nucleic acid or vector according to any of the above aspects of the invention and a nucleic acid comprising the exogenous nucleic acid or a vector comprising said nucleic acid.
[0115] In some embodiments, the exogenous nucleic acid is contained in a nucleic acid or vector of the invention, or may be administered separately.
[0116] In one embodiment, the nucleic acid or vector of the invention or the exogenous nucleic acid is provided to the cell using a method selected from the group consisting of electroporation, microinjection, lipoprotein particles, virus-like particles.
[0117] The exogenous nucleic acid can be DNA or RNA.
[0118] In another aspect, the invention relates to a polypeptide, a nucleic acid, a vector or a cell as described for use in medicine, in particular in gene therapy.
[0119] In certain embodiments, gene therapy includes, but is not limited to, autologous or allogeneic T cell therapy, gene therapy targeting any cell type in the blood, hematopoietic stem cell therapy, liver gene therapy, central nervous system gene therapy, eye gene therapy, muscle gene therapy, skin gene therapy, and / or gene therapy for the treatment of cancer.
[0120] In general, the therapeutic applications of the present invention are multifaceted, therefore the polypeptides, nucleic acids, vectors, cells and in particular the methods for integrating an exogenous nucleic acid into the genome of a cell according to the invention may also be useful in therapeutic applications, i.e. gene therapy applications, where the above are used to stably integrate a therapeutic nucleic acid ("nucleic acid of therapeutic interest"), e.g. a gene (nucleic acid of therapeutic interest), into the genome of a target cell. This is also for tumor vaccination or for the treatment of infectious diseases from pathogens, e.g. leprosy, tetanus, whooping cough, typhoid, paratyphoid, cholera, plague, tuberculosis, meningitis, bacterial pneumonia, anthrax, botulism, bacterial dysentery, diarrhea, food poisoning, syphilis, gastroenteritis, trench fever, influenza, scarlet fever, diphtheria, gonorrhea, toxic shock syndrome, Lyme disease, typhus, listeriosis, peptic ulcer and legionellosis; e.g. acquired immune deficiency syndrome, adenoviridae infections, alphavirus infections, arbovirus infections, Borna disease, and the like. disease), Bunyaviridae infections, Caliciviridae infections, Chickenpox, Genital warts, Coronaviridae infections, Coxsackievirus infections, Cytomegalovirus infections, Dengue, DNA virus infections, Ecthyma, Contagious diseases, Encephalitis, Arbovirus, Epstein-Barr virus infections, Erythema infectiosum, Hantavirus infections, Hemorrhagic fever, Virus, Hepatitis, Virus, Human, Herpes simplex, Shingles, Shingles oticus, Herpesviridae infections, Infectious mononucleosis, Avian influenza, Influenza, Human, Lassa fever, Measles, Contagious diseases Treatment of viral infections resulting in molluscum contagiosum, mumps, paramyxoviridae infections, papatasii fever, polyomavirus infections, rabies, respiratory syncytial virus infections, Rift Valley fever, RNA virus infections, rubella, slow-onset viral disease, smallpox, subacute sclerosing panencephalitis, tumor virus infections, warts, West Nile fever, viral diseases, yellow fever; vaccine therapy may also be of interest, incorporating antigens into antigen-presenting cells, such as specific tumor antigens, e.g., MAGE-1, for pathological antigens for the treatment of protozoan infections resulting in malaria. A wide range of therapeutic nucleic acids can be delivered using the methods of the present invention. Therapeutic nucleic acids of the present invention include genes that replace defective genes in target host cells, such as those responsible for disease states based on genetic defects; genes with therapeutic utility in the treatment of cancer, and the like.
[0121] In a further aspect, the invention relates to a pharmaceutical composition comprising the described polypeptide, nucleic acid, vector or cell and a pharma- ceutically acceptable carrier, adjuvant or vehicle.
[0122] The pharmaceutical composition may further comprise one or more carriers and / or additives, all of which are preferably pharma- ceutically acceptable. According to the present invention, the pharmaceutical composition comprises an effective amount of an active agent, such as a polypeptide, a nucleic acid, a vector or a cell described herein, to produce a desired reaction or a desired effect. The pharmaceutical composition of the present invention is preferably sterile. The pharmaceutical composition may be provided in a homogenous dosage form and may be prepared in a manner known per se. The pharmaceutical composition of the present invention may be, for example, in the form of a solution or suspension.
[0123] Pharmaceutically acceptable carriers, adjuvants or vehicles that may be used in the compositions of the present invention include, but are not limited to, ion exchangers, alumina, aluminum stearate, lecithin, serum proteins such as human serum albumin, buffers such as phosphate, glycine, sorbic acid, potassium sorbate, partial glyceride mixtures of saturated vegetable fatty acids, water, salts or electrolytes such as protamine sulfate, disodium hydrogen phosphate, potassium hydrogen phosphate, sodium chloride, zinc salts, colloidal silica, magnesium trisilicate, polyvinylpyrrolidone, cellulose-based substances, polyethylene glycol, sodium carboxymethylcellulose, polyacrylic acid, wax, polyethylene-polyoxypropylene-block polymers, polyethylene glycol and wool fat. The pharmaceutical compositions of the present invention may be administered orally, parenterally, by inhalation spray, topically, rectally, nasally, bucally, vaginally or via an implanted reservoir. The term parenteral as used herein includes subcutaneous, intravenous, intramuscular, intra-articular, intra-synovial, intrasternal, intrathecal, intrahepatic, intralesional and intracranial injection or infusion techniques. Preferably, the pharmaceutical composition is administered orally, intraperitoneally or intravenously. Sterile injectable forms of the pharmaceutical composition of the present invention include aqueous or oily suspensions. These suspensions are formulated by techniques known in the art using suitable dispersing or wetting agents and suspending agents. Sterile injectable preparations can also be sterile injectable solutions or suspensions in non-toxic, parenterally acceptable diluents or solvents, such as solutions in 1,3-butanediol. Among the acceptable vehicles and solvents that can be used are water, Ringer's solution and isotonic sodium chloride solution. In addition, sterile, fixed oils are commonly used as solvents or suspending media.
[0124] The pharmaceutical composition of the present invention is preferably used to treat diseases, in particular diseases caused by genetic defects, such as cystic fibrosis, hypercholesterolemia, hemophilia, e.g. A, B, C or XIII, immunodeficiencies including HIV, Huntington's disease, α-antitrypsin deficiency, as well as colon cancer, melanoma, kidney cancer, lymphoma, acute myeloid leukemia (AML), acute lymphoid leukemia (ALL), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), gastrointestinal tumors, lung cancer, glioma, thyroid cancer, breast cancer, prostate tumors, hepatocellular carcinoma, various virus-induced tumors, such as papillomavirus-induced tumors (e.g. cervical cancer), adenocarcinoma, herpes virus-induced tumors (e.g. Burkitt's lymphoma, EBV-induced B-cell lymphoma), hepatitis B-induced tumors. and the like.
[0125] There are a variety of alternative techniques and methodologies available to those skilled in the art that similarly enable successful practice of the intended invention. EXAMPLES
[0126] Example 1: Saturation mutagenesis of H187, P247, and K248 in Sleeping Beauty transposase identifies mutants with altered transposition efficiency To assess the relative effects of single amino acid substitutions at positions 187, 247, and 248 on transposition, the SB100X transposase was subjected to saturation mutagenesis by incorporation of all possible amino acids by site-directed PCR mutagenesis. All constructs encoding mutant SB100X transposases showed protein expression levels comparable to SB100X by Western blot analysis.
[0127] The mutants were then assessed for their relative transposition activity to SB100X by application of a cell-based transposition assay in human cells fine-tuned to yield one transposon integration per cell. Briefly, a transposon-bearing donor plasmid marked with a puromycin (puro) resistance gene was co-transfected with a helper plasmid encoding SB100X, an inactive E279D transposase (D3) or a mutant variant of the SB100X transposase.
[0128] In contrast to the H187 mutation, the vast majority of P247 mutations were completely inactive or showed severe reductions in both global transposition (Figure 3B) and excision (data not shown) activities. P247A was the most active mutant, with 109% transposition activity compared to SB100X. Of note, the relative transposition activity of P247A exceeded the calculated relative excision activity. This discrepancy between excision and integration activity may be due to the fact that both values were determined by two independent assays. The second most active mutant was P247S, which retained 73% of the transposon integration activity compared to SB100X, followed by P247R (27%, p=0.017), P247C (7%, p=0.002) and P247K (3%, p=0.003) (Figure 3B).
[0129] Similar to P247, amino acid exchange at position 248 strongly reduced transposition in nearly all 19 mutants (Figure 3C). K248R was the most active mutant to date, with 77% integration activity compared to SB100X. Three additional K248 mutants, including K248C (4%, p=0.003), K248I (4%, p=0.003), and K248V (9%, p=0.003), showed lower levels of transposition activity compared to SB100X. As recently published, some of the K248 mutants showed uncoupling of transposon excision and integration activities (data not shown).
[0130] Example 2: Target site preference alteration of H187, P247 and K248 transposase mutants To investigate whether the mutations have an effect on target site selection, we used some of the above SB mutants to generate a transposon insertion site library in human HepG2 cells and compared the local attributes as well as the genome-wide distribution of these insertion sites with those generated by SB100X. For position 187, we selected six mutants (H187R, H187E, H187S, H187T, H187V, H187P) that had a significant effect on transposition efficiency (Figure 3A), while for positions 247 and 248, we selected all mutants that showed measurable transposition activity (P247A, P247R, P247C, P247S, P247K and K248C, K248I, K248V, K248R, respectively) (Figures 3B and C). Transposon integration sites were visualized by SeqLogo analysis, which not only reports the consensus sequence but also provides information on the overall sequence conservation (measured in bits) and the relative frequency of the nucleotide at each position (measured as symbol height ordered by relative frequency within the logo). The analysis confirmed that the highly preferred TA target site dinucleotide is embedded in A / T-rich DNA in all mutants, as shown previously for SB (Figure 4). The SB100X transposase encodes an 8-bp palindromic AT repeat sequence ATA centered on the actual TA target dinucleotide. T.A.TAT. Because SB integrates almost exclusively into TA, the central TA approaches a maximum score of 2 bits in the logo. The SB100X control logo also showed a relatively strong overall conservation of the A base at the -3 position of the consensus and the matching T at the +3 position, with an overall score of 0.6 bits (Figure 4). These integration preferences appeared similar in all transposase mutants (Figure 4 and data not shown); however, there were notable differences. First, the P247S and P247R mutants showed slight deviations of the consensus target site [AGATATCT for P247S and ATTATATAAT for P247R (Figure 4); note that the palindromic nature of the consensus was maintained]. Second, conservation of the alternating AT bases upstream and downstream of the 8 bp consensus was more evident in mutants H187P, H187V and K248R, manifesting itself as a "shoulder" adjacent to the 8 bp consensus (Figure 4). Finally, the conservation of A at position -3 and T at position +3 of the consensus dramatically increased by about 1.5 bits in the H187P, H187V and K248R mutants (Figure 4). These findings indicate that the amino acid substitutions at positions 187, 247 and 248 of the SB transposase do indeed have an effect on the target site selection properties of each mutant transposase, with the most striking modification being a more pronounced overall preference for A / T richness of tDNA and a stronger conservation of palindromic bases within the consensus sequence. Thus, the data indicate that some of the mutants have become more constrained in their target site selection properties. This was clearly confirmed by the analysis of the overall frequency of transposon insertions into the 8bp ATATATAT sequence.
[0131] While SB100X and P247S and P247R mutants only targeted this particular sequence 2-3% of the time, some other mutants showed significantly higher integration frequencies into this motif (18% for H187P, 21% for H187V and 39% for K248R, Figure 4). Taken together, certain amino acid substitutions at positions 187, 247 and 248 of the SB transposase result in a highly preferential integration phenotype into AT repeats.
[0132] Example 3: Alteration of genome-wide distribution of insertions catalyzed by H187, P247 and K248 transposase mutants The relative frequency of integration into genomic features including genes and non-genic regions, oncogenes, exons, introns, 5'- and 3'-UTRs, and upstream and downstream sequence flanking genes within a 10 kb window was determined for computer-generated random data sets. In addition to comparing insertions into these genomic features produced by SB100X and its variants, the analysis also included MLV gammaretroviral and HIV lentiviral integration sites. Both of these viral systems are versatile vectors in gene therapy. As previously established, SB transposon insertions show only a slight bias towards genes and their flanking regions, in contrast to MLV and HIV insertions, which are enriched in loci adjacent to transcription start sites (TSSs) and within actively transcribed genes, respectively (Figure 5). Of these three gene vector systems, the overall insertion frequency of SB is closest to the expected random distribution (Figure 5).
[0133] We next selected a small subset of mutants; namely, the P247R, H187V, and K248R mutants, to see whether they produced genome-wide distribution profiles that were appreciably different from that of SB100X. The data shown in Figure 5A are significant. In particular, the H187V and K248R mutants show a significant depletion of insertions into genes. The most striking differences are within exons, including 5'- and 3'-UTRs and coding exons. Importantly, transposon integrations by H187V and K248R into these genomic regions are depleted not only compared to SB100X, but also compared to a random dataset. The overall percentage of insertions into these genomic regions is shown in Figure 5E. The most dramatic change in integration frequency is the four-fold reduction seen with K248R compared to SB100X in coding exons (p<0.001) (Figure 5A). The preference for integration into AT repeats by our variants, as well as the avoidance of integration into exons, indicates that these insertions are enriched in repetitive DNA. In fact, a dramatic increase in simple repeats is detected for integrations catalyzed by H187V and K248R; enrichment in this compartment of the genome is approximately 12-fold (p<0.001) over the random data set for H187V and approximately 4-fold (p<0.001) over SB100X and approximately 21-fold (p<0.001) over the random data set for K248R and approximately 7-fold (p<0.001) over SB100X (Figure 5C). Consistent with the preference for integration into ATATATAT, the H187V variant shows a 24-fold and approximately 5-fold enrichment over random and SB100X, respectively, in TA-rich simple repeats (p<0.001), whereas K248R-mediated insertions are enriched in TA-rich simple repeats. The enrichment was approximately 43-fold and 8-fold higher than random and SB100X, respectively (Figure 5D).
[0134] We next profiled the insertions produced by SB100X and P247R, H187V and K248R mutants in functional genomic segments. These segments are defined by the occurrence of epigenetic signal patterns that computationally cluster together to include various functional compartments of the human genome. We used a 25-state chromatin model of human HepG2 cells. As described above, we included MLV and HIV insertions in the analysis. First, consistent with previous observations, SB transposon insertions show only a slight bias towards promoters, TSSs, enhancers and transcriptional regulatory regions with open chromatin structures, in contrast to MLV and HIV insertions, which show the highest enrichment in promoter regions and transcribed regions containing TSSs, respectively (Figure 5B). Second, dramatic depletion in enhancers, promoters containing TSSs, and open regulatory regions was seen in integrations catalyzed by H187V and K248R; the most striking changes were the approximately 50-fold, 3-fold, and 5-fold depletion in promoters containing TSSs by H187V over MLV, SB100X (p<0.001), and random data sets, respectively, and the approximately 14-fold, 4-fold, and 3-fold depletion in enhancers by K248R over MLV, SB100X (p<0.001), and random data sets, respectively (Figure 5B). Collectively, the data indicate that the P247R, H187V, and K248R mutants detarget an alternative proportion of transposon integrations away from exons and transcriptional regulatory regions, including promoters and enhancers.
[0135] Example 4: Enrichment of insertions into the genomic safe harbor with H187V, P247R and K248R transposase mutants Integration of therapeutic gene constructs into safe sites in the human genome prevents the risk of insertional mutagenesis and concomitant oncogenesis in gene therapy. Genomic "safe harbors" (GSHs) are regions of the human genome that can accommodate predictable expression of newly integrated DNA without deleterious effects on the host cell or organism. GSHs can be bioinformatically assigned to chromosomal sites or regions if they meet the following criteria: (i) no overlap with transcription units, (ii) at least 50 kb away from the 5' end of any gene, iii) at least 300 kb away from cancer-related genes, and (iv) regions outside microRNA genes and (v) ultraconserved elements (UCEs). We have previously established that the SB transposon system has a much more favorable insertion profile than MLV- and HIV-based viral integration systems in terms of insertion frequency into GSHs.
[0136] The above data show that some of our SB transposase mutants significantly detarget insertions from exons of genes as well as transcriptional regulatory regions, essentially implying that the majority of insertions catalyzed by these enzymes are located in GSH. We analyzed the insertion site data set with P247R, H187V and K248R transposase variants for the relative frequency of integration into GSH, and included MLV and HIV insertions in the analysis as described above (Figure 6). Figure 6A shows three clusters in the dendrogram: the first one is represented by MLV and HIV, the second by SB100X, P247R and H187V, and the third by K248R and random. There is a gradual shift towards a random-like distribution of insertions by SB100X transposase and its mutants. For example, only 17% of HIV insertions are extragenic, whereas this value is 41% for SB100X, 42% for P247R, 44% for H187V, 46% for K248R and 49% for the random data set (Figure 6A). In agreement with these observations, the analysis confirmed an increase in the frequency of insertions into GSH with the simultaneous application of all five GSH criteria. Among the mutants tested, the K248R transposase variant approached randomness the most: overall, 29% of all K248R insertions and 34% of all random insertions fall into GSH (Figure 6B). Taken together, the analysis predicts 1) that the SB system is safer than MLV and HIV-based vectors, hence a favorable transgene insertion profile, and 2) an enhanced safety of the P247R, H187V and K248R transposase variants in therapeutic gene transfer in human cells.
[0137] Example 5: H187V and K248R transposase variants avoid nucleosomal DNA for integration High-density integration profiling of Hermes transposons in yeast confirmed a strong correlation between integration sites and nucleosome-free chromatin. Furthermore, recent evidence indicates that Tc1 / mariner transposons preferentially integrate at internucleosomal linker regions. The insertion dataset was mapped with respect to nucleosome occupancy as determined by MNase-Seq data. As shown previously, transposon insertions mediated by SB100X are only underrepresented in nucleosomal DNA (Figure 7). Notably, among the three transposase mutants tested, the H187V and K248R variants showed a dramatic decrease in the correlation between transposon integration sites and nucleosome occupancy, not only compared to the random control but also compared to the SB100X dataset (Figure 7). Thus, the H187V and K248R mutations dramatically increase the phenotype of SB transposase; i.e., nucleosomal DNA avoidance for transposon integration.
[0138] Example 6: Molecular characteristics associated with preferential genomic target sites of H187, P247 and K248 transposase mutants The above data confirm that H187V and K248R transposase variants detarget exons and transcriptional regulatory regions of genes and avoid nucleosomal DNA for integration. However, these data do not necessarily shed light on the causal relationship between these independent observations. For example, exons tend to associate with nucleosomes; therefore, the depletion of integration in exonic sequences by H187V and K248R may simply reflect the avoidance of nucleosomal DNA by these transposase variants. However, transcriptional regulatory regions, including enhancers and TSSs, were clearly depleted in nucleosomes; nevertheless, we found that H187V and K248R detarget these genomic regions, and nucleosome occupancy by itself cannot fully explain why H187V and K248R integration is depleted in exons as well as in regulatory regions. As genomic segments are defined by chromatin marks, the question arose as to whether it is the local chromatin structure or the underlying primary DNA sequence that regulates the integration frequency in these segments. Consistent with previous findings, SB100X transposase has a near random insertion profile in human cells, with a slight bias towards euchromatin marks (including H3K4me1, H3K27ac, H3K36me3 and H3K29me2) (Figure 8A). We detected significant (p<0.001) depletion of chromatin marks by H187V and K248R variants compared to SB100X insertions (Figure 8C); however, depletion is not restricted to euchromatin marks. Rather, H187V and K248R insertions are generally depleted in histone modifications, regardless of whether the marks are transcriptionally active or repressive chromatin (Figure 8A); therefore, we considered differential interactions of H187V and K248R transposase variants with chromatin compared to SB100X unlikely. The above findings led us to hypothesize that DNA sequence composition is a major determinant of preferred genomic targeting by the H187, P247, and K248 transposase mutants: exons, especially coding exons, are relatively deficient in TA dinucleotides, which are preferred target sites for SB transposition (Figure 8B).Given the enhanced preference for the ATATATAT sequence motif by some of our mutants, especially H187V and K248R, we indicate that the detargeting of exonic sequences by these mutants is driven by the rare occurrence of this sequence motif in exonic sequences. In fact, ATATATAT is remarkably underrepresented in exons, especially in coding exons (Figure 8B), thereby suggesting that primary DNA sequence composition is the major determinant driving the avoidance of exonic sequences by the H187V and K248R insertions. Similarly, open control regions, TSSs and enhancers are relatively TA-poor compared to the average base composition of the human genome (Figure 8C). As seen in coding exons, the ATATATAT sequence motif, which is preferentially integrated by the H187V and K248R mutants, is remarkably underrepresented in these three genomic segments (Figure 8C). The data indicate that a major determinant of detargeting of exons and transcriptional regulatory regions by the H187V and K248R mutants is the low availability of the highly preferred ATATATAT sequence in these genomic regions.
[0139] Example 7: Genome-wide distribution alteration of insertions catalyzed by H187P, H187R, H187E, H187S, H187T, P247S, P247A, P247C, P247S, K248C, K248I and K248V transposase mutants As in Example 3, the relative frequency of integration of genomic features including genic and nongenic regions, oncogenes, exons, introns, 5'- and 3'-UTRs, and upstream and downstream sequence flanking genes within a 10 kb window was determined for computer-generated random data sets. Similarly, MLV gammaretroviral and HIV lentiviral integration sites were included in the analysis.
[0140] This time, we analyzed a further subset of mutants: H187P, H187R, H187E, H187S, H187T, P247S, P247A, P247C, P247S, K248C, K248I and K248V to see whether they also produced genome-wide distribution profiles that were appreciably different from SB100X (Figure 9). All mutants tested, except for P247A and P247C, showed more or less pronounced depletion of insertions into genes. Similarly, the largest differences were found within exons, including 5'- and 3'-UTRs and coding exons. The most dramatic change in integration frequency was the 4-fold reduction seen with K248I and K248C compared to SB100X in coding exons (Figure 9). Overall, these two mutants exerted a similar effect to the K248R mutant tested in Example 3 (Figure 5A). Interestingly, the P247A variant appears to be a clear exception from the other variants, as it actually increased the frequency of integration into exons and transcriptional regulatory regions (Figure 9), thus leading not only to a specificity gain of the insertion, but also to a specificity gain in the opposite direction compared to the other variants.
Claims
1. A polypeptide having transposition activity comprising or consisting of a transposase of the Tc1 / mariner superfamily, in which the amino acid positions corresponding to amino acid positions 248, 247, and / or 187 of the sleeping beauty (SB) transposase of SEQ ID NO: 1 are substituted with different amino acids, wherein the transposase has acquired specificity for integration into the genome.
2. The polypeptide of claim 1, wherein gain of specificity means a reduction of at least 25% in the number of exon integration events due to transposition in the polypeptide compared to the number of exon integration events observed when using the SB of sequence number 1.
3. The polypeptide of claim 1, wherein the transposase is a sleeping beauty transposase or a variant thereof having at least 70% sequence identity to SEQ ID NO:
1.
4. containing a substitution at an amino acid position corresponding to amino acid 248 of the SB transposase, wherein preferably the substitution is selected from the group consisting of K248R, K248S, K248V, K248I and K248C. The polypeptide of claim 1.
5. containing a substitution at an amino acid position corresponding to amino acid 247 of the SB transposase; wherein preferably the substitution is selected from the group consisting of P247R, P247C, P247A and P247S. The polypeptide of claim 1.
6. containing a substitution at an amino acid position corresponding to amino acid 187 of the SB transposase; wherein preferably said substitutions are selected from the group consisting of H187A, H187N, H187C, H187Q, H187G, H187I, H187L, H187M, H187S, H187V, H187W, H187K, H187R, H187E, H187P, H187T and H187S. The polypeptide of claim 1.
7. 10. The polypeptide of claim 1, comprising one or more of the following substitutions: (i) H187V, H187P, H187R, H187E, H187T or H187S, H187A, H187N, H187C, H187Q, H187G, H187I, H187L, H187M, H187S, H187W, H187K, preferably H187P, H187V or H187T, more preferably H187V; and / or (ii) P247R, P247S, P247C or P247A, preferably P247R or P247S; and / or (iii) K248R, K248S, K248V, K248I or K248C; preferably K248R; or a substitution at the corresponding position in its variant.
8. 2. The polypeptide of claim 1, comprising the following secondary structure elements: α1 helix-α2 helix-β1 sheet-β2 sheet-β3 sheet-β4 sheet-β5 sheet-α3 helix-β6 sheet-η1-α4 helix-η2-α5 helix-α6 helix-α7 helix-α8 helix.
9. The polypeptide of claim 1, wherein the substitution enhances the translocation activity of the polypeptide.
10. The polypeptide of claim 1 comprising SEQ ID NO: 20, 21 or 22.
11. A polypeptide having transposition activity comprising or consisting of a variant of a naturally occurring transposase having the following secondary structure elements: α1 helix - α2 helix - β1 sheet - β2 sheet - β3 sheet - β4 sheet - β5 sheet - α3 helix - β6 sheet - η1-α4 helix - η2-α5 helix - α6 helix - α7 helix - α8 helix (where, (i) a naturally occurring H in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, L, M, F, P, R, S, T, W, K, E or Y, preferably V, P or T; or the naturally occurring F in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, L, M, P, R, S, T, W, K, H, E or Y, preferably V, P or T; or the naturally occurring Y in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, L, M, F, P, R, S, T, W, H, E or K, preferably V, P or T; or the naturally occurring L in the loop connecting the β3 and β4 sheets is replaced by V, A, N, C, Q, G, I, M, F, P, R, S, T, W, H, Y, E or K, preferably V, P or T, or the naturally occurring S in the loop connecting the β3 and β4 sheets is replaced with V, A, N, C, Q, G, I, L, M, F, P, R, T, W, H, Y, E or K, preferably V, P or T; and / or (ii) at least one naturally occurring Q, P, S, or T in the β6 sheet connecting the β6 sheet and the α4 helix is replaced with R, S, C, or A, preferably R or S; and / or (iii) at least one naturally occurring K, I, S, T, V, P, or C in the β6 sheet connecting the β6 sheet and the α4 helix is replaced with R, I, C, or V, preferably R; wherein the naturally occurring transposase variant comprises at least one of substitutions (i), (ii) or (iii).
12. 12. The polypeptide of claim 11, wherein the naturally occurring transposase is selected from the group consisting of Sleeping Beauty (SEQ ID NO: 1), Tdr1 (SEQ ID NO: 2), ZB (SEQ ID NO: 3), FP (SEQ ID NO: 4), Passport (SEQ ID NO: 5), TCB2 (SEQ ID NO: 6), S (SEQ ID NO: 7), Quetzal (SEQ ID NO: 8), Paris (SEQ ID NO: 9), Tc1 (SEQ ID NO: 10), Minos (SEQ ID NO: 11), Uhu (SEQ ID NO: 12), Bari (SEQ ID NO: 13), Tc3 (SEQ ID NO: 14), Impala (SEQ ID NO: 15), Himar (SEQ ID NO: 16) or Mos1 (SEQ ID NO: 17), preferably the transposase is SB (SEQ ID NO: 1).
13. (i) a first amino acid stretch: X 1 X 2 X 3 X 12 X 4 X 5 GX 6 (SEQ ID NO: 18) (where, X 1 is T, F, P or R, X 2 is I, T, N, R, K, Q or V; X 3 is K, N, L, R or Q, preferably K; X 4 is G, N, P, A or Q, preferably G; X 5 is G, K, A or absent, preferably X 5 is G, X 6 is S or absent, and X 12 is F, H, Y, L, or S in the wild-type protein, and X 12 X in naturally occurring transposases 12 is unsubstituted or substituted with A, N, C, Q, G, I, L, M, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X 12 X in naturally occurring transposases 12 is unsubstituted or substituted with A, N, C, Q, G, I, L, M, F, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X 12 X in naturally occurring transposases 12 is unsubstituted or substituted with A, N, C, Q, G, I, L, M, F, P, R, S, T, V, W, E or K, preferably V, P or T; X 12 X in naturally occurring transposases 12 is unsubstituted or substituted with A, N, C, Q, G, I, F, M, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X 12 X in naturally occurring transposases 12 wherein the substituted group is unsubstituted or substituted with A, N, C, Q, G, I, L, F, M, P, R, T, V, W, K, E or Y, preferably V, P or T; and / or (ii) a second amino acid stretch: X 7 X 8 DX 9 X 10 X 13 X 11 X 14 H (SEQ ID NO: 19) (where, X 7 is Q, L, M, H or Y, and X 8 is M, Q or H, X 9 is N, H or G, X 10 is A or D, X 11 is V, T or K or absent, preferably absent, and X 13 is Q, P, S, or T in the naturally occurring transposase, unsubstituted or substituted with R, S, C, or A, preferably R or S; and / or X 14 is K, I, S, T, V, P, or C in the naturally occurring transposase, and is unsubstituted or substituted with R, A, P, or Q, I, C, or V, preferably R; Here, X 12 , X 13 and X 14 at least one of The polypeptide of claim 11, comprising:
14. X 1 is T, X 2 is V or I, preferably V; X 3 is K or Q, preferably K; X 4 is G, X 5 is G, X 6 is S, and X 12 is F, H, Y, L, or S in the wild-type protein, and X 12 X in naturally occurring transposases 12 is unsubstituted or substituted with A, N, C, Q, G, I, L, M, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X 12 X in naturally occurring transposases 12 is unsubstituted or substituted with A, N, C, Q, G, I, L, M, F, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X 12 X in naturally occurring transposases 12 is unsubstituted or substituted with A, N, C, Q, G, I, L, M, F, P, R, S, T, V, W, E or K, preferably V, P or T; X 12 X in naturally occurring transposases 12 is unsubstituted or substituted with A, N, C, Q, G, I, F, M, P, R, S, T, V, W, K, E or Y, preferably V, P or T; X 12 X in naturally occurring transposases 12 is unsubstituted or substituted with A, N, C, Q, G, I, L, F, M, P, R, T, V, W, K, E or Y, preferably V, P or T; and / or X 7 is Q, X 8 is M, Q or H, X 9 is N, H or G, X 10 is A or D, X 11 is absent, and T, V, or K X 13 is P in the naturally occurring transposase and is unsubstituted or substituted with R, S, C or A, preferably R or S; and / or X 14 is K in the naturally occurring transposase and is unsubstituted or substituted with R, I, C or V, A, P or Q, preferably R, and Here, X 12 , X 13 and X 14 wherein at least one of: The polypeptide of claim 13.
15. SEQ ID NO: 1 or a variant thereof having transposition activity and having at least 70% sequence identity with SEQ ID NO: 1, wherein SEQ ID NO: 1 or its variant contains one or more of the following substitutions: (i) H187A, H187N, H187C, H187Q, H187G, H187I, H187L, H187M, H187F, H187S, H187V, H187W, H187K, H187Y, H187R, H187E, H187P or H187T, preferably H187P, H187V or H187T, more preferably H187V; and / or (ii) P247R, P247C, P247A or P247S, preferably P247R or P247S and / or (iii) K248R, K248A, K248S, K248V, K248I or K248C; preferably K248R; or a substitution at the corresponding position of a variant thereof, A polypeptide of claim 1 or 11.
16. and at least one of the following substitutions or groups of substitutions: (1) K14R, K13D, K13A, K30R, K33A, T83A, I100L, R115H, R143L, R147E, A205K / H207V / K208R / D210E, H207V / K208R / D210E, R214D / K215A / E216V / N217Q; M243Q, E267D, T314N and / or G317E; (2)K14R / / R214D / K215A / E216V / N217Q; (3) K33A / R115H / R214D / K215A / E216V / N217Q / M243H; (4) K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H; (5)K13D / K33A / T83N / H207V / K208R / D210E / M243Q; (6)K13A / K33A / R214D / K215A / E216V / N217Q; (7)K33A / T83N / R214D / K215NE216V / N217Q / / G317E; (8)K14R / T83A / M243Q; (9)K14R / T83A / I100L / M243Q; (10)K14R / T83A / R143L / M243Q; (11)K14R / T83A / R147E / M243Q; (12)K14R / T83A / M243Q / E267D; (13)K14R / T83A / M243Q / T314N; (14)K14R / K30R / I110L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H; (15)K14R / K30R / R143L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H; (16)K14R / K30R / R147E / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H; (17)K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / E267D; (18)K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / M243H / T314N; (19)K14R / K30R / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / G317E; (20)K14R / K33A / R115H / R214D / K215A / E216V / N217Q / M243H; (21)K14R / K30R / R147E / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / T314N; (22)K14R / K30R / R143U / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / E267D; (23)K14R / K30R / R143L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / T314N; (24)K14R / K30R / R143L / A205K / H207V / K208R / D210E / R214D / K215A / E216V / N217Q / / M243H / G317E; (25)K14R / K33A / R115H / R143L / R214D / K215A / E216V / N217Q / M243H; (26)K14R / K33A / R115H / R147E / / R214D / K215A / E216V / N217Q / M243H; (27)K14R / K33A / R115H / R214D / K215A / E216V / N217Q / / M243H / E267D; (28)K14R / K33A / R115H / R214D / K215A / E216V / N217Q / / M243H / T314N; (29)K14R / K33A / R115H / R214D / K215A / E216V / N217Q / / M243H / G317E; (30) K14R / T83A / M243Q / G317E; or (31)K13A / K33A / T83N / R214D / K215A / E216V / N217Q; or a variant thereof, comprising at least one of said substitutions or groups of substitutions at the corresponding position(s), A polypeptide of claim 1 or 11.
17. A nucleic acid comprising a nucleic acid sequence encoding the polypeptide of claim 1 or 11, optionally wherein the nucleic acid sequence is operably linked to at least one transcription control unit.
18. A vector comprising the nucleic acid molecule of claim 17.
19. A cell comprising the nucleic acid of claim 17.
20. 1. An in vitro method for integrating an exogenous nucleic acid into the genome of a cell, comprising: a) providing an isolated cell; b) providing the cell with a polypeptide of claim 1 or 11; and c) providing a cell with a nucleic acid containing an exogenous nucleic acid or a vector containing the same; A method comprising the steps of:
21. (a) a polypeptide having transposition activity is provided to a cell by introducing into the cell a nucleic acid comprising a nucleic acid sequence encoding a polypeptide comprising or consisting of a transposase of the Tc1 / mariner superfamily, wherein the amino acid positions corresponding to amino acid positions 248, 247, and / or 187 of sleeping beauty (SB) transposase of SEQ ID NO: 1 are substituted with different amino acids, wherein the transposase has acquired specificity for integration into the genome; and / or (b) a nucleic acid or vector containing the exogenous nucleic acid is provided separately to the cell; 21. The in vitro method of claim 20.
22. An in vivo method for integrating an exogenous nucleic acid into the genome of a subject's cells, comprising the step of administering to a subject a nucleic acid comprising the polypeptide of claim 1 or 11, the nucleic acid of claim 17 and the exogenous nucleic acid, or a vector comprising the exogenous nucleic acid.
23. 20. A polypeptide of claim 1 or 11, a nucleic acid of claim 17, a vector of claim 18 or a cell of claim 19 for use in medicine.
24. 20. The polypeptide of claim 1 or 11, the nucleic acid of claim 17, the vector of claim 18 or the cell of claim 19 for use in gene therapy, including but not limited to autologous or allogeneic T cell therapy, gene therapy targeting any cell type in the blood, hematopoietic stem cell therapy, liver gene therapy, central nervous system gene therapy, eye gene therapy, muscle gene therapy, and skin gene therapy.
25. 20. A pharmaceutical composition comprising the polypeptide of claim 1 or 11, the nucleic acid of claim 17, the vector of claim 18 or the cell of claim 19 and a pharmaceutically acceptable carrier or vehicle.