Enhanced hAT family transposon-mediated gene transfer and related compositions, systems, and methods

Mutant TcBuster transposases with specific amino acid substitutions and fusion designs, combined with optimized delivery methods, address efficiency limitations in DNA transposon systems, achieving high gene transfer and delivery to human hematopoietic and immune cells.

JP7733865B2Active Publication Date: 2025-09-04R&D SYST INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024001350
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-06-21
Filing Date
2024-01-09
Publication Date
2025-09-04
Estimated Expiration
2039-06-21

AI Technical Summary

Technical Problem

Existing DNA transposon systems, such as TcBuster, face limitations in transposition efficiency and delivery methods, particularly in human hematopoietic and immune system cells, necessitating improved transposase designs and delivery strategies.

Method used

Development of mutant TcBuster transposases with specific amino acid substitutions and fusion transposases incorporating DNA sequence-specific binding domains, along with optimized transposon vector designs and delivery methods using chemically modified mRNA and minicircle plasmids, enhance transposition efficiency and gene transfer.

Benefits of technology

The enhanced transposases and delivery systems achieve unexpectedly high levels of gene transfer and improved delivery to target cell types, including human hematopoietic and immune cells, with increased transposition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007733865000047
    Figure 0007733865000047
  • Figure 0007733865000048
    Figure 0007733865000048
  • Figure 0007733865000049
    Figure 0007733865000049
Patent Text Reader

Abstract

To provide various devices, systems and methods relating to synergistic approaches to enhance gene transfer into human hematopoietic and immune system cells using hAT family transposon components that are widespread in plants and animals.SOLUTION: The present invention provides a mutant TcBuster transposase comprising an amino acid sequence at least 90% identical to the full-length TcBuster transposase sequence and having an amino acid substitution, situated at a specific site, and also having at least one additional amino acid substitution.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Detailed Description of the Invention

[0001] [Technical field] cross reference This application also claims the benefit of U.S. Provisional Patent Application No. 62 / 688,278, filed June 21, 2018, which is incorporated herein by reference in its entirety. [Background technology]

[0002] Transposable genetic elements, also called transposons, are segments of DNA that can move from one location in the genome to another within a single cell. Transposons can be divided into two major groups according to their mechanism of transposition: (1) transposition can occur via reverse transcription of an RNA intermediate, an element called a retrotransposon, and (2) via direct transposition of DNA flanked by terminal inverted repeats (TIRs) for DNA transposons. Active transposons encode one or more proteins required for transposition. Naturally occurring active DNA transposons harbor the gene for the transposase enzyme.

[0003] The hAT family of DNA transposons is widely distributed in plants and animals. Numerous active hAT transposon systems have been identified and found to be functional, including, but not limited to, Hermes, Ac, hobo, and Tol2 transposons. The hAT family consists of two families, classified into the AC and Buster subfamilies based on the primary sequence of the transposase. Members of the hAT family belong to class II transposable elements. Class II mobile elements use a cut-and-paste mechanism of transposition. hAT elements share a similar transposase, short terminal inverted repeats, and eight-base-pair overlap of genomic targets. [Summary of the Invention]

[0004] Described herein, in one aspect, is a mutant TcBuster transposase comprising an amino acid sequence that is at least 70% identical to full-length SEQ ID NO:1 and has one or more amino acid substitutions in Table 1.1. In some embodiments, the mutant TcBuster transposase comprises an amino acid substitution that increases the net charge at neutral pH relative to SEQ ID NO:1. In some embodiments, the amino acid substitution that increases the net charge at neutral pH comprises a substitution of lysine or arginine. In some embodiments, the amino acid substitution that increases the net charge at neutral pH comprises a substitution of aspartic acid or glutamic acid with the neutral amino acids lysine or arginine. In some embodiments, the mutant TcBuster transposase comprises one or more amino acid substitutions in Table 4.1. In some embodiments, the mutant TcBuster transposase further comprises one or more amino acid substitutions in Table 4. In some embodiments, the mutant TcBuster transposase comprises an amino acid substitution in the DNA binding and oligomerization domain, the insertion domain, the Zn-BED domain, or a combination thereof. In some embodiments, the mutant TcBuster transposase comprises an amino acid substitution within or near the catalytic domain that increases the net charge at neutral pH compared to SEQ ID NO: 1. In some embodiments, the mutant TcBuster transposase comprises an amino acid substitution that increases the net charge at neutral pH compared to SEQ ID NO: 1; wherein, when numbered according to SEQ ID NO: 1, the one or more amino acids are located near D223, D289, or E589. In some embodiments, the vicinity is a distance of about 80, 75, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids. In some embodiments, the vicinity is a distance of about 70-80 amino acids. In some embodiments, the amino acid sequence of the mutant TcBuster transposase is at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the full-length SEQ ID NO: 1. In some embodiments, the mutant TcBuster transposase further comprises one or more amino acid substitutions from Table 2. In some embodiments, the mutant TcBuster transposase further comprises one or more amino acid substitutions in Table 3.In some embodiments, when numbered according to SEQ ID NO: 1, the mutant TcBuster transposase further comprises the amino acid substitutions V377T, E469K, and D189A. In some embodiments, when numbered according to SEQ ID NO: 1, the mutant TcBuster transposase further comprises the amino acid substitutions K573E and E578L. In some embodiments, when numbered according to SEQ ID NO: 1, the mutant TcBuster transposase further comprises the amino acid substitution I452K. In some embodiments, when numbered according to SEQ ID NO: 1, the mutant TcBuster transposase further comprises the amino acid substitution A358K. In some embodiments, when numbered according to SEQ ID NO: 1, the mutant TcBuster transposase further comprises the amino acid substitution V297K. In some embodiments, when numbered according to SEQ ID NO: 1, the mutant TcBuster transposase further comprises the amino acid substitution N85S. In some embodiments, when numbered according to SEQ ID NO: 1, the mutant TcBuster transposase further comprises the amino acid substitutions I452F, V377T, E469K, and D189A. In some embodiments, when numbered according to SEQ ID NO: 1, the mutant TcBuster transposase further comprises the amino acid substitutions A358K, V377T, E469K, and D189A. In some embodiments, when numbered according to SEQ ID NO: 1, the mutant TcBuster transposase further comprises the amino acid substitutions V377T, E469K, D189A, K573E, and E578L. In some embodiments, the mutant TcBuster transposase further comprises one or more amino acid substitutions in Table 1. In some embodiments, the mutant TcBuster transposase has increased transposition efficiency compared to a wild-type TcBuster transposase having the amino acid sequence SEQ ID NO: 1. In some embodiments, transposition efficiency is measured by an assay comprising introducing a mutant TcBuster transposase or a wild-type TcBuster transposase and a TcBuster transposon containing a reporter cargo cassette into a cell population and detecting transposition of the reporter cargo cassette in the genome of the cell population.

[0005] Described herein, in one aspect, is a fusion transposase comprising a TcBuster transposase sequence and one or more additional nuclear localization signal sequences, wherein the TcBuster transposase sequence has at least 70% identity to full-length SEQ ID NO:1. In some embodiments, the TcBuster transposase sequence has at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to full-length SEQ ID NO:1. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions that increase the net charge at neutral pH relative to SEQ ID NO:1. In some embodiments, the one or more amino acid substitutions comprise substitutions with lysine or arginine. In some embodiments, the one or more amino acid substitutions comprise substitutions of aspartic acid or glutamic acid with neutral amino acids, lysine or arginine. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions from Table 4, Table 4.1, or both. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions in the DNA-binding and oligomerization domain; the insertion domain; the Zn-BED domain; or a combination thereof. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions from Table 1, Table 1.1, or both. In some embodiments, the TcBuster transposase sequence has a higher transposition efficiency compared to a wild-type TcBuster transposase having the amino acid sequence SEQ ID NO:1. In some embodiments, the transposition efficiency of the TcBuster transposase sequence is measured by an assay comprising introducing a fusion TcBuster transposase or a wild-type TcBuster transposase and a TcBuster transposon containing a reporter cargo cassette into a cell population and detecting transposition of the reporter cargo cassette in the genome of the cell population. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions within or near the catalytic domain compared to SEQ ID NO:1 that increase the net charge at neutral pH.In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions that increase the net charge at neutral pH compared to SEQ ID NO:1, and the one or more amino acid substitutions are located near D223, D289, or E589 when numbered according to SEQ ID NO:1. In some embodiments, the vicinity is at a distance of about 80, 75, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids. In some embodiments, the vicinity is at a distance of about 70-80 amino acids. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions in Table 2. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions in Table 3. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO:1, comprises the amino acid substitutions V377T, E469K, and D189A. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO:1, comprises the amino acid substitutions K573E and E578L. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO: 1, comprises the amino acid substitution I452K. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO: 1, comprises the amino acid substitution A358K. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO: 1, comprises the amino acid substitution V297K. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO: 1, comprises the amino acid substitution N85S. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO: 1, comprises the amino acid substitutions I452F, V377T, E469K, and D189A. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO: 1, comprises the amino acid substitutions A358K, V377T, E469K, and D189A. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO:1, comprises the amino acid substitutions V377T, E469K, D189A, K573E, and E578L.In some embodiments, the TcBuster transposase sequence has 100% identity to the full length SEQ ID NO:1.

[0006] Described herein, in one aspect, is a fusion transposase comprising a TcBuster transposase sequence and a DNA sequence-specific binding domain, wherein the TcBuster transposase sequence has the amino acid sequence of a mutant TcBuster of any one of claims 1-26. In some embodiments, the DNA sequence-specific binding domain comprises a TALE domain, a zinc finger domain, an AAV Rep DNA binding domain, or any combination thereof. In some embodiments, the DNA sequence-specific binding domain comprises a TALE domain. In some embodiments, the TcBuster transposase sequence and the DNA sequence-specific binding domain are separated by a linker. In some embodiments, the linker comprises at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, or at least 50 amino acids. In some embodiments, the linker comprises SEQ ID NO:9.

[0007] Described herein, in one aspect, is a polynucleotide comprising a nucleic acid sequence at least about 80%, 85%, 90%, 95%, or 98% identical to or complementary to the full-length SEQ ID NO: 204 or 207. Described herein, in one aspect, is a polynucleotide encoding a mutant TcBuster transposase described herein. Described herein, in one aspect, is a polynucleotide encoding a fusion transposase described herein. In some embodiments, the polynucleotide comprises DNA encoding a mutant TcBuster transposase or fusion transposase. In some embodiments, the polynucleotide comprises messenger RNA (mRNA) encoding the mutant TcBuster transposase or fusion transposase. In some embodiments, the mRNA is chemically modified. In some embodiments, the polynucleotide comprises a nucleic acid sequence encoding a transposon recognizable by the mutant TcBuster transposase or fusion transposase. In some embodiments, the polynucleotide is present in a DNA vector. In some embodiments, the DNA vector comprises a minicircle plasmid. In some embodiments, the polynucleotide is codon-optimized for expression in human cells. In some embodiments, the polynucleotide comprises a nucleic acid sequence that is at least about 70%, 75%, 80%, 85%, 90%, 95%, or 98% identical to or complementary to the full-length SEQ ID NO:204 or 207. In some embodiments, the polynucleotide comprises a nucleic acid sequence that is at least about 80%, 85%, 90%, 95%, or 98% identical to or complementary to the full-length SEQ ID NO:204 or 207. In some embodiments, the polynucleotide comprises a nucleic acid sequence that is at least about 95% identical to or complementary to the full-length SEQ ID NO:204 or 207. In some embodiments, the polynucleotide comprises a nucleic acid sequence that is 100% identical to or complementary to the full-length SEQ ID NO:204 or 207.

[0008] One aspect of the disclosure provides a mutant TcBuster transposase comprising an amino acid sequence that is at least 70% identical to full-length SEQ ID NO:1 and has one or more amino acid substitutions that increase the net charge at neutral pH relative to SEQ ID NO:1. In some embodiments, the mutant TcBuster transposase has a higher transposition efficiency relative to a wild-type TcBuster transposase having the amino acid sequence SEQ ID NO:1. Another aspect of the disclosure provides a mutant TcBuster transposase comprising an amino acid sequence that is at least 70% identical to full-length SEQ ID NO:1 and has another amino acid substitution in the DNA-binding and oligomerization domain; the insertion domain; the Zn-BED domain; or a combination thereof. In some embodiments, the mutant TcBuster transposase has a higher transposition efficiency relative to a wild-type TcBuster transposase having the amino acid sequence SEQ ID NO:1. Yet another aspect of the disclosure provides a mutant TcBuster transposase comprising an amino acid sequence that is at least 70% identical to full-length SEQ ID NO:1 and has one or more amino acid substitutions from Table 1. In some embodiments, the mutant TcBuster transposase comprises one or more amino acid substitutions within or near the catalytic domain that increase the net charge at neutral pH relative to SEQ ID NO:1. In some embodiments, the mutant TcBuster transposase comprises one or more amino acid substitutions that increase the net charge at neutral pH relative to SEQ ID NO:1, wherein the one or more amino acids are located near D223, D289, or E589, when numbered according to SEQ ID NO:1. In some embodiments, the vicinity is a distance of about 80, 75, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids. In some embodiments, the vicinity is a distance of about 70-80 amino acids. In some embodiments, the amino acid sequence of the mutant TcBuster transposase is at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the full-length SEQ ID NO:1. In some embodiments, the one or more amino acid substitutions comprise a substitution to lysine or arginine.In some embodiments, the one or more amino acid substitutions comprise a substitution of aspartic acid or glutamic acid with a neutral amino acid, lysine, or arginine. In some embodiments, the mutant TcBuster transposase comprises one or more amino acid substitutions in Table 4. In some embodiments, the mutant TcBuster transposase comprises one or more amino acid substitutions in Table 2. In some embodiments, the mutant TcBuster transposase comprises one or more amino acid substitutions in Table 3. In some embodiments, the mutant TcBuster transposase, when numbered according to SEQ ID NO: 1, comprises the amino acid substitutions V377T, E469K, and D189A. In some embodiments, the mutant TcBuster transposase, when numbered according to SEQ ID NO: 1, comprises the amino acid substitutions K573E and E578L. In some embodiments, the mutant TcBuster transposase, when numbered according to SEQ ID NO: 1, comprises the amino acid substitution I452K. In some embodiments, the mutant TcBuster transposase, when numbered according to SEQ ID NO: 1, comprises the amino acid substitution A358K. In some embodiments, the mutant TcBuster transposase, when numbered according to SEQ ID NO: 1, comprises the amino acid substitution V297K. In some embodiments, the mutant TcBuster transposase, when numbered according to SEQ ID NO: 1, comprises the amino acid substitution N85S. In some embodiments, the mutant TcBuster transposase, when numbered according to SEQ ID NO: 1, comprises the amino acid substitutions I452F, V377T, E469K, and D189A. In some embodiments, the mutant TcBuster transposase, when numbered according to SEQ ID NO: 1, comprises the amino acid substitutions A358K, V377T, E469K, and D189A. In some embodiments, the mutant TcBuster transposase, when numbered according to SEQ ID NO: 1, comprises the amino acid substitutions V377T, E469K, D189A, K573E, and E578L.In some embodiments, transposition efficiency is measured by an assay comprising introducing a mutant TcBuster transposase and a TcBuster transposon containing a reporter cargo cassette into a cell population and detecting transposition of the reporter cargo cassette in the genome of the cell population.

[0009] Yet another aspect of the present disclosure provides a fusion transposase comprising a TcBuster transposase sequence and a DNA sequence-specific binding domain. In some embodiments, the TcBuster transposase sequence has at least 70% identity to the full-length SEQ ID NO:1. In some embodiments, the DNA sequence-specific binding domain comprises a TALE domain, a zinc finger domain, an AAV Rep DNA binding domain, or any combination thereof. In some embodiments, the DNA sequence-specific binding domain comprises a TALE domain. In some embodiments, the TcBuster transposase sequence has at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to the full-length SEQ ID NO:1. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions that increase the net charge at neutral pH compared to SEQ ID NO:1. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions in the DNA binding and oligomerization domain; the insertion domain; the Zn-BED domain; or a combination thereof. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions from Table 1. In some embodiments, the TcBuster transposase sequence has increased transposition efficiency compared to a wild-type TcBuster transposase having the amino acid sequence SEQ ID NO: 1. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions in or near the catalytic domain that increase the net charge at neutral pH compared to SEQ ID NO: 1. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions that increase the net charge at neutral pH compared to SEQ ID NO: 1, wherein the one or more amino acid substitutions are located near D223, D289, or E589, when numbered according to SEQ ID NO: 1. In some embodiments, the vicinity is a distance of about 80, 75, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids. In some embodiments, the vicinity is a distance of about 70-80 amino acids.In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions from Table 2. In some embodiments, the TcBuster transposase sequence comprises one or more amino acid substitutions from Table 3. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO:1, comprises the amino acid substitutions V377T, E469K, and D189A. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO:1, comprises the amino acid substitutions K573E and E578L. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO:1, comprises the amino acid substitution I452K. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO:1, comprises the amino acid substitution A358K. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO:1, comprises the amino acid substitution V297K. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO:1, comprises the amino acid substitution N85S. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO: 1, comprises the amino acid substitutions I452F, V377T, E469K, and D189A. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO: 1, comprises the amino acid substitutions A358K, V377T, E469K, and D189A. In some embodiments, the TcBuster transposase sequence, when numbered according to SEQ ID NO: 1, comprises the amino acid substitutions V377T, E469K, D189A, K573E, and E578L. In some embodiments, the TcBuster transposase sequence has 100% identity to the full-length SEQ ID NO: 1. In some embodiments of the fusion transposase, the TcBuster transposase sequence and the DNA sequence-specific binding domain are separated by a linker.In some embodiments, the linker comprises at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, or at least 50 amino acids. In some embodiments, the linker comprises SEQ ID NO:9.

[0010] Yet another aspect of the present disclosure provides a polynucleotide encoding a mutant TcBuster transposase as described herein. Yet another aspect of the present disclosure provides a polynucleotide encoding a fusion transposase as described herein. In some embodiments, the polynucleotide comprises DNA encoding the mutant TcBuster transposase or fusion transposase. In some embodiments, the polynucleotide comprises messenger RNA (mRNA) encoding the mutant TcBuster transposase or fusion transposase. In some embodiments, the mRNA is chemically modified. In some embodiments, the polynucleotide comprises a nucleic acid sequence encoding a transposon recognizable by the mutant TcBuster transposase or fusion transposase. In some embodiments, the polynucleotide is present in a DNA vector. In some embodiments, the DNA vector comprises a minicircle plasmid.

[0011] Yet another aspect of the present disclosure provides a cell that produces a mutant TcBuster transposase or fusion transposase as described herein. Yet another aspect of the present disclosure provides a cell containing a polynucleotide as described herein. Yet another aspect of the present disclosure provides a method comprising introducing into a cell a mutant TcBuster transposase as described herein and a transposon recognizable by the mutant TcBuster transposase. Yet another aspect of the present disclosure provides a method comprising introducing into a cell a fusion transposase as described herein and a transposon recognizable by the fusion transposase. In some embodiments of the method, the introducing comprises contacting the cell with a polynucleotide encoding the mutant TcBuster transposase or fusion transposase. In some embodiments, the polynucleotide comprises DNA encoding the mutant TcBuster transposase or fusion transposase. In some embodiments, the polynucleotide comprises messenger RNA (mRNA) encoding the mutant TcBuster transposase or fusion transposase. In some embodiments, the mRNA is chemically modified. In some embodiments of the method, the introducing comprises contacting the cell with a DNA vector containing a transposon. In some embodiments, the DNA vector comprises a minicircle plasmid. In some embodiments, the introducing comprises contacting the cell with a plasmid vector containing both a transposon and a polynucleotide encoding the mutant TcBuster transposase or fusion transposase. In some embodiments, the introducing comprises contacting the cell with the mutant TcBuster transposase or fusion transposase as a purified protein. In some embodiments of the method, the transposon comprises a cargo cassette positioned between two inverted repeats.In some embodiments, the left inverted repeat sequence of the two inverted repeats comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO: 3. In some embodiments, the left inverted repeat sequence of the two inverted repeats comprises SEQ ID NO: 3. In some embodiments, the right inverted repeat sequence of the two inverted repeats comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO: 4. In some embodiments, the left inverted repeat sequence of the two inverted repeats comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO: 5. In some embodiments, the left inverted repeat of the two inverted repeat sequences comprises SEQ ID NO: 5. In some embodiments, the right inverted repeat of the two inverted repeat sequences comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO: 6. In some embodiments, the right inverted repeat of the two inverted repeat sequences comprises SEQ ID NO: 6. In some embodiments, the left inverted repeat of the two inverted repeat sequences comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO: 205. In some embodiments, the left inverted repeat of the two inverted repeat sequences comprises SEQ ID NO: 205. In some embodiments, the right inverted repeat of the two inverted repeat sequences comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO:206.In some embodiments, the right inverted repeat of the two inverted repeats comprises SEQ ID NO: 206. In some embodiments, the cargo cassette comprises a promoter selected from the group consisting of CMV, EFS, MND, EF1α, CAGCs, PGK, UBC, U6, H1, and cumate. In some embodiments, the vector comprises a CMV promoter. In some embodiments, the cargo cassette is present in a forward orientation. In some embodiments, the cargo cassette is present in a reverse orientation. In some embodiments, the cargo cassette comprises a transgene. In some embodiments, the transgene encodes a protein selected from the group consisting of a cellular receptor, an immune checkpoint protein, a cytokine, and any combination thereof. In some embodiments, the transgene encodes a cellular receptor selected from the group consisting of a T cell receptor (TCR), a B cell receptor (BCR), a chimeric antigen receptor (CAR), or any combination thereof. In some embodiments, the introducing comprises transfecting the cells with the aid of electroporation, microinjection, calcium phosphate precipitation, cationic polymers, dendrimers, liposomes, biolistics, FuGene, direct sonication, cell squeezing, optical transfection, protoplast fusion, imparefection, magnetofection, nucleofection, or any combination thereof. In some embodiments, the introducing comprises electroporating the cells. In some embodiments of the method, the cells are primary cells isolated from a subject. In some embodiments, the subject is human. In some embodiments, the subject is a diseased patient. In some embodiments, the subject has been diagnosed with cancer or a tumor. In some embodiments, the cells are isolated from the subject's blood. In some embodiments, the cells comprise primary immune cells. In some embodiments, the cells comprise primary leukocytes. In some embodiments, the cells comprise primary T cells. In some embodiments, the primary T cells comprise gamma delta T cells, helper T cells, memory T cells, natural killer T cells, effector T cells, or any combination thereof. In some embodiments, the primary immune cells comprise CD3+ cells.In some embodiments, the cells comprise stem cells. In some embodiments, the stem cells are selected from the group consisting of embryonic stem cells, hematopoietic stem cells, epidermal stem cells, epithelial stem cells, bronchoalveolar stem cells, mammary stem cells, mesenchymal stem cells, intestinal stem cells, endothelial stem cells, neural stem cells, olfactory adult stem cells, neural crest stem cells, testicular cells, and any combination thereof. In some embodiments, the stem cells comprise induced pluripotent stem cells.

[0012] Another aspect of the present disclosure provides methods of treatment comprising: (a) introducing a transposon and a mutant TcBuster transposase or fusion transposase as described herein that recognizes the transposon into a cell, thereby generating an engineered cell; and (b) administering the engineered cell to a patient in need of treatment. In some embodiments, the engineered cell comprises a transgene introduced by the transposon. In some embodiments, the patient has been diagnosed with cancer or a tumor. In some embodiments, administering comprises infusing the engineered cell into the patient's bloodstream.

[0013] Yet another aspect of the present disclosure provides a system for genome editing comprising a mutant TcBuster transposase or fusion transposase as described herein and a transposon recognizable by the mutant TcBuster transposase or fusion transposase. Yet another aspect of the present disclosure provides a system for genome editing comprising a polynucleotide encoding a mutant TcBuster transposase or fusion transposase as described herein and a transposon recognizable by the mutant TcBuster transposase or fusion transposase. In some embodiments of the system, the polynucleotide comprises DNA encoding the mutant TcBuster transposase or fusion transposase. In some embodiments, the polynucleotide comprises messenger RNA (mRNA) encoding the mutant TcBuster transposase or fusion transposase. In some embodiments, the mRNA is chemically modified. In some embodiments, the transposon is present in a DNA vector. In some embodiments, the DNA vector comprises a minicircle plasmid. In some embodiments, the polynucleotide and the transposon are present in the same plasmid. In some embodiments, the transposon comprises a cargo cassette disposed between two inverted repeat sequences. In some embodiments, the left inverted repeat sequence of the two inverted repeat sequences comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO:3. In some embodiments, the left inverted repeat sequence of the two inverted repeat sequences comprises SEQ ID NO:3. In some embodiments, the right inverted repeat sequence of the two inverted repeat sequences comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO:4. In some embodiments, the right inverted repeat sequence of the two inverted repeat sequences comprises SEQ ID NO:4.In some embodiments, the left inverted repeat of the two inverted repeat sequences comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO: 5. In some embodiments, the left inverted repeat of the two inverted repeat sequences comprises SEQ ID NO: 5. In some embodiments, the right inverted repeat of the two inverted repeat sequences comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO: 6. In some embodiments, the left inverted repeat of the two inverted repeat sequences comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO: 205. In some embodiments, the left of the two inverted repeats comprises SEQ ID NO: 205. In some embodiments, the right of the two inverted repeats comprises a sequence having at least 50%, at least 60%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO: 206. In some embodiments, the right of the two inverted repeats comprises SEQ ID NO: 206. In some embodiments, the cargo cassette comprises a promoter selected from the group consisting of CMV, EFS, MND, EF1α, CAGCs, PGK, UBC, U6, H1, and cumate. In some embodiments, the vector comprises a CMV promoter. In some embodiments, the cargo cassette comprises a transgene. In some embodiments, the transgene encodes a protein selected from the group consisting of a cellular receptor, an immune checkpoint protein, a cytokine, and any combination thereof. In some embodiments, the transgene encodes a cellular receptor selected from the group consisting of a T cell receptor (TCR), a B cell receptor (BCR), a chimeric antigen receptor (CAR), or any combination thereof.In some embodiments, the cargo cassette is present in a forward orientation. In some embodiments, the cargo cassette is present in a reverse orientation.

[0014] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that a term incorporated by reference conflicts with a term defined herein, the present specification controls.

[0015] The novel features of the present disclosure are set forth with particularity in the appended claims. A better understanding of these features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings. [Brief description of the drawing]

[0016]

[0014] FIG. 1 shows the transposition efficiency of several exemplary TcBuster transposon vector constructs as measured by the percentage of mCherry-positive cells in cells transfected with wild-type (WT) TcBuster transposase and exemplary TcBuster transposons. FIG. 2 shows a comparison of the nucleotide sequences of exemplary TcBuster IR / DR sequence 1 (SEQ ID NOS: 3-4, respectively, in order of appearance) and sequence 2 (SEQ ID NOS: 5-6, respectively, in order of appearance). Figure 3A shows representative brightfield and fluorescent images of HEK-293T cells two weeks after transfection with an exemplary TcBuster transposon, Tn-8 (containing a puro-mCherry cassette; shown in Figure 1), and wild-type or V596A mutant transposases (containing a V596A substitution). Transfected cells were plated in 6-well plates with 1 μg / mL puromycin two days after transfection and fixed and stained with crystal violet two weeks after transfection for colony quantification. B shows representative photographs of transfected cell colonies in 6-well plates two weeks after transfection. C is a graph showing colony quantification for each transfection condition two weeks after transfection.

[0023] FIG. 4 shows a comparison of the amino acid sequence of the TcBuster transposase compared to several transposases in the AC subfamily, with only regions of amino acid conservation shown (SEQ ID NOS: 89-194, respectively, in order of appearance). Figure 5 shows a comparison of the amino acid sequence of the TcBuster transposase relative to a number of other transposase members in the Buster subfamily (SEQ ID NOS: 195-203, respectively, in order of appearance). Specific exemplary amino acid substitutions are indicated above the protein sequence, with the percentage shown above the comparison indicating the percentage of other Buster subfamily members that contain the amino acid that is believed to be substituted in the TcBuster sequence, and the percentage shown below indicating the percentage of other Buster subfamily members that contain the standard TcBuster amino acid at that position.

[0023] FIG. 6 shows the vector map of the exemplary expression vector pcDNA-DEST40 used to test TcBuster transposase variants. [Figure 7] Graph quantifying the transposition efficiency of exemplary TcBuster transposase variants, as measured by the percentage of mCherry-positive cells in HEK-293T cells transfected with the TcBuster transposon Tn-8 (shown in Figure 1) along with the exemplary transposase variants. [Figure 8] Shows one exemplary fusion transposase containing a DNA sequence-specific binding domain and a TcBuster transposase sequence linked by an optional linker. [Figure 9] Graph quantifying the transposition efficiency of exemplary TcBuster transposases containing different tags, as measured by the percentage of mCherry-positive cells in HEK-293T cells transfected with TcBuster transposon Tn-8 (shown in Figure 1) together with exemplary transposases containing the tags. Figure 10A is a graph quantifying the transduction efficiency of an exemplary TcBuster transduction system in human CD3+ T cells as measured by the percentage of GFP-positive cells. B is a graph quantifying the viability of transfected T cells 2 and 7 days after transfection by flow cytometry. Data relate to pulse control.

[0031] FIG. 11 shows the amino acid sequence of wild-type TcBuster transposase with specific amino acids annotated (SEQ ID NO: 1).

[0033] FIG. 12 shows the amino acid sequence of a mutant TcBuster transposase containing the amino acid substitutions D189A / V377T / E469K (SEQ ID NO: 78).

[0033] FIG. 13 shows the amino acid sequence of a mutant TcBuster transposase containing the amino acid substitutions D189A / V377T / E469K / I452K (SEQ ID NO: 79).

[0033] FIG. 14 shows the amino acid sequence of a mutant TcBuster transposase containing the amino acid substitutions D189A / V377T / E469K / N85S (SEQ ID NO: 80).

[0033] FIG. 15 shows the amino acid sequence of a mutant TcBuster transposase containing the amino acid substitutions D189A / V377T / E469K / A358K (SEQ ID NO: 81). FIG. 16 shows the amino acid sequence of a mutant TcBuster transposase containing the amino acid substitutions D189A / V377T / E469K / K573E / E578L (SEQ ID NO: 13). [Mode for Carrying Out the Invention]

[0017] overview

[0018] DNA transposons can translocate via a non-replicative "cut-and-paste" mechanism. This requires the recognition of two terminal inverted repeats by a catalytic enzyme, a transposase, which can cleave its target and thus liberate the DNA transposon from its donor template. Upon excision, the DNA transposon may then integrate into acceptor DNA that is cleaved by the same transposase. In some of their natural configurations, DNA transposons are flanked by two inverted repeats and may contain a gene encoding the transposase that catalyzes transposition.

[0019] For the application of DNA transposon-mediated genome editing, it is desirable to engineer transposons to develop a binary system based on two different plasmids, whereby the transposase is physically separated from the transposon DNA containing the gene of interest flanked by inverted repeats. Co-delivery of the transposon and transposase plasmids into target cells allows transposition via the traditional cut-and-paste mechanism.

[0020] TcBuster is a member of the hAT family of DNA transposons. Other members of the family include Sleeping Beauty and PiggBac. Discussed herein are various devices, systems, and methods related to a synergistic approach for enhancing gene transfer into human hematopoietic and immune system cells using hAT family transposon components. This disclosure relates to improved hAT transposases, transposon vector sequences, transposase delivery methods, and transposon delivery methods. In one embodiment, this work identified specific universal sites for generating hyperactive hAT transposases. In another embodiment, a method for generating minimally sized inverted terminal repeats (ITRs) of hAT transposon vectors that conserve genomic space is described. In another embodiment, an improved method for delivering hAT family transposases as chemically modified in vitro transcribed mRNA is described. In another embodiment, a method for delivering hAT family transposon vectors as "miniature" circles of DNA, with substantially all prokaryotic sequences recombinantly removed, is described. In another embodiment, a method is described for fusing DNA sequence-specific binding domains with transcription activator-like (TAL) domains fused to the hAT transposase. These improvements, individually or in combination, can result in unexpectedly high levels of gene transfer into cell types of interest and improved delivery of transposon vectors to sequences of interest.

[0021] Mutant TcBuster transposase

[0022] One aspect of the present disclosure provides a mutant TcBuster transposase, which may comprise one or more amino acid substitutions compared to the wild-type TcBuster transposase (SEQ ID NO: 1).

[0023] A mutant TcBuster transposase can comprise an amino acid sequence that has at least 70% sequence identity to the full-length sequence of the wild-type TcBuster transposase (SEQ ID NO: 1). In some embodiments, a mutant TcBuster transposase can comprise an amino acid sequence that has at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the full-length sequence of the wild-type TcBuster transposase (SEQ ID NO: 1). In some cases, the mutant TcBuster transposase can comprise an amino acid sequence having at least 98%, at least 98.5%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or at least 99.95% sequence identity to the full-length sequence of the wild-type TcBuster transposase (SEQ ID NO: 1).

[0024] As used herein, the term "percent identity" refers to the percentage of amino acid (or nucleic acid) residues in a candidate sequence that are identical to those in a reference sequence, after aligning the sequences to achieve the maximum percentage identity and, if necessary, introducing gaps (i.e., gaps can be introduced in one or both of the candidate and reference sequences for optimal alignment, and non-homologous sequences can be ignored for comparison purposes). Sequence comparison for purposes of determining percent identity can be accomplished in a variety of ways that are within the skill of the art, for example, using publicly available computer software such as BLAST, ALIGN, or Megalign (DNASTAR) software. The percent identity of two sequences can be calculated by aligning a test sequence with a comparison sequence using BLAST, determining the number of amino acids or nucleotides in the aligned test sequence that are identical to amino acids or nucleotides at the same positions in the comparison sequence, and dividing the number of identical amino acids or nucleotides by the number of amino acids or nucleotides in the comparison sequence.

[0025] As used herein, the terms "complement," "complements," "complementary," and "complementarity" can refer to a sequence that is perfectly complementary to and hybridizable with a given sequence. In some cases, a sequence that hybridizes to a given nucleic acid is referred to as the "complement" or "reverse complement" of the given molecule because its sequence of bases across a given region can complementarily bind to its binding partner base sequence, e.g., if AT, AU, GC, and GU base pairs are formed. Generally, a first sequence that can hybridize to a second sequence is specifically or selectively hybridizable to the second sequence, such that hybridization to a second sequence or set of second sequences (e.g., thermodynamically more stable under a given set of conditions, such as stringent conditions commonly used in the art) is favored over hybridization to non-target sequences during a hybridization reaction. Typically, hybridizable sequences share sequence complementarity over all or part of their respective lengths, e.g., a degree of complementarity between 25% and 100%, including at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, and 100% sequence complementarity. Sequence identity, such as for purposes of assessing percent complementarity, can be measured by a suitable sequence comparison algorithm, including, but not limited to, the Needleman-Wunsch algorithm (see, e.g., the EMBOSS Needle aligner available at www.ebi.ac.uk / Tools / psa / emboss_needle / nucleotide.html, optionally with default settings), the BLAST algorithm (see, e.g., the BLAST sequence comparison tool available at blast.ncbi.nlm.nih.gov / Blast.cgi, optionally with default settings), or the Smith-Waterman algorithm (see, e.g., the EMBOSS Water aligner available at www.ebi.ac.uk / Tools / psa / emboss_water / nucleotide.html, optionally with default settings).Optimal sequence comparison can be assessed using appropriate parameters of a selected algorithm, including default parameters.

[0026] Complementarity can be perfect or substantial / sufficient. Perfect complementarity between two nucleic acids means that the two nucleic acids can form a duplex in which all bases in the duplex bind to complementary bases through Watson-Crick pairing. Substantial or sufficient complementarity can mean that the sequence of one strand is not completely and / or perfectly complementary to the sequence of the opposite strand, but sufficient binding occurs between the bases of the two strands to form a stable hybrid complex under a set of hybridization conditions (e.g., salt concentration and temperature). Such conditions can be predicted by predicting the Tm of the hybridized strands using the sequences and standard mathematical calculations, or by empirically determining the Tm using routine methods.

[0027] A mutant TcBuster transposase can comprise an amino acid sequence having at least one amino acid that differs from the full-length sequence of a wild-type TcBuster transposase (SEQ ID NO: 1). In some embodiments, a mutant TcBuster transposase can comprise an amino acid sequence having at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more amino acids that differ from the full-length sequence of a wild-type TcBuster transposase (SEQ ID NO: 1). In some cases, a mutant TcBuster transposase can comprise an amino acid sequence having at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, or at least 300 amino acids that differ from the full-length sequence of a wild-type TcBuster transposase (SEQ ID NO: 1). In some cases, the mutant TcBuster transposase can comprise an amino acid sequence having at most 3, at most 6, at most 12, at most 25, at most 35, at most 45, at most 55, at most 65, at most 75, at most 85, at most 95, at most 150, at most 250, or at most 350 amino acids that differ from the full-length sequence of the wild-type TcBuster transposase (SEQ ID NO: 1).

[0028] As shown in FIG. 4 , wild-type TcBuster transposases can generally be considered to comprise, from N- to C-terminus, a ZnF-BED domain (amino acids 76-98), a DNA-binding and oligomerization domain (amino acids 112-213), a first catalytic domain (amino acids 213-312), an insertion domain (amino acids 312-543), and a second catalytic domain (amino acids 583-620), as well as at least four interdomain regions between these annotated domains. Unless otherwise specified, all numerical references to amino acids as used herein are in accordance with SEQ ID NO: 1. Mutant TcBuster transposases can contain one or more amino acid substitutions in any one of these domains, or in any combination thereof. In some cases, mutant TcBuster transposases can contain one or more amino acid substitutions in the ZnF-BED domain, the DNA-binding and oligomerization domain, the first catalytic domain, the insertion domain, or a combination thereof. The mutant TcBuster transposase can contain one or more amino acid substitutions in at least one of the two catalytic domains.

[0029] Exemplary mutant TcBuster transposases can include one or more amino acid substitutions in Table 1 or Table 1.1. Sometimes, mutant TcBuster transposases can include at least one of the amino acid substitutions in Table 1 or Table 1.1. Mutant TcBuster transposases can include at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least 20, at least 30 or more of the amino acid substitutions in Table 1 or Table 1.1. [Table 1] JPEG0007733865000002.jpg235153 JPEG0007733865000003.jpg220153 [Table 2] JPEG0007733865000005.jpg234153 JPEG0007733865000006.jpg234153 JPEG0007733865000007.jpg51153

[0030] Exemplary mutant TcBuster transposases comprise one or more amino acid substitutions, or combinations of substitutions, in Table 2. In some cases, mutant TcBuster transposases can comprise at least one or combination of amino acid substitutions in Table 2. Mutant TcBuster transposases can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30 or more amino acid substitutions or combinations of substitutions in Table 2. [Table 3] JPEG0007733865000009.jpg171153

[0031] Exemplary mutant TcBuster transposases comprise one or more amino acid substitutions, or combinations of substitutions, in Table 3. In some cases, mutant TcBuster transposases can comprise at least one or combination of amino acid substitutions in Table 3. Mutant TcBuster transposases can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30 or more amino acid substitutions or combinations of substitutions in Table 3. [Table 4]

[0032] Hyperactive mutant TcBuster transposase

[0033] Another aspect of the present disclosure is to provide hyperactive mutant TcBuster transposases. As used herein, a "hyperactive" mutant TcBuster transposase can refer to any mutant TcBuster transposase that has increased transposition efficiency compared to the wild-type TcBuster transposase having the amino acid sequence SEQ ID NO:1.

[0034] In some embodiments, a hyperactive mutant TcBuster transposase may have a higher transposition efficiency under certain circumstances than a wild-type TcBuster transposase having the amino acid sequence SEQ ID NO: 1. For example, a hyperactive mutant TcBuster transposase may have a better transposition efficiency than a wild-type TcBuster transposase when used to catalyze the transposition of transposons having certain types of inverted repeats. With some other transposons having other types of inverted repeats, a hyperactive mutant TcBuster transposase may not have a higher transposition efficiency than a wild-type TcBuster transposase. In some other non-limiting examples, a hyperactive mutant TcBuster transposase may have a higher transposition efficiency than a wild-type TcBuster transposase having the amino acid sequence SEQ ID NO: 1 under certain transfection conditions. Without limitation, when compared to a wild-type TcBuster transposase, a hyperactive mutant TcBuster transposase may have better transposition efficiency if the temperature is higher than normal cell culture temperatures; a hyperactive mutant TcBuster transposase may have better transposition efficiency in relatively acidic or basic aqueous media; or a hyperactive mutant TcBuster transposase may have better transposition efficiency if certain types of transfection techniques (e.g., electroporation) are performed.

[0035] Transposition efficiency can be measured by the rate of successful transposition events occurring in a population of host cells, normalized by the amount of transposon and transposase introduced into the population of host cells. Often, when comparing the transposition efficiencies of two or more transposases, the same transposon construct is paired with each of the two or more transposases for transfection of host cells under the same or similar transfection conditions. The amount of transposition events in a host cell can be determined by various approaches. For example, a transposon construct may be designed to contain a reporter gene located between the inverted repeats, and transfected cells that test positive for the reporter gene can be counted as cells in which successful transposition events occur, which can provide an estimate of the amount of transposition events. Another non-limiting example includes sequencing the host cell genome to determine the insertion of the transposon cassette cargo. In some embodiments, when comparing the transposition efficiencies of two or more different transposons, the same transposase can be paired with each of the different transposons for transfection of host cells under the same or similar transfection conditions. A similar approach can be used to measure transposition efficiency. Other methods known to those skilled in the art may also be performed to compare transposition efficiencies.

[0036] Also provided herein are methods for obtaining hyperactive mutant TcBuster transposases.

[0037] One exemplary method can involve systematically mutating amino acids in a TcBuster transposase to increase the net charge of the amino acid sequence. Sometimes, the method can involve performing a systematic alanine scan to mutate aspartic acid (D) or glutamic acid (E), which are negatively charged at neutral pH, to alanine residues. The method can also involve systematically mutating lysine (K) or arginine (R), which are positively charged at neutral pH.

[0038] Without wishing to be bound by any particular theory, increasing the net charge of an amino acid sequence at neutral pH may increase the transposition efficiency of TcBuster transposase. In particular, increasing the net charge near the catalytic domain of the transposase is expected to increase transposition efficiency. It is possible that positively charged amino acids form contact points with the DNA target, allowing the catalytic domain to act on the DNA target. It is also possible that loss of these positively charged amino acids may reduce either the excision or integration activity of the transposase.

[0039] Figure 11 shows the amino acid sequence of wild-type TcBuster transposase, highlighting amino acids that may be contact points with DNA. In Figure 11, large bold letters indicate catalytic triad amino acids; boxed letters indicate amino acids that increase translocation when substituted with a positively charged amino acid; italicized and lowercase letters indicate positively charged amino acids that decrease translocation when substituted with a different amino acid; bold italicized and underlined letters indicate amino acids that increase translocation when substituted with a positively charged amino acid and decrease translocation when substituted with a negatively charged amino acid; and underlined letters indicate amino acids that are likely positively charged based on alignment of protein sequences to the Buster subfamily.

[0040] A mutant TcBuster transposase can comprise one or more amino acid substitutions that increase the net charge at neutral pH compared to SEQ ID NO: 1. Sometimes, a mutant TcBuster transposase that comprises one or more amino acid substitutions that increase the net charge at neutral pH compared to SEQ ID NO: 1 can be hyperactive. Sometimes, a mutant TcBuster transposase can comprise one or more substitutions with a positively charged amino acid, such as, for example, lysine (K) or arginine (R). A mutant TcBuster transposase can comprise one or more substitutions of a negatively charged amino acid, such as, but not limited to, aspartic acid (D) or glutamic acid (E), with a neutral amino acid or a positively charged amino acid.

[0041] One non-limiting example includes a mutant TcBuster transposase that contains one or more amino acid substitutions that increase the net charge at neutral pH within or near the catalytic domain relative to SEQ ID NO: 1. The catalytic domain can be the first catalytic domain or the second catalytic domain. The catalytic domain can also include both catalytic domains of the transposase.

[0042] Exemplary methods of the present disclosure can include mutating amino acids predicted to be close to or in direct contact with DNA. These amino acids can be replacement amino acids identified as conserved in other members of the hAT family (e.g., other members of the Buster and / or Ac subfamilies). Amino acids predicted to be close to or in direct contact with DNA can be identified, for example, by reference to crystal structures, predicted structures, mutational analysis, functional analysis, sequence comparison with other members of the hAT family, or other suitable methods.

[0043] While not wishing to be bound by theory, TcBuster transposases, like other members of the hAT transposase family, possess a DDE motif, which may be the active site that catalyzes transposon movement. D223, D289, and E589 are thought to constitute the active site, a triad of acidic residues. The DDE motif may coordinate divalent metal ions and may be important in catalysis. In some embodiments, mutant TcBuster transposases can contain one or more amino acid substitutions that increase the net charge at neutral pH compared to SEQ ID NO:1, and the one or more amino acids are located near D223, D289, or E589, when numbered according to SEQ ID NO:1.

[0044] In certain embodiments, mutant TcBuster transposases as provided herein do not comprise a disruption of the catalytic triad, i.e., D223, D289, or E589. A mutant TcBuster transposase may not comprise an amino acid substitution at D223, D289, or E589. A mutant TcBuster transposase may comprise an amino acid substitution at D223, D289, or E589, but such substitution does not disrupt the catalytic activity conferred by the catalytic triad.

[0045] In some cases, the term "proximity" can refer to a linear distance measurement in the primary structure of a transposase. For example, the distance between D223 and D289 in the primary structure of wild-type TcBuster transposase is 66 amino acids. In certain embodiments, proximity can refer to a distance of about 70-80 amino acids. Often, proximity can refer to a distance of about 80, 75, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids.

[0046] In some cases, the term "vicinity" can refer to a spatial relationship in the secondary or tertiary structure of a transposase, i.e., when the transposase folds into its three-dimensional configuration. The secondary structure of a protein can refer to the three-dimensional form of a local segment of the protein. Common secondary structure elements include alpha helices, beta sheets, beta turns, and omega loops. Secondary structure elements may form as intermediates before the protein folds into its three-dimensional tertiary structure. The tertiary structure of a protein can refer to the three-dimensional shape of the protein. The tertiary structure of a protein may exhibit dynamic conformational changes under physiological or other conditions. The tertiary structure will have a single polypeptide chain "backbone" with one or more protein domains, which are the secondary structures of the protein. Amino acid side chains may interact and bond in numerous ways. The interactions and bonds of side chains within a particular protein determine its tertiary structure. In many embodiments, vicinity can refer to a distance of about 1 Å, about 2 Å, about 5 Å, about 8 Å, about 10 Å, about 15 Å, about 20 Å, about 25 Å, about 30 Å, about 35 Å, about 40 Å, about 50 Å, about 60 Å, about 70 Å, about 80 Å, about 90 Å, or about 100 Å.

[0047] Neutral pH can be a pH value of about 7. Sometimes, neutral pH can be a pH value of 6.9-7.1, 6.8-7.2, 6.7-7.3, 6.6-7.4, 6.5-7.5, 6.4-7.6, 6.3-7.7, 6.2-7.8, 6.1-7.9, 6.0-8.0, 5-8, or any range derived therefrom.

[0048] Non-limiting exemplary mutant TcBuster transposases comprising one or more amino acid substitutions that increase the net charge at neutral pH relative to SEQ ID NO:1 include TcBuster transposases that comprise at least one combination of amino acid substitutions in Table 4, Table 4.1, or both. The mutant TcBuster transposase can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30 or more amino acid substitutions in Table 4, Table 4.1, or both.

[0049] In some embodiments, the mutant TcBuster transposase can comprise one or more amino acid substitutions relative to SEQ ID NO: 1 that increase the net charge at non-neutral pH. In some cases, the net charge is increased within or near the catalytic domain at non-neutral pH. Often, the net charge is increased near D223, D289, or E589 at non-neutral pH. A non-neutral pH can be a pH value less than 7, less than 6.5, less than 6, less than 5.5, less than 5, less than 4.5, less than 4, less than 3.5, less than 3, less than 2.5, less than 2, less than 1.5, or less than 1. A non-neutral pH can be a pH value greater than 7, greater than 7.5, greater than 8, greater than 8.5, greater than 9, greater than 9.5, or greater than 10. [Table 5] JPEG0007733865000012.jpg16153 [Table 6]

[0050] In an exemplary embodiment, the method can include systematically mutating amino acids in the DNA-binding and oligomerization domain. Without wishing to be bound by theory, mutations in the DNA-binding and oligomerization domain can increase binding affinity to the DNA target and promote oligomerization activity of the transposase, thereby promoting transposition efficiency. More specifically, the method can include systematically mutating amino acids one by one within or near the DNA-binding and oligomerization domain (e.g., amino acids 112-213). The method can also include mutating more than one amino acid within or near the DNA-binding and oligomerization domain. The method can also include mutating one or more amino acids within or near the DNA-binding and oligomerization domain along with one or more amino acids outside the DNA-binding and oligomerization domain.

[0051] In some embodiments, the methods can involve making rational substitutions of selective amino acid residues based on multiple sequence comparisons of TcBuster with other hAT family transposases (Ac, Hermes, Hobo, Tag2, Tam3, Hermes, Restless, and Tol2) or other members of the Buster subfamily (e.g., AeBuster1, AeBuster2, AeBuster3, BtBuster1, BtBuster2, CfBuster1, and CfBuster2). Without being bound by theory, the conservation of certain amino acids among other hAT family transposases, particularly among active transposases, may indicate their importance for the catalytic activity of the transposase. Thus, replacing non-conserved amino acids in the wild-type TcBuster sequence (SEQ ID NO: 1) with amino acids that are conserved among other hAT family members may result in hyperactive mutant TcBuster transposases. The method may include obtaining the sequence of other hAT family transposases, as well as TcBuster; aligning the sequences and identifying amino acids of the TcBuster transposase that differ with conserved counterparts among other hAT family transposases; and performing site-directed mutagenesis to create mutant TcBuster transposases containing the mutation(s) therein.

[0052] A hyperactive mutant TcBuster transposase can comprise one or more amino acid substitutions based on sequence comparison to other members of the Buster subfamily or other members of the hAT family. Often, the one or more amino acid substitutions can be a substitution of a conserved amino acid for a non-conserved amino acid in the wild-type TcBuster sequence (SEQ ID NO: 1). Non-limiting examples of mutant TcBuster transposases include TcBuster transposases that comprise at least one amino acid substitution in Table 5, Table 5.1, or both. A mutant TcBuster transposase can comprise at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least 20, at least 30, or more amino acid substitutions in Table 5, Table 5.1, or both.

[0053] Another exemplary method can involve systematically mutating acidic amino acids to basic amino acids and identifying hyperactive mutant transposases.

[0054] In some cases, the mutant TcBuster transposase can comprise the amino acid substitutions V377T, E469K, and D189A. The mutant TcBuster transposase can comprise the amino acid substitutions K573E and E578L. The mutant TcBuster transposase can comprise the amino acid substitution I452K. The mutant TcBuster transposase can comprise the amino acid substitution A358K. The mutant TcBuster transposase can comprise the amino acid substitution V297K. The mutant TcBuster transposase can comprise the amino acid substitution N85S. The mutant TcBuster transposase can comprise the amino acid substitutions N85S, V377T, E469K, and D189A. The mutant TcBuster transposase can comprise the amino acid substitutions I452F, V377T, E469K, and D189A. The mutant TcBuster transposase can include the amino acid substitutions A358K, V377T, E469K, and D189A.The mutant TcBuster transposase can include the amino acid substitutions V377T, E469K, D189A, K573E, and E578L. [Table 7] JPEG0007733865000015.jpg143153 [Table 8] JPEG0007733865000017.jpg234153 JPEG0007733865000018.jpg37153

[0055] Nuclear localization signal

[0056] Another aspect of the present disclosure provides fusion TcBuster transposases with one or more additional nuclear localization signal (NLS) sequences. Wild-type TcBuster transposase (SEQ ID NO: 1) contains two putative monokaryotic NLS sequences, RKKR and KKRK. In some embodiments of the present disclosure, fusion TcBuster transposases as provided herein can include an additional monokaryotic NLS sequence created via amino acid substitution. In some cases, the additional monokaryotic NLS sequence can have the sequence K(K / R)X(K / R), where X represents any amino acid. In some cases, fusion TcBuster transposases as provided herein that include an additional monokaryotic NLS sequence can have a higher transposition efficiency than an otherwise identical TcBuster transposase that does not have the additional monokaryotic NLS sequence.

[0057] In some embodiments, a fusion TcBuster transposase as provided herein comprises at least one, at least two, at least three, at least four, or at least five additional NLS sequences. In some cases, the additional NLS sequences include those listed in Table 6. As provided herein, the additional NLS sequences can be fused to the N-terminus, C-terminus, and / or internal portions of the TcBuster transposase.

[0058] Exemplary TcBuster transposases comprising a bipartite NLS sequence as provided herein can have an NLS sequence such as K(K / R)XXXXXXXXXXXX(K / R)(K / R)(K / R), K(K / R)XXXXXXXXXXX(K / R)(K / R)(K / R)(K / R), K(K / R)XXXXXXXXXX(K / R)(K / R)(K / R)(K / R), K(K / R)XXXXXXXXX(K / R)(K / R)(K / R), K(K / R)XXXXXXXX(K / R)(K / R)(K / R)(K / R), or K(K / R)XXXXXXX(K / R)(K / R)(K / R)(K / R), where X represents any amino acid. [Table 9]

[0059] Fusion transposase with a DNA-binding domain

[0060] Another aspect of the present disclosure provides a fusion transposase, which can comprise a TcBuster transposase sequence and a DNA sequence-specific binding domain.

[0061] The TcBuster transposase sequence of the fusion transposase can comprise the amino acid sequence of any of the mutant TcBuster transposases described herein. The TcBuster transposase sequence of the fusion transposase can also comprise the amino acid sequence of a wild-type TcBuster transposase having the amino acid sequence SEQ ID NO:1.

[0062] A DNA sequence-specific binding domain, as described herein, can refer to a protein domain adapted to bind to a DNA molecule at a sequence region containing a particular sequence motif (a "target sequence"). For example, an exemplary DNA sequence-specific binding domain may selectively bind to the sequence motif TATA, while another exemplary DNA sequence-specific binding domain may selectively bind to a different sequence motif ATGCNTAGAT (SEQ ID NO: 82), where N represents any one of A, T, G, and C.

[0063] Fusion transposases such as those provided herein may direct sequence-specific insertion of transposons. For example, a DNA sequence-specific binding domain may induce the fusion transposase to bind to a target sequence based on the binding specificity of the binding domain. Binding to or being confined to a specific sequence region may spatially restrict the interaction between the fusion transposase and the transposon, thereby restricting catalyzed transposition to a sequence region near the target sequence. Depending on the size, three-dimensional organization, and sequence binding affinity of the DNA-binding domain, as well as the spatial relationship between the DNA-binding domain and the TcBuster transposase sequence and the flexibility of the connection between the two domains, the distance of the actual transposition site to the target sequence may vary. Appropriate design of the fusion transposase construct can direct transposition to a desired target genomic region.

[0064] The target genomic region for transposition can be any specific genomic region depending on the purpose of the application. For example, it is sometimes desirable to avoid the transcription start site of transposition, which may cause undesirable or even harmful changes in the expression level of certain important endogenous gene(s) of the cell. The fusion transposase may contain a DNA sequence-specific binding domain that can direct transposition to a safe zone in the host genome. Non-limiting examples of safe zones include HPRT, AAVS sites (e.g., AAVS1, AAVS2, ETC), CCR5, or Rosa26. A safe zone site generally refers to a site for transgene insertion whose use has little or no disruptive effect on the integrity of the cell's genome or the health and function of the cell.

[0065] The DNA sequence-specific binding domain may be derived from or a variant of any DNA-binding protein with sequence specificity. In many cases, the DNA sequence-specific binding domain may comprise an amino acid sequence at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to a naturally occurring sequence-specific DNA-binding protein. The DNA sequence-specific binding domain may comprise an amino acid sequence at least 70% identical to a naturally occurring sequence-specific DNA-binding protein. Non-limiting examples of naturally occurring sequence-specific DNA-binding proteins include, but are not limited to, transcription factors, specific sequence nucleases, and viral replication proteins from various sources. The naturally occurring sequence-specific DNA-binding protein may also be any other protein with specific binding ability from various sources. The selection and prediction of DNA-binding proteins can be performed by a variety of approaches, including, but not limited to, the use of computational prediction databases available online, such as DP-Bind (http: / / lcg.rit.albany.edu / dp-bind / ) or DNABIND (http: / / dnabind.szialab.org / ).

[0066] The term "transcription factor" can refer to a protein that controls the rate of transcription of genetic information from DNA to messenger DNA by binding to specific DNA sequences. Transcription factors that can be used in the fusion transposases described herein can be based on prokaryotic or eukaryotic transcription factors, as long as they confer sequence specificity when binding to target DNA molecules. Transcription factor prediction databases, such as DBD (http: / / www.transcriptionfactor.org), can be used to select appropriate transcription factors for application of the present disclosure.

[0067] As used herein, a DNA sequence-specific binding domain can include one or more DNA binding domains derived from naturally occurring transcription factors. Non-limiting examples of DNA binding domains from transcription factors include DNA binding domains belonging to families such as basic helix-loop-helix, basic leucine zipper (bZIP), bipartite response regulator C-terminal effector domains, AP2 / ERF / GCC box, helix-turn-helix, homeodomain proteins, lambda repressor-like, srf-like (serum response factor), pairing box, winged helix, zinc finger, multidomain Cys2His2 (C2H2) zinc finger, Zn2 / Cys6, or Zn2 / Cys8 nuclear receptor zinc finger.

[0068] The DNA sequence-specific binding domain can be an artificially engineered amino acid sequence that binds to a specific DNA sequence. Non-limiting examples of such artificially designed amino acid sequences include sequences created based on frameworks such as transcription activator-like effector nuclease (TALE) DNA binding domains, zinc finger nucleases, adeno-associated virus (AAV) Rep proteins, and other suitable DNA binding proteins as described herein.

[0069] Natural TALEs are proteins secreted by Xanthomonas bacteria that aid in the infection of plant species. Natural TALEs can aid infection by binding to specific DNA sequences and activating the expression of host genes. Generally, TALE proteins consist of a central repeat domain, which determines the specificity of DNA targeting but can be rapidly synthesized de novo. TALEs have a modular DNA-binding domain (DBD) containing a repeat sequence of residues. In some TALEs, each repeat region contains 34 amino acids. As used herein, the term "TALE domain" can refer to the modular DBD of a TALE. The pair of residues at positions 12 and 13 of each repeat region can determine nucleotide specificity and is called the repeat variable dipeptide (RVD). The final repeat region, called the hemi-repeat, is usually truncated to 20 amino acids. Combining these repeat regions allows the synthesis of sequence-specific synthetic TALEs. The C-terminus typically contains a nuclear localization signal (NLS) that targets the TALE to the nucleus, as well as functional domains that regulate transcription, such as an acidic activation domain (AD). The endogenous NLS can be replaced with an organism-specific localization signal. For example, the NLS derived from simian virus 40 large T antigen can be used in mammalian cells. The RVDs HD, NG, NI, and NN target C, T, A, and G / A, respectively. A list of RVDs and their binding preferences under specific nucleotide contexts can be found in Table 7. Additional TALE RVDs can also be used for custom degenerate TALE-DNA interactions. For example, NA has high affinity for all four bases of DNA. Furthermore, * is an RVD with a deletion of the 13th residue, N * can correspond to all DNA letters, including methylated cytosines. * may have the ability to bind to any DNA nucleotide.

[0070] Numerous online tools are available for designing TALEs that target specific DNA sequences, such as TALE-NT (https: / / tale-nt.cac.cornell.edu / ) and MojoHand (http: / / www.talendesign.org / ). Commercially available kits may also be useful for creating custom TALE repeat regions between the N- and C-termini of proteins. These methods can be used to construct custom DBDs, which can then be cloned into expression vectors containing functional domains, such as the TcBuster transposase sequence. [Table 10] JPEG0007733865000021.jpg17153

[0071] TALEs can be synthesized de novo in the laboratory by combining the digestion and ligation steps in a Golden Gate reaction with type II restriction enzymes, for example. Alternatively, TALEs can be constructed by a number of different approaches, including, but not limited to, ligation-independent cloning (LIC), rapid ligation-based automatable solid-phase high-throughput (FLASH) assembly, and iterative cap assembly (ICA).

[0072] Zinc fingers (ZFs) are approximately 30 amino acids long and can bind to a limited set of approximately three nucleotides. The ZF domain of C2H2 may be the most common type of ZF and appears to be one of the most abundantly expressed proteins in eukaryotic cells. ZFs are small, functional, independently folding domains that coordinate their structure with zinc molecules. The amino acids in each ZF can have affinity for specific nucleotides, allowing each finger to selectively recognize three to four nucleotides on DNA. Multiple ZFs can be arranged in a tandem array and can recognize sets of nucleotides on DNA. By using different combinations of zinc fingers, unique DNA sequences within the genome can be targeted. A variety of ZFPs of various lengths can be generated, allowing recognition of almost any desired DNA sequence from 64 possible triplet subsites.

[0073] Zinc fingers used in connection with the present disclosure can be generated using established modular assembly fingers, such as the set of modular assembly finger domains developed by Barbas and coworkers and another set of modular assembly finger domains by ToolGen. Both sets of domains cover all 3-bp GNNs, most ANNs, many CNNs, and some TNN triplets (where N can be any of the four nucleotides). Both have different sets of fingers, allowing for the retrieval and encoding of different ZF modules as needed. The combinatorial selection-based oligomerization pool engineering (OPEN) strategy can also be employed to minimize background-dependent effects on modular assembly, including the position of the finger in the protein and the sequence of adjacent fingers. OPEN ZF arrays are publicly available from the Zinc Finger Consortium database.

[0074] The AAV Rep DNA-binding domain is another DNA sequence-specific binding domain that can be used in connection with the subject matter of the present disclosure. The viral cis-acting inverted terminal repeats (ITRs) and the trans-acting viral Rep protein (Rep) are thought to be factors mediating the preferential integration of AAV into the AAVS1 site of the host genome in the absence of helper virus. The AAV Rep protein can bind to specific DNA sequences in the AAVS1 site. Therefore, site-specific DNA-binding domains can be fused together with the TcBuster transposase domain as described herein.

[0075] Fusion transposases as provided herein can include a TcBuster transposase sequence and a tag sequence. As provided herein, a tag sequence can refer to any protein sequence that can be used as a detection tag for the fusion protein, such as, but not limited to, a reporter protein and affinity tag that can be recognized by an antibody. Reporter proteins include, but are not limited to, fluorescent proteins (e.g., GFP, RFP, mCherry, YFP), β-galactosidase (β-Gal), alkaline phosphatase (AP), chloramphenicol acetyltransferase (CAT), and horseradish peroxidase (HRP). Non-limiting examples of affinity tags include polyhistidine (His tag), glutathione S-transferase (GST), maltose binding protein (MBP), calmodulin binding peptide (CBP), intein-chitin binding domain (intein-CBD), streptavidin / biotin-based tags, epitope tags such as FLAG, HA, c-myc, T7, Glu-Glu, and the like.

[0076] Fusion transposases as provided herein can include a TcBuster transposase sequence and a DNA sequence-specific binding domain or tag sequence fused together without any intermediate sequences (e.g., "back-to-back"). In some cases, fusion transposases as provided herein can include a TcBuster transposase sequence and a DNA sequence-specific binding domain or tag sequence connected by a linker sequence. Figure 8 is a schematic diagram of an exemplary fusion transposase including a DNA sequence-specific binding domain and a TcBuster transposase sequence connected by a linker. In exemplary fusion transposases, the linker may primarily serve as a spacer between the first and second polypeptides. The linker can be a short amino acid sequence that separates multiple domains in a single polypeptide. The linker sequence can include a linker present in a naturally occurring multidomain protein. In some cases, the linker sequence can include an artificially created linker. The selection of the linker sequence can be based on the application of the fusion transposase. The linker sequence can comprise 3, 4, 5, 6, 7, 8, 9, 10, or more amino acids. In some embodiments, the linker sequence can comprise at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, or at least 50 amino acids. In some embodiments, the linker sequence can comprise at most 4, at most 5, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, or at most 15 amino acids. It can comprise at most 20, at most 30, at most 40, at most 50, or at most 100 amino acids. In certain cases, it may be desirable to use flexible linker sequences, such as, but not limited to, stretches of Gly and Ser residues ("GS" linkers), such as (GGGGS)n (n = 2-8) (SEQ ID NO: 83), (Gly)8 (SEQ ID NO: 84), GSAGSAAGSGEF (SEQ ID NO: 85), and (GGGGS)4 (SEQ ID NO: 86).It may be desirable to use rigid linker sequences such as, but not limited to, pro-rich sequences such as (EAAAK)n (n=2-7) (SEQ ID NO: 87), (XP)n, where X designates any amino acid.

[0077] In the exemplary fusion transposases provided herein, the TcBuster transposase sequence can be fused to the N-terminus of a DNA sequence-specific binding domain or tag sequence. Alternatively, the TcBuster transposase sequence can be fused to the C-terminus of a DNA sequence-specific binding domain or tag sequence. In some embodiments, depending on the application of the fusion transposase, a third domain sequence or multiple other sequences can be present between the TcBuster transposase and the DNA sequence-specific binding domain or tag sequence.

[0078] Nucleotide sequence of TcBuster

[0079] Another aspect of the present disclosure provides polynucleotides, e.g., nucleotide sequences, encoding TcBuster transposases as provided herein. In some embodiments, polynucleotides as provided herein include one or more codons that are favored by the translation system of an organism whose cells the polynucleotide is delivered to. For example, polynucleotides as provided herein can include one or more codons that are favored by the human (e.g., Homo Sapiens) translation system when the polynucleotide is delivered to a human cell for genome editing purposes. In some embodiments, one or more codons in a polynucleotide encoding a TcBuster transposase as provided herein can be codons that are more frequently found in the organism whose cells the polynucleotide is delivered to. Without being bound by theory, in some cases, a TcBuster transposase as provided herein is delivered to a target cell in the form of a polynucleotide encoding it, and the frequently occurring codons in the target cell can be more efficiently utilized by the cell's translation system compared to the native codons in the DNA encoding the TcBuster transposase, thereby resulting in increased expression of the TcBuster transposase in the target cell.

[0080] Certain embodiments of polynucleotides as provided herein can include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 codons replaced by codons that are advantageous in the organism whose cells the polynucleotide is delivered to. Can include at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, or at least 200 codons. Certain embodiments of polynucleotides can include one or more codons found at high frequency in Homo sapiens, e.g., those with high frequency / 1000 (or a small fraction) listed in Table 8. In some cases, a codon from Table 8 is selected if its frequency / 1000 in Homo Sapiens is at least 5, at least 6, at least 8, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 25, or at least 30. In some cases, a codon from Table 8 is selected if its fraction in Homo Sapiens is at least 0.1, at least 0.12, at least 0.14, at least 0.16, at least 0.18, at least 0.2, at least 0.22, at least 0.24, at least 0.26, at least 0.28, at least 0.3, at least 0.35, at least 0.4, at least 0.45, at least 0.5, or at least 0.55.

[0081] In some embodiments, the polynucleotides provided herein are codon optimized for expression in cells of a target species, for example, human cells. The polynucleotide can be codon-optimized for expression in cells of the target species, e.g., at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons in the polynucleotide that are present at high frequency in the target species (e.g., at least 5, at least 6, at least 8, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 25, or at least 30 frequencies / 1000 in the target species, or, e.g., at least 0.1, at least 0.12, at least 0.14, at least 0.16, at least 0.18, at least 0.2, at least 0.22, at least 0.24, at least 0.26, at least 0.28, at least 0.3, at least 0.35, at least 0.4, at least 0.45, at least 0.5, or at least 0.55 fraction in the target species). In some embodiments, the polynucleotide is codon-optimized for expression in cells of the target species, e.g., at least 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons in the polynucleotide that are present at a high frequency in the target species (e.g., at least 20, at least 25, or at least 30 frequencies / 1000 in the target species, or at least 0.2, at least 0.22, at least 0.24, at least 0.26, at least 0.28, at least 0.3, at least 0.35, at least 0.4, at least 0.45, at least 0.5, or at least 0.55 fraction in the target species).

[0082] SEQ ID NO:204 is an exemplary DNA sequence that is codon-optimized for expression in human cells and encodes a wild-type TcBuster transposase. SEQ ID NO:207 is an exemplary mRNA sequence that is codon-optimized for expression in human cells and encodes a wild-type TcBuster transposase. The polynucleotides provided herein can comprise a nucleotide sequence that is at least about 70%, 75%, 80%, 85%, 90%, 95%, or 98% identical to or complementary to the full-length SEQ ID NO:204 or 207. In some embodiments, the polynucleotide has a nucleotide sequence that is at least about 80%, 85%, 90%, 95%, or 98% identical to or complementary to the full-length SEQ ID NO:204 or 207. In some embodiments, the polynucleotide has a nucleotide sequence that is at least about 95% identical to or complementary to the full-length SEQ ID NO:204 or 207. [Table 11] JPEG0007733865000023.jpg23153

[0083] TcBuster transposon

[0084] Another aspect of the present disclosure provides a TcBuster transposon comprising a cassette cargo positioned between two inverted repeats, wherein the TcBuster transposon can be recognized by a TcBuster transposase as described herein, e.g., the TcBuster transposase can recognize the TcBuster transposon and catalyze the transposition of the TcBuster transposon into a DNA sequence.

[0085] The terms "inverted repeat," "terminal inverted repeat," and "inverted terminal repeat" used interchangeably herein can refer to short sequence repeats flanking the transposase gene in natural transposons or the cassette cargo in artificially engineered transposons. Two inverted repeat sequences are generally required for transposon mobilization in the presence of the corresponding transposase. Inverted repeat sequences as described herein may contain one or more direct repeat (DR) sequences. These sequences are typically embedded in the terminal inverted repeat (TIR) ​​sequences of the element. The term "cargo cassette" as used herein can refer to a nucleotide sequence other than the native nucleotide sequence between the inverted repeat sequences containing the TcBuster transposase gene. Cargo cassettes can be artificially engineered.

[0086] The transposons described herein may contain a cargo cassette flanked by IR / DR sequences. In some embodiments, at least one of the repeat sequences contains at least one direct repeat sequence. As shown in Figures 1 and 2, the transposon may contain a cargo cassette flanked by IRDR-L-Seq1 (SEQ ID NO: 3) and IRDR-R-Seq1 (SEQ ID NO: 4). In many cases, the left inverted repeat sequence can comprise a sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-L-Seq1 (SEQ ID NO: 3). In some cases, the right inverted repeat sequence can comprise a sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-R-Seq1 (SEQ ID NO: 4). In other cases, the right inverted repeat sequence can comprise a sequence at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-L-Seq1 (SEQ ID NO: 3). Sometimes, the left inverted repeat sequence can comprise a sequence at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-R-Seq1 (SEQ ID NO: 4). As used herein, the terms "left" and "right" can refer to the 5' and 3' sides of a cargo cassette on the sense strand of a double-stranded transposon, respectively. It is also possible that the transposon may contain a cargo cassette flanked by IRDR-L-Seq2 (SEQ ID NO: 5) and IRDR-R-Seq2 (SEQ ID NO: 6).In many cases, the left inverted repeat sequence can comprise a sequence at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-L-Seq2 (SEQ ID NO: 5). In some cases, the right inverted repeat sequence can comprise a sequence at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-R-Seq2 (SEQ ID NO: 6). In other cases, the right inverted repeat sequence can comprise a sequence at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-L-Seq2 (SEQ ID NO: 5). Sometimes, the left inverted repeat can comprise a sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-R-Seq2 (SEQ ID NO:6). Alternatively, the transposon can contain a cargo cassette flanked by IRDR-L-Seq3 (SEQ ID NO:205) and IRDR-R-Seq3 (SEQ ID NO:206). Often, the left inverted repeat sequence can comprise a sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-L-Seq3 (SEQ ID NO:205). In some cases, the right inverted repeat sequence can comprise a sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-R-Seq3 (SEQ ID NO: 206).In other cases, the right inverted repeat can comprise a sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-L-Seq3 (SEQ ID NO: 205). Sometimes, the left inverted repeat can comprise a sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to IRDR-R-Seq3 (SEQ ID NO: 206). A transposon may contain a cargo cassette flanked by two inverted repeats with nucleotide sequences different from those provided in Figure 2, or various sequence combinations known to those skilled in the art. At least one of the two inverted repeat sequences of the transposons described herein may contain a sequence that is at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to any one of SEQ ID NOs: 3-6. At least one of the inverted repeat sequences of the transposons described herein may contain a sequence that is at least 80% identical to SEQ ID NO: 3 or 4. At least one of the inverted repeat sequences of the transposons described herein may contain a sequence that is at least 80% identical to SEQ ID NO: 5 or 6. The choice of inverted repeat sequence may vary depending on the expected transposition efficiency, the type of cell being engineered, the transposase used, and many other factors.

[0087] In many embodiments, the smallest possible inverted terminal repeat sequences of transposon-based vectors that preserve genomic space may be used. The ITRs of hAT family transposons vary significantly due to differences between the right and left ITRs. In many cases, small ITRs consisting of only 100-200 nucleotides are as active as the long native ITRs of hAT transposon-based vectors. These sequences may be consistently reduced while mediating transposition of hAT family transposons. These short ITRs can preserve genomic space within hAT transposon-based vectors.

[0088] The inverted repeat sequences of the transposons provided herein may be about 50 to 2000 nucleotides, about 50 to 1000 nucleotides, about 50 to 800 nucleotides, about 50 to 600 nucleotides, about 50 to 500 nucleotides, about 50 to 400 nucleotides, about 50 to 350 nucleotides, about 50 to 300 nucleotides, about 50 to 250 nucleotides, about 50 to 200 nucleotides, about 50 to 180 nucleotides, about 50 to 160 nucleotides, about 50 to 140 nucleotides, about 50 to 120 nucleotides, about 50 to 110 nucleotides, about 50 to 100 nucleotides, about 50 to 90 nucleotides, about 50 to 110 nucleotides, about 50 to 120 nucleotides, about 50 to 130 nucleotides, about 50 to 140 nucleotides, about 50 to 150 nucleotides, about 50 to 160 nucleotides, about 50 to 170 nucleotides, about 50 to 180 nucleotides, about 50 to 190 nucleotides, about 50 to 210 nucleotides, about 50 to 220 nucleotides, about 50 to 240 nucleotides, about 50 to 250 nucleotides, about 50 to 260 nucleotides, about 50 to 270 nucleotides, about 50 to 280 nucleotides, about 50 to 290 nucleotides, about 50 to 310 nucleotides, about 50 to 320 nucleotides, about 50 to 330 nucleotides, about 50 to 340 nucleotides, about The amino acid sequence can be 0 nucleotides, about 50 to 80 nucleotides, about 50 to 70 nucleotides, about 50 to 60 nucleotides, about 75 to 750 nucleotides, about 75 to 450 nucleotides, about 75 to 325 nucleotides, about 75 to 250 nucleotides, about 75 to 150 nucleotides, about 75 to 95 nucleotides, about 100 to 500 nucleotides, about 100 to 400 nucleotides, about 100 to 350 nucleotides, about 100 to 300 nucleotides, about 100 to 250 nucleotides, about 100 to 220 nucleotides, about 100 to 200 nucleotides, or any range derivable therein.

[0089] In some cases, the cargo cassette may contain a promoter, a transgene, or a combination thereof. In cargo cassettes containing both a promoter and a transgene, expression of the transgene may be directed by the promoter. The promoter may be any type of promoter available to those skilled in the art. Non-limiting examples of promoters that can be used in the TcBuster transposon include EFS, CMV, MND, EF1α, CAGGs, PGK, UBC, U6, H1, and cumate. The choice of promoter to be used in TcBuster transposition depends on many factors, including, but not limited to, the promoter's expression efficiency, the type of cell being genetically engineered, and the desired transgene expression level.

[0090] The transgene in the TcBuster transposon can be any gene of interest and is available to those skilled in the art. The transgene can be derived from a natural gene or its variant, or can be artificially designed. The transgene can be of the same species origin as the engineered cell, or of a different species. The transgene can be a prokaryotic gene or a eukaryotic gene. Sometimes, the transgene can be a gene derived from a non-human animal, a plant, or a human. The transgene can contain an intron. Alternatively, the transgene can be intron-deleted or absent.

[0091] In some embodiments, the transgene can encode a protein. Exemplary proteins include, but are not limited to, cell receptors, immune checkpoint proteins, cytokines, or any combination thereof. Sometimes, the cell receptors described herein can include, but are not limited to, T cell receptors (TCRs), B cell receptors (BCRs), chimeric antigen receptors (CARs), or any combination thereof.

[0092] A cargo cassette as described herein may not contain a transgene encoding any kind of protein product, but may be useful for other purposes. For example, when inserted into an exon of a gene in a host genome, the cargo cassette may be used to create a frameshift at the insertion site. This may lead to a truncated or null mutation in the gene product. Sometimes, a cargo cassette may be used to replace an endogenous genomic sequence with an exogenous nucleotide sequence, thereby manipulating the host genome.

[0093] The transposons described herein may have a cargo cassette in either a forward or reverse orientation. In many cases, the cargo cassette has its own orientation. For example, a cargo cassette containing a transgene has a 5' to 3' coding sequence. A cargo cassette containing a promoter and a gene insert has the promoter at the 5' site of the gene insert. As used herein, the term "forward orientation" can refer to a situation in which the cargo cassette maintains its orientation on the sense strand of a double-stranded transposon. As used herein, the term "reverse orientation" can refer to a situation in which the cargo cassette maintains its orientation on the antisense strand of a double-stranded transposon.

[0094] Systems for genome editing and methods of use

[0095] Another aspect of the present disclosure provides a system for genome editing.The system can include TcBuster transposase and TcBuster transposon.The system can be used for genome editing of host cells, such as disrupting or modifying the endogenous genomic region of a host cell, inserting an exogenous gene into the host genome, replacing an endogenous nucleotide sequence with an exogenous nucleotide sequence, or any combination thereof.

[0096] A genome editing system can include a mutant TcBuster transposase or fusion transposase as described herein and a transposon that can be recognized by the mutant TcBuster transposase or fusion transposase. The mutant TcBuster transposase or fusion transposase can be provided as a purified protein. Protein production and purification techniques are known to those skilled in the art. The purified protein can be stored in a container different from the transposon, or in the same container.

[0097] In many cases, a system for genome editing can include a polynucleotide encoding a mutant TcBuster transposase or fusion transposase as described herein and a transposon recognizable by the mutant TcBuster transposase or fusion transposase. Sometimes, the system polynucleotide can include DNA encoding the mutant TcBuster transposase or fusion transposase. Alternatively or additionally, the system polynucleotide can include messenger RNA (mRNA) encoding the mutant TcBuster transposase or fusion transposase. mRNA can be produced by numerous approaches well known to those of skill in the art, including, but not limited to, in vivo transcription and RNA purification, in vitro transcription, and de novo synthesis. In many cases, mRNA can be chemically modified. Chemically modified mRNA may be more resistant to degradation or may degrade more rapidly than unmodified or native mRNA. In many cases, chemical modification of mRNA can result in more efficient translation of the mRNA. Chemical modification of mRNA can be performed by well-known techniques available to those of skill in the art or by commercial suppliers.

[0098] For many applications, safety dictates that the duration of hAT transposase expression be only long enough to mediate safe transposon delivery. Furthermore, a pulse of hAT transposase expression consistent with high levels of transposon vector can achieve maximal gene delivery. Embodiments are performed using available techniques for in vitro transcription of RNA molecules from DNA plasmid templates. RNA molecules can be synthesized using a variety of methods for in vitro (e.g., cell-free) transcription from DNA copies. Methods for doing this have been described and are commercially available, such as the mMessage Machine in vitro transcription kit available through Life Technologies.

[0099] There are also many companies that offer in vitro transcription for a service-based fee. We have also found that chemically modified RNA for hAT expression is particularly convenient for gene transfer. These chemically modified RNAs are generated using a proprietary method that does not induce or circumvent cellular immune responses. These RNA preparations remove RNA dimers (Clean-Cap) and cellular reactivity (incorporation of pseudouridine) results in better transient gene expression in human T cells without toxicity (data not shown). RNA molecules can be introduced into cells using any of the many methods described for RNA transfection, which are generally non-toxic to most cells. Methods for doing this have been described and are commercially available, such as the Amaxa Nucleofector, Neon Electroporator, and Maxcyte platform.

[0100] A transposon as described herein may be present in an expression vector. In many cases, the expression vector may be a DNA plasmid. Sometimes, the expression vector may be a minicircle vector. As used herein, the term "minicircle vector" may refer to a small circular plasmid derivative that does not contain most, if not all, prokaryotic vector components (e.g., regulatory sequences or nonfunctional sequences of prokaryotic origin). Under certain circumstances, toxicity to cells produced by transfection or electroporation can be reduced by using a "minicircle" as described herein.

[0101] Minicircle vectors can be prepared using well-known molecular cloning techniques. First, a "parent plasmid" (a bacterial plasmid with an insert, such as a transposon construct) can be generated in bacteria such as E. coli, followed by the induction of a site-specific recombinase. Following these steps, the prokaryotic vector can be excised via two recombinase target sequences at either end of the insert, and the resulting minicircle vector can be recovered. The purified minicircle can then be transferred into recipient cells by transfection or lipofection, or into differentiated tissues, for example, by jet injection. Minicircles containing TcBuster transposons can have a size of about 1.5kb, about 2kb, about 2.2kb, about 2.4kb, about 2.6kb, about 2.8kb, about 3kb, about 3.2kb, about 3.4kb, about 3.6kb, about 3.8kb, about 4kb, about 4.2kb, about 4.4kb, about 4.6kb, about 4.8kb, about 5kb, about 5.2kb, about 5.4kb, about 5.6kb, about 5.8kb, about 6kb, about 6.5kb, about 7kb, about 8kb, about 9kb, about 10kb, about 12kb, about 25kb, about 50kb, or a value between any two of these figures. Sometimes, minicircles containing TcBuster transposons as provided herein can have a size of at most 2.1 kb, at most 3.1 kb, at most 4.1 kb, at most 4.5 kb, at most 5.1 kb, at most 5.5 kb, at most 6.5 kb, at most 7.5 kb, at most 8.5 kb, at most 9.5 kb, at most 11 kb, at most 13 kb, at most 15 kb, at most 30 kb, or at most 60 kb.

[0102] In certain embodiments, a system as described herein may contain a polynucleotide encoding a mutant TcBuster transposase or fusion transposase as described herein and a transposon present on the same expression vector, e.g., a plasmid.

[0103] Yet another aspect of the present disclosure provides a method of genetic engineering. The method of genetic engineering can include introducing a TcBuster transposase and a transposon recognizable by the TcBuster transposase into a cell. The method of genetic engineering can also be performed in a cell-free environment. The method of genetic engineering in a cell-free environment can include combining the TcBuster transposase, a transposon recognizable by the transposase, and a target nucleic acid in a container such as a well or test tube.

[0104] The methods described herein can include introducing into a cell a mutant TcBuster transposase provided herein and a transposon recognizable by the mutant TcBuster transposase. A method of genome editing can include introducing into a cell a fusion transposase provided herein and a transposon recognizable by the fusion transposase.

[0105] The mutant TcBuster transposase or fusion transposase can be introduced into a cell as a protein or via a polynucleotide encoding the mutant TcBuster transposase or fusion transposase. As discussed above, the polynucleotide can comprise DNA or mRNA encoding the mutant TcBuster transposase or fusion transposase.

[0106] Often, the TcBuster transposase or fusion transposase can be transfected into host cells as a protein, and the concentration of the protein can be at least 0.05 nM, at least 0.1 nM, at least 0.2 nM, at least 0.5 nM, at least 1 nM, at least 2 nM, at least 5 nM, at least 10 nM, at least 50 nM, at least 100 nM, at least 200 nM, at least 500 nM, at least 1 μM, at least 2 μM, at least 5 μM, at least 7.5 μM, at least 10 μM, at least 15 μM, at least 20 μM, at least 25 μM, at least 50 μM, at least 100 μM, at least 200 μM, at least 500 μM, or at least 1 μM. Sometimes, the concentration of the protein can be from about 1 μM to about 50 μM, from about 2 μM to about 25 μM, from about 5 μM to about 12.5 μM, or from about 7.5 μM to about 10 μM.

[0107] In many cases, the TcBuster transposase or fusion transposase can be transfected into a host cell via a polynucleotide, and the polynucleotide concentration is at least about 5ng / ml, 10ng / ml, 20ng / ml, 40ng / ml, 50ng / ml, 60ng / ml, 80ng / ml, 100ng / ml, 120ng / ml, 150ng / ml, 180ng / ml, 200ng / ml, 220ng / ml, 250ng / ml, 280ng / ml, or , 300ng / ml, 500ng / ml, 750ng / ml, 1μg / ml, 2μg / ml, 3μg / ml, 5μg / ml, 50μg / ml, 100μg / ml, 150μg / ml, 200μg / ml, 250μg / ml, 300μg / ml, 350μg / ml, 400μg / ml, 450μg / ml, 500μg / ml, 550μg / ml, 600μg / ml, 650μg / ml, 700μg / ml, 750μg / ml, or 800μg / ml. Sometimes, the concentration of the polynucleotide can be between about 5-25 μg / ml, 25-50 μg / ml, 50-100 μg / ml, 100-150 μg / ml, 150-200 μg / ml, 200-250 μg / ml, 250-500 μg / ml, 5-800 μg / ml, 200-800 μg / ml, 250-800 μg / ml, 400-800 μg / ml, 500-800 μg / ml, or any range derivable therein.In many cases, the transposon will be present on a separate expression vector from the transponase, and the transposon will be present at a concentration of at least about 5ng / ml, 10ng / ml, 20ng / ml, 40ng / ml, 50ng / ml, 60ng / ml, 80ng / ml, 100ng / ml, 120ng / ml, 150ng / ml, 180ng / ml, 200ng / ml, 220ng / ml, 250ng / ml, 280ng / ml, 300ng / ml, 500ng / ml / ml, 750ng / ml, 1μg / ml, 2μg / ml, 3μg / ml, 5μg / ml, 50μg / ml, 100μg / ml, 150μg / ml, 200μg / ml, 250μg / ml, 300μg / ml, 350μ g / ml, 400 μg / ml, 450 μg / ml, 500 μg / ml, 550 μg / ml, 600 μg / ml, 650 μg / ml, 700 μg / ml, 750 μg / ml, or 800 μg / ml. Sometimes, the transposon concentration can be between about 5-25 μg / ml, 25-50 μg / ml, 50-100 μg / ml, 100-150 μg / ml, 150-200 μg / ml, 200-250 μg / ml, 250-500 μg / ml, 5-800 μg / ml, 200-800 μg / ml, 250-800 μg / ml, 400-800 μg / ml, 500-800 μg / ml, or any range derivable therein. The ratio of transposons to polynucleotides encoding transposases can be at most 10,000, at most 5,000, at most 1,000, at most 500, at most 200, at most 100, at most 50, at most 20, at most 10, at most 5, at most 2, at most 1, at most 0.1, at most 0.05, at most 0.01, at most 0.001, at most 0.0001, or any number between any two thereof.

[0108] In some other cases, the transposon and the polynucleotide encoding the transposase are present in the same expression vector, and the concentration of the expression vector containing both the transposon and the polynucleotide encoding the transposase is at least about 5 ng / ml, 10 ng / ml, 20 ng / ml, 40 ng / ml, 50 ng / ml, 60 ng / ml, 80 ng / ml, 100 ng / ml, 120 ng / ml, 150 ng / ml, 180 ng / ml, 200 ng / ml, 220 ng / ml, 250 ng / ml, 300 ng / ml, 350 ng / ml, 360 ng / ml, 370 ng / ml, 380 ng / ml, 390 ng / ml, 410 ng / ml, 420 ng / ml, 430 ng / ml, 440 ng / ml, 450 ng / ml, 460 ng / ml, 470 ng / ml, 480 ng / ml, 490 ng / ml, 500 ng / ml, 510 ng / ml, 520 ng / ml, 530 ng / ml, 540 ng / ml, 550 ng / ml, 560 ng / ml, 570 ng / ml, 580 ng / ml, 590 ng / ml, 600 ng / ml, 610 ng / ml, 620 ng / ml, 630 ng / ml, 640 ng / ml, 650 ng / ml, 660 ng / ml, 670 ng / ml, 680 ng / ml, 690 ng / ml, 700 ng / ml, 710 ng / ml, 720 ng / ml ng / ml, 280ng / ml, 300ng / ml, 500ng / ml, 750ng / ml, 1μg / ml, 2μg / ml, 3μg / ml, 5μg / ml, 50μg / ml, 100μg / ml, 150μg / ml, 200μg / ml, 250μg / m l, 300 μg / ml, 350 μg / ml, 400 μg / ml, 450 μg / ml, 500 μg / ml, 550 μg / ml, 600 μg / ml, 650 μg / ml, 700 μg / ml, 750 μg / ml, or 800 μg / ml. Sometimes, the concentration of the expression vector containing both the transposon and the polynucleotide encoding the transponase can be between about 5 to 25 μg / ml, 25 to 50 μg / ml, 50 to 100 μg / ml, 100 to 150 μg / ml, 150 to 200 μg / ml, 200 to 250 μg / ml, 250 to 500 μg / ml, 5 to 800 μg / ml, 200 to 800 μg / ml, 250 to 800 μg / ml, 400 to 800 μg / ml, 500 to 800 μg / ml, or any range derivable therein.

[0109] In some cases, the amount of polynucleic acid that may be introduced into cells by electroporation may be varied to optimize transfection efficiency and / or cell viability. In some cases, less than about 100 pg of nucleic acid may be added to each cell sample (e.g., one or more cells to be electroporated). In some cases, at least about 100 pg, at least about 200 pg, at least about 300 pg, at least about 400 pg, at least about 500 pg, at least about 600 pg, at least about 700 pg, at least about 800 pg, at least about 900 pg, at least about 1 microgram, at least about 1.5 μg, at least about 2 μg, at least about 2.5 μg, at least about 3 μg, at least about 3.5 μg, at least about 4 μg, at least about 4.5 μg, at least about 5 μg, at least about 5.5 μg, at least about 6 μg, at least about 6.5 μg, or less. At least about 7 μg, at least about 7.5 μg, at least about 8 μg, at least about 8.5 μg, at least about 9 μg, at least about 9.5 μg, at least about 10 μg, at least about 11 μg, at least about 12 μg, at least about 13 μg, at least about 14 μg, at least about 15 μg, at least about 20 μg, at least about 25 μg, at least about 30 μg, at least about 35 μg, at least about 40 μg, at least about 45 μg, or at least about 50 μg of nucleic acid may be added to each cell sample (e.g., one or more cells to be electroporated). For example, 1 microgram of dsDNA may be added to each cell sample for electroporation. In some cases, the amount of polynucleic acid (e.g., dsDNA) required for optimal transfection efficiency and / or cell viability may be cell type specific.

[0110] The subject matter disclosed herein may be used in genome editing of a wide variety of host cells. In preferred embodiments, the host cells may be derived from eukaryotes. In some embodiments, the cells may be derived from mammalian sources. In some embodiments, the cells may be derived from human sources.

[0111] In general, the cells may be derived from an immortalized cell line or primary cells.

[0112] The terms "cell line" and "immortalized cell line," when used interchangeably herein, can refer to a population of cells derived from an organism that do not normally proliferate indefinitely, but which may, due to mutations, avoid normal cellular senescence and instead continue to undergo division. The subject matter provided herein may be used with a variety of common established cell lines, including, but not limited to, human BC-1 cells, human BJAB cells, human IM-9 cells, human Jiyoye cells, human K-562 cells, human LCL cells, mouse MPC-11 cells, human Raji cells, human Ramos cells, mouse Ramos cells, human RPMI8226 cells, human RS4-11 cells, human SKW6.4 cells, human dendritic cells, mouse P815 cells, mouse RBL-2H3 cells, human HL-60 cells, human NAMALWA cells, human macrophage cells, mouse RAW264.7 cells, human KG-1 cells, mouse M1 cells, human PBMC cells, mouse BW5147(T200-A)5.2 cells, human CCRF-CEM cells, mouse EL4 cells, human Jurkat cells, human SCID.adh cells, human U-937 cells, or any combination of these cells.

[0113] As used herein, the term "primary cells" and its grammatical equivalents can refer to cells taken directly from the living tissue of an organism, typically a multicellular organism such as an animal or plant. Often, primary cells may be established for in vitro growth. In some cases, primary cells may have just been removed from an organism and may not yet be established for in vitro growth prior to transfection. In some embodiments, primary cells may also be grown in vitro, i.e., primary cells may also include progeny cells generated from the growth of cells taken directly from an organism. In these cases, the progeny cells do not exhibit the indefinite growth characteristics of cells in established cell lines. For example, the host cells may be human primary T cells, but prior to transfection, they have been exposed to stimulatory factors that may result in T cell proliferation and expansion of the cell population.

[0114] The cells to be genetically engineered may be primary cells derived from tissues or organs such as, but not limited to, the brain, lung, liver, heart, spleen, pancreas, small intestine, large intestine, skeletal muscle, smooth muscle, skin, bone, adipose tissue, hair, thyroid, trachea, gallbladder, kidney, ureter, urinary bladder, aorta, vein, esophagus, diaphragm, stomach, rectum, adrenal gland, bronchi, ear, eye, retina, reproductive organs, hypothalamus, larynx, nose, tongue, spinal cord, or ureter, uterus, ovary, testis, and any combination thereof. In certain embodiments, the cells include blood cells, hair follicles, keratinocytes, gonadotrophs, corticotrophs, thyrotrophs, somatotrophs, prolactinocytes, chromaffin cells, parafollicular cells, glomus cells, melanocytes, nevus cells, Merkel cells, odontoblasts, cementoblasts, corneal keratinocytes, retinal Müller cells, retinal pigment epithelial cells, neurons, glia, ependymal cells, pineal cells, lung cells, Clara cells, goblet cells, G cells, D cells, enterochromaffin-like cells, gastric chief cells, parietal cells, pituitary cells, K cells, D cells, I cells, Paneth cells, intestinal cells, fold cells, hepatocytes, hepatic stellate cells, gallbladder cells, cochlear cells, pancreatic stellate cells, pancreatic alpha cells, pancreatic beta cells. The present invention may include, but is not limited to, a fibroblast, a pancreatic delta cell, a pancreatic F cell, a pancreatic epsilon cell, a thyroid parathyroid gland, an eosinophil cell, a urothelial cell, an osteoblast, an osteocyte, a chondroblast, a chondrocyte, a fibroblast, a fibrocyte, a myoblast, a muscle cell, a muscle satellite cell, a tenocyte, a cardiomyocyte, a lipoblast, an adipocyte, an interstitial cell of Cajal, angioblast, an endothelial cell, a mesangial cell, a juxtaglomerular cell, a macula densa cell, an interstitial cell, an interstitial cell, a simple epithelial cell, a podocyte, a kidney proximal tubule brush border cell, a Sertoli cell, a Leydig cell, a granulosa cell, a PEG cell, a germ cell, a testicular ovum, a lymphocyte, a bone marrow cell, an endothelial progenitor cell, an endothelial stem cell, angioblast, a mesoangioblast, a pericyte, and any combination thereof. Often, the cells being engineered may be stem cells, such as, but not limited to, embryonic stem cells, hematopoietic stem cells, epidermal stem cells, epithelial stem cells, bronchoalveolar stem cells, mammary stem cells, mesenchymal stem cells, intestinal stem cells, endothelial stem cells, neural stem cells, olfactory adult stem cells, neural crest stem cells, testicular cells, and any combination thereof.Sometimes the cells can be induced pluripotent stem cells derived from any type of tissue.

[0115] In some embodiments, the genetically engineered cells may be mammalian cells. In some embodiments, the cells may be immune cells. Non-limiting examples of cells include B cells, basophils, dendritic cells, eosinophils, gamma delta T cells, granulocytes, helper T cells, Langerhans cells, lymphoid cells, innate lymphoid cells (ILCs), macrophages, mast cells, megakaryocytes, memory T cells, monocytes, myeloid cells, natural killer T cells, neutrophils, progenitor cells, plasma cells, progenitor cells, regulatory T cells, T cells, thymocytes, any differentiated or dedifferentiated cells thereof, or any mixture or combination of such cells. In some embodiments, the cells may be T cells. In some embodiments, the cells may be primary T cells. In some embodiments, the cells may be antigen-presenting cells (APCs). In some embodiments, the cells may be primary APCs. APCs relevant to the present disclosure may be dendritic cells, macrophages, B cells, other non-professional APCs, or any combination thereof.

[0116] In some embodiments, the cells may be ILCs (innate lymphoid cells), and the ILCs can be group 1 ILCs, group 2 ILCs, or group 3 ILCs. Group 1 ILCs may generally be described as cells that are regulated by the T-bet transcription factor and secrete type 1 cytokines, such as IFN-γ and TNF-α, in response to intracellular pathogens. Group 2 ILCs may generally be described as cells that depend on GATA-3 and ROR-alpha transcription factors and produce type 2 cytokines in response to extracellular parasite infection. Group 3 ILCs may generally be described as cells that are regulated by the ROR-gamma t transcription factor and produce IL-17 and / or IL-22.

[0117] In some embodiments, the cells may be cells that are positive or negative for a given factor. In some embodiments, the cells may be CD3+ cells, CD3- cells, CD5+ cells, CD5- cells, CD7+ cells, CD7- cells, CD14+ cells, CD14- cells, CD8+ cells, CD8- cells, CD103+ cells, CD103- cells, CD11b+ cells, CD11b- cells, BDCA1+ cells, BDCA1- cells, L-selectin+ cells, L-selectin- cells, CD25+, CD25- cells, CD27+, CD27- cells, CD28+ cells, CD28- cells, CD44+ cells, CD The cells may be CD44- cells, CD56+ cells, CD56- cells, CD57+ cells, CD57- cells, CD62L+ cells, CD62L- cells, CD69+ cells, CD69- cells, CD45RO+ cells, CD45RO- cells, CD127+ cells, CD127- cells, CD132+ cells, CD132- cells, IL-7+ cells, IL-7- cells, IL-15+ cells, IL-15- cells, lectin-like receptor G1-positive cells, lectin-like receptor G1-negative cells, or differentiated or dedifferentiated cells thereof. The examples of factors expressed by the cells are not intended to be limiting, and one of skill in the art will understand that the cells may be positive or negative for any factor known in the art. In some embodiments, the cells may be positive for more than one factor. For example, the cells may be CD4+ and CD8+. In some embodiments, the cells may be negative for more than one factor. For example, the cells may be CD25-, CD44-, and CD69-. In some embodiments, the cells may be positive for one or more factors and negative for one or more factors, for example, the cells may be CD4+ and CD8-.

[0118] It should be understood that the cells used in any of the methods disclosed herein can be a mixture of any of the cells disclosed herein (e.g., two or more different cells). For example, the methods of the present disclosure can include cells, where the cells are a mixture of CD4+ cells and CD8+ cells. In another example, the methods of the present disclosure can include cells, where the cells are a mixture of CD4+ cells and naive cells.

[0119] As provided herein, transposases and transposons can be introduced into cells via a number of approaches. As used herein, the term "transfection" and its grammatical equivalents generally refer to the process by which nucleic acids are introduced into eukaryotic cells. Transfection methods that can be used in connection with the subject matter include, but are not limited to, electroporation, microinjection, calcium phosphate precipitation, cationic polymers, dendrimers, liposomes, biolistics, FuGene, direct sonication, cell squeezing, optical transfection, protoplast fusion, imparefection, magnetofection, nucleofection, or any combination thereof. In many cases, the transposases and transposons described herein can be transfected into host cells via electroporation. Sometimes, transfection can also be achieved via variations of the electroporation method, such as nucleofection (also known as Nucleofector™ technology). As used herein, the term "electroporation" and its grammatical equivalents can refer to a process in which an electric field is applied to cells to increase the permeability of the cell membrane, allowing chemicals, drugs, or DNA to be introduced into the cells. During electroporation, the electric field is often provided in the form of a "pulse" of very short duration, e.g., 5 milliseconds, 10 milliseconds, and 50 milliseconds. As those skilled in the art will understand, electroporation uses a pulsed rotating electric field to temporarily open pores in the outer membrane of a cell. Methods and devices used for in vitro and in vivo electroporation are also well known. Various electrical parameters, such as pulse strength, pulse length, and pulse number, can be selected depending on the type of cell being electroporated and the physical properties of the molecules to be taken up by the cell.

[0120] Applicable

[0121] The subject matter provided herein, e.g., compositions (e.g., mutant TcBuster transposases, fusion transposases, TcBuster transposases), systems, and methods, may be used in a wide range of applications related to genome editing in various aspects of modern life.

[0122] In certain circumstances, the advantages of the subject matter described herein may include, but are not limited to, reduced costs, regulatory considerations, low immunogenicity, and reduced complexity. In some cases, a key advantage of the present disclosure is high transposition efficiency. Another advantage of the present disclosure is that in many cases, the transposition system provided herein can be "tunable," for example, transposition can be designed to target selected genomic regions rather than random insertion.

[0123] One non-limiting example relates to creating genetically engineered cells for research and clinical applications. For example, as discussed above, genetically engineered T cells can be created using the subject matter provided herein, which may be used to help people fight various diseases, such as, but not limited to, cancer and infectious diseases.

[0124] One specific example includes generating genetically engineered primary leukocytes using the methods provided herein and administering the genetically engineered primary leukocytes to a patient in need thereof. Generating genetically engineered primary leukocytes can include introducing a transposon and a mutant TcBuster transposase or fusion transposase, as described herein, capable of recognizing the transposon into leukocytes, thereby generating genetically engineered leukocytes. In many cases, the transposon may comprise a transgene. The transgene can be a cellular receptor, an immune checkpoint protein, a cytokine, or any combination thereof. Sometimes, the cellular receptor can include, but is not limited to, a T cell receptor (TCR), a B cell receptor (BCR), a chimeric antigen receptor (CAR), or any combination thereof. In other cases, the transposon and transposase are designed to delete or modify an endogenous gene, such as a cytokine, an immune checkpoint protein, an oncogene, or any combination thereof. Genetic manipulation of primary white blood cells can be designed to promote immunity against infectious pathogens or cancer cells that cause morbidity in patients.

[0125] Another non-limiting example relates to creating genetically engineered organisms for agriculture, food production, medicine, and pharmaceutics. Species that can be genetically engineered span a wide range, including, but not limited to, plants and animals. Genetically engineered organisms, such as genetically engineered crops and livestock, may be modified in specific aspects of their physiological characteristics. Examples of food crops include resistance to specific pests, diseases, or environmental conditions, reduced decay, or resistance to chemical treatments (e.g., herbicide resistance), or improved nutritional profiles of crops. Examples of non-food crops include the production of pharmaceuticals, biofuels, and other industrially useful commodities, as well as bioremediation. Examples of livestock include resistance to specific parasites, production of specific nutritional elements, increased growth rates, and increased milk production.

[0126] As used herein, the term "about" and its grammatical equivalents in connection with a reference numerical value can include a range of values ​​from that value plus or minus 10%. For example, the amount "about 10" includes the amount 9 to 11. The term "about" in connection with a reference numerical value can also include a range of values ​​from that value plus or minus 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%. [Example]

[0127] The following examples further illustrate the described embodiments without limiting the scope of the disclosure.

[0128] Example 1. Materials and Methods

[0129] This example describes several methods utilized to generate and evaluate exemplary mutant TcBuster transposases.

[0130] Site-directed mutagenesis for preparing TcBuster variants

[0131] Putative hyperactive TcBuster (TcB) transposase variants were identified by nucleotide and amino acid sequence comparisons with the hAT and Buster subfamilies. The Q5 Site-Directed Mutagenesis Kit (New England BioLabs) was used for all site-directed mutagenesis. Following PCR mutagenesis, PCR products were purified with a GeneJET PCR Purification Kit (Thermo Fisher Scientific). A 20 μL ligation reaction of the purified PCR product was performed using T4 DNA ligase (New England BioLabs). A 5 μL ligation reaction was used to transform DH10Beta cells. The presence of the desired mutation was confirmed using direct colony sequencing by Sequetech. DNA of the confirmed mutations was prepared using a ZymoPURE Plasmid Miniprep Kit (Zymo Research).

[0132] Measurement of transfection efficiency of HEK-293T cells

[0133] HEK-293T cells were plated at 300,000 cells per well in a 6-well plate one day before transfection. Cells were transfected with 500 ng of a transposon carrying an mCherry-puromycin cassette and 62.5 ng of TcB transposase using TransIT X2 reagent according to the manufacturer's instructions (Mirus Bio). Two days after transfection, cells were replated in triplicate in 6-well plates at a density of 3,000 cells / well in DMEM complete medium with or without puromycin (1 μg / mL). Stable integration of the transgene was assessed by colony counting of puromycin-treated cells (individual cells that survived drug selection formed colonies) or by flow cytometry. For colony counting, the DMEM complete + puromycin medium was removed after 2 weeks of puromycin selection. The cells were washed with 1x PBS and stained with 1x crystal violet solution for 10 minutes. The plates were washed twice with PBS and the colonies were counted.

[0134] For flow cytometry analysis, stable integration of the transgene was assessed by detecting mCherry fluorescence in cells grown without drug selection. Transfected cells were harvested at the indicated time points posttransfection, washed once with PBS, and resuspended in 200 μL of RDFII buffer for analysis. Cells were analyzed using Novocyte (Acea Biosciences), and mCherry expression was assessed using the PE-Texas Red channel.

[0135] Screening for TcB transposase mutants in HEK-293T cells

[0136] HEK-293T cells were plated at 75,000 cells per well in a 24-well plate one day before transfection. Cells were transfected with 500 ng transposon and 125 ng transposase in duplicate using TransIT X2 reagent according to the manufacturer's instructions (Mirus Bio). Stable integration of the transgene was assessed by detecting mCherry fluorescence. 14 days after transfection, cells were harvested, washed once with PBS, and resuspended in 200 μL of RDF1I buffer. Cells were analyzed using Novocyte (Acea Biosciences), and mCherry expression was assessed using the PE-Texas Red channel.

[0137] Transfection of TcBuster transposon and transposase in CD3+ T cells

[0138] CD3+ T cells were enriched from Leukopacks (StemCell Technologies) and cryopreserved. CD3+ T cells were thawed and activated for 2 days using CD3 / CD28 Dynabeads (ThermoFisher) in X-Vivo-15 medium supplemented with human serum and the cytokines IL-2, IL-15, and IL-7. Prior to transfection, the CD3 / CD28 beads were removed, the cells were washed, and electroporated with TcBuster transposons (minicircles carrying the TcBuster and Sleeping Beauty IR / DR and GFP cargos) and TcBuster or Sleeping Beauty transposase in RNA form using the Neon transfection system (ThermoFisher). As a viability control, cells were "pulse-electroporated" without DNA or RNA. Electroporated cells were expanded for 21 days posttransfection, and viability and stable integration of the GFP cargo were assessed by flow cytometry. Viability was measured by SSC-A vs. FSC-A and normalized to the pulse-only control, and GFP expression was assessed using the FITC channel at 2, 7, 14, and 21 days.

[0139] Example 2. Exemplary transposon constructs

[0140] The purpose of this study was to examine the transposition efficiency of various exemplary TcBuster transposon constructs. We compared the structures of 10 TcBuster (TcB) transposons (Tn) (Figure 1A) to determine their transposition efficiency in mammalian cells. These 10 TcB Tn constructs differed in the promoters they used (EFS and CMV), IR / DR sequences, and the orientation of the transposon cargo. Each transposon contained an identical cassette encoding mCherry linked to the drug resistance gene puromycin by 2A, allowing transfected cells to be identified by fluorescence and / or puromycin selection. HEK-293T cells were transfected with one of the 10 TcB Tn constructs and the TcB wild-type transposase (at a ratio of one transposon:one transposase). Stable integration of the transgene was assessed by flow cytometry by detection of mCherry fluorescence 10–30 days after transfection (Figure 1B).

[0141] Under the experimental conditions, stable expression of the transgene mCherry was found to be significantly enhanced using the CMV promoter compared to EFS. Transposition appeared to occur only when the 1 IR / DR sequence was used. Transcription of the cargo in the reverse orientation was also found to promote even greater transposition activity compared to the forward orientation.

[0142] TcB Tn-8 demonstrated the highest transposition efficiency by flow cytometry among the 10 Tns tested. To confirm the transposition efficiency of TcB Tn-8, HEK-293T cells were transfected with TcB Tn-8 together with the wild-type transposase or the V596A mutant transposase. Two days after transfection, cells were replated in triplicate in 6-well plates in DMEM complete medium with puromycin (1 μg / mL) at a density of 3,000 cells / well. After 2 weeks of selection, individual cells that survived drug selection formed colonies, which were assessed for mCherry expression (Figure 3A) and counted to confirm stable integration of the transgene (Figure 3B-C). The transposition efficiency of TcB-Tn8 was confirmed by mCherry expression and puromycin-resistant colonies in HEK-293T cells.

[0143] Example 3. Exemplary Transposase Variants

[0144] The aim of this study was to generate TcBuster transposase mutants and examine their transposition efficiency.

[0145] To this end, we generated a consensus sequence by comparing the cDNA and amino acid sequences of wild-type TcB transposase with other similar transposases. For comparison, Sleeping Beauty was reconstituted by aligning the sequences of 13 similar transposases, and SPIN was reconstituted by aligning the sequences of SPIN-like transposases from eight different organisms. SPIN and TcBuster are part of the abundant hAT family of transposases.

[0146] The hAT transposon family consists of two subfamilies: the AC subfamily, which includes hobo, hermes, and Tol2, and the Buster subfamily, which includes SPIN and TcBuster. We aligned the amino acid sequence of TcBuster against those of both AC and Buster subfamily members to identify key amino acids not conserved in TcBuster that could be targets for hyperactivity substitutions. Sequence comparison of TcBuster to AC subfamily members Hermes, Hobo, Tag2, Tam3, Herves, Restless, and Tol2 allowed us to identify amino acids within highly conserved regions that could be substituted in TcBuster (Figure 4). Furthermore, sequence comparison of TcBuster to the Buster subfamily yielded a large number of candidate amino acids that could be substituted (Figure 5). Candidate TcB transposase mutants were generated using oligonucleotides containing the site mutations listed in Table 9. The mutants were then sequence verified and cloned into the pcDNA-DEST40 expression vector (Figure 6), followed by miniprep prior to transfection. [Table 12] JPEG0007733865000025.jpg240153 JPEG0007733865000026.jpg40153

[0147] To examine the transposition efficiency of TcB transposase mutants, HEK-293T cells were transfected in duplicate with either the wild-type or V596A mutant transposase, or with TcB Tn-8 (mCherry-puromycin cassette) along with the candidate transposase mutants. Cells were grown in DMEM complete medium (without drug selection), and mCherry expression was assessed by flow cytometry 14 days after transfection. Over 20 TcB transposase mutants with transposition efficiencies greater than those of the wild-type transposase were identified (Figure 7). Among these mutants, one mutant transposase containing a combination of three amino acid substitutions, D189A, V377T, and E469K, was found to confer substantially increased transposition activity compared to mutants containing each single substitution. Mutants with high transposition activity also included K573E / E578L, I452F, A358K, V297K, N85S, S447E, E247K, and Q258T, among others.

[0148] Among these mutants examined, most substitutions of positively charged amino acids, such as lysine (K) or arginine (R), near one of the catalytic triad amino acids (D234, D289, and E589) increased translocation. In addition, removing a positive charge or adding a negative charge reduced translocation. These data suggest that amino acids near the catalytic domain may help promote the translocation activity of TcB, especially when these amino acids are mutated to positively charged amino acids.

[0149] The amino acid sequence of the hyperactive TcBuster variant D189A / V377T / E469K (SEQ ID NO: 78) is shown in Figure 12. Further mutational analysis of this variant will be performed. As illustrated in Figure 13, the TcBuster variant D189A / V377T / E469K / I452F (SEQ ID NO: 79) will be constructed. As illustrated in Figure 14, the TcBuster variant D189A / V377T / E469K / N85S (SEQ ID NO: 80) will be constructed. As illustrated in Figure 15, the TcBuster variant D189A / V377T / E469K / S358K (SEQ ID NO: 81) will be constructed. As illustrated in Figure 16, the TcBuster variant D189A / V377T / E469K / K573E / E578L (SEQ ID NO: 13) will be constructed. In each of Figures 12-16, the domains of TcBuster are indicated as follows: ZnF-BED (lowercase), DNA-binding / oligonucleotide domain (bold), catalytic domain (underlined), and insertion domain (italic); the core D189A / V377T / E469K substitutions are indicated in large, bold, italic, and underlined letters; additional substitutions are indicated in large, bold. Each of these constructs was tested as previously described and is expected to exhibit hyperactivity compared to wild-type TcBuster.

[0150] Example 4. Exemplary fusion transposases containing tags

[0151] The goal of this study was to generate fusion TcBuster transposases and examine their transposition efficiency. As an example, a protein tag, GST, or PEST domain was fused to the N-terminus of the TcBuster transposase to generate the fusion TcBuster transposase. A flexible linker, GGSGGSGGSGGSGTS (SEQ ID NO: 9), encoded by SEQ ID NO: 10, was used to separate the GST / PEST domain from the TcBuster transposase. The presence of this flexible linker minimizes nonspecific interactions in the fusion protein and may enhance its activity. The exemplary fusion transposase was co-transfected with TcB Tn-8 as described above, and transposition efficiency was measured by mCherry expression on day 14 via flow cytometry. Transposition efficiency was not affected by tagging with either the GFP or PEST domain ( Figure 9 ), suggesting that fusing the DNA-binding domain of a transposase to direct integration of TcBuster cargo to select genomic sites such as safe zone sites may be a viable option for TcBuster to enable a safe integration profile.

[0152] Example 5. Exemplary Fusion Transposases Containing TALE Domains

[0153] The purpose of this study is to generate a fusion TcBuster transposase containing a TALE domain and examine the transposition activity of the fusion transposase. The TALE sequence (SEQ ID NO: 11) is designed to target the human AAVS1 (hAAVS1) site in the human genome. Therefore, the TALE sequence is fused to the N-terminus of wild-type TcBuster transposase (SEQ ID NO: 1) to generate the fusion transposase. A flexible linker, Gly4Ser2 (SEQ ID NO: 88), encoded by SEQ ID NO: 12, is used to separate the TALE domain and the TcBuster transposase sequence. An exemplary fusion transposase has the amino acid sequence of SEQ ID NO: 8.

[0154] The exemplary fusion transposase will be transfected into HeLa cells with TcB Tn-8, as described above, using electroporation. TcB Tn-8 contains the reporter gene mCherry. Transfection efficiency can be examined by flow cytometry two days after transfection to count mCherry-positive cells. Furthermore, next-generation sequencing will be performed to evaluate the insertion site of the mCherry gene in the genome. The designed TALE sequence is expected to mediate targeted insertion of the mCherry gene at a genomic site near the hAAVS1 site.

[0155] Example 6. Transduction efficiency in primary human T cells

[0156] The goal of this study was to develop a TcBuster transposon system for engineering primary CD3+ T cells. To this end, we incorporated an exemplary TcBuster transposon carrying a GFP transgene into a minicircle plasmid. Activated CD3+ T cells were electroporated with the TcB minicircle transposon and an RNA transposase, such as the wild-type TcBuster transposase, and exemplary mutants were selected as described in Example 2. Transgene expression was monitored by flow cytometry for 21 days after electroporation.

[0157] Transposition of the TcB transposon was found to be nearly twofold improved using the exemplary mutants V377T / E469K and V377T / E469K / D189A compared with the wild-type TcBuster transposase and the V596A mutant transposase 14 days after transfection (Figure 10A). Furthermore, the average transposition efficiency with the hyperactive mutants V377T / E469K and V377T / E469K / D189A was two-fold (average = 20.2) and three-fold (average = 24.1) more efficient than SB11 (average = 8.4), respectively.

[0158] Next, the viability of CD3+ T cells was assessed with minicircle TcB transposons and RNA transposase two days after electroporation. Transfection of CD3+ T cells with TcB minicircles and RNA transposase was found to result in a moderate decrease in viability; however, cells rapidly recovered viability by day 7 (Figure 10B). These experiments demonstrate the capability of the TcBuster transposon system according to some embodiments of the present disclosure for cell engineering of primary T cells.

[0159] Example 7. Generation of chimeric antigen receptor-modified T cells for the treatment of cancer patients.

[0160] Minicircle plasmids containing the TcB Tn-8 construct described above can be engineered to harbor a chimeric antigen receptor (CAR) gene between the inverted repeats of the transposon. The CAR can be engineered to have specificity for the B cell antigen CD19 in combination with the signaling domains of CD137 (a costimulatory receptor for T cells [4-1BB]) and CD3-zeta (a signaling component of the T cell antigen receptor).

[0161] Autologous T cells will be obtained from the peripheral blood of patients with cancer, e.g., leukemia. T cells can be isolated by lysing red blood cells and depleting monocytes by centrifugation through a PERCOLL™ gradient. CD3+ T cells can be isolated by flow cytometry using anti-CD3 / anti-CD28 conjugated beads, such as DYNABEAD M-450 CD3 / CD28T. Isolated T cells will be cultured under standard conditions in accordance with GMP guidelines.

[0162] Primary T cells will be genetically engineered using a mutant TcBuster transposase (SEQ ID NO: 13) containing the amino acid substitutions V377T, E469K, D189A, K573E, and E578L, as described above, and a TcBuster Tn-8 transposase containing a CAR. T cells will be electroporated in the presence of the mutant TcBuster transposase and the CAR-containing Tn-8 transposase. After transfection, T cells will be treated with immune stimulatory reagents (e.g., anti-CD3 antibodies, IL-2, IL-7, IL-15) for activation and proliferation. Verification of transfection will be performed by next-generation sequencing two weeks after transfection. The transfection efficiency and transgene load in transfected T cells can be determined to aid in the design of treatment regimens. If the insertion site of a potentially dangerous transgene is revealed by sequencing, specific measures will be taken to eliminate safety concerns.

[0163] Infusion of chimeric antigen receptor-modified T cells (CAR-T cells) back into cancer patients will begin after verification of transgene insertion and in vitro expansion of CAR-T cells to clinically desirable levels.

[0164] The infusion dose will be determined by many factors, including, but not limited to, the stage of the cancer, the patient's treatment history, CBC (complete blood count), and the patient's vital signs on the day of treatment. The infusion dose may be increased or decreased depending on the progression of the disease, the patient's adverse response, and many other medical factors. Meanwhile, during the treatment plan, quantitative polymerase chain reaction (qPCR) analysis will be performed to detect chimeric antigen receptor T cells in the blood and bone marrow. qPCR analysis can be used to make medical decisions regarding dosing strategies and other treatment plans. [Table 13] JPEG0007733865000028.jpg236153 JPEG0007733865000029.jpg236153 JPEG0007733865000030.jpg236153 JPEG0007733865000031.jpg236153 JPEG0007733865000032.jpg237153 JPEG0007733865000033.jpg237153 JPEG0007733865000034.jpg237153 JPEG0007733865000035.jpg237153 JPEG0007733865000036.jpg236153 JPEG0007733865000037.jpg237153 JPEG0007733865000038.jpg237153 JPEG0007733865000039.jpg237153 JPEG0007733865000040.jpg237153 JPEG0007733865000041.jpg237153 JPEG0007733865000042.jpg237153 JPEG0007733865000043.jpg237153 JPEG0007733865000044.jpg237153 JPEG0007733865000045.jpg237153 JPEG0007733865000046.jpg196153

[0165] While preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the subject matter described herein may be employed in practicing the subject matter disclosed herein. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby. [Sequence table free text]

[0166] The contents of array 82 are as follows: <210> 82 <211> 10 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic oligonucleotide <220> <221> modified_base <222> (5)..(5) <223> a, c, t, or g <220> <221> misc_feature <222> (5)..(5) <223> n is a, c, g, or t <400> 82 atgcntagat 10

[0167] The contents of array 208 are as follows: <210> 208 <211> 18 <212> PRT <213> Artificial sequence <220> <223> Synthetic <220> <221> MISC_FEATURE <222> (2)..(2) <223> X = K or R <220> <221> misc_feature <222> (3)..(14) <223> Xaa can be any naturally occurring amino acid <220> <221> MISC_FEATURE <222> (15)..(18) <223> X = K or R <400> 208 Lys Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa 1 5 10 15 Xaa Xaa

[0168] The content of Sequence 209 is as follows. <210> 209 <211> 17 <212> PRT <213> Artificial sequence <220> <223> Synthetic <220> <221> MISC_FEATURE <222> (2)..(2) <223> X = K or R <220> <221> misc_feature <222> (3)..(13) <223> Xaa can be any naturally occurring amino acid <220> <221> MISC_FEATURE <222> (14)..(17) <223> X = K or R <400> 209 Lys Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa 1 5 10 15 Xaa

[0169] The content of array 210 is as follows. <210> 210 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Synthetic <220> <221> MISC_FEATURE <222> (2)..(2) <223> X= K or R <220> <221> misc_feature <222> (3)..(12) <223> Xaa can be any naturally occurring amino acid <220> <221> MISC_FEATURE <222> (13)..(16) <223> X= K or R <400> 210 Lys Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa 1 5 10 15

[0170] The content of array 211 is as follows. <210> 211 <211> 15 <212> PRT <213> Artificial sequence <220> <223> Synthetic <220> <221> MISC_FEATURE <222> (2)..(2) <223> X= K or R <220> <221> misc_feature <222> (3)..(11) <223> Xaa can be any naturally occurring amino acid <220> <221> MISC_FEATURE <222> (12)..(15) <223> X= K or R <400> 211 Lys Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa 1 5 10 15

[0171] The content of Sequence 212 is as follows. <210> 212 <211> 14<​​​​​​​​​​​​​​​​​​​​​​<223> Xaa can be any naturally occurring amino acid <220> <221> MISC_FEATURE <222> (11)..(14) <223> X = K or R <400> 212 Lys Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa 1 5 10

[0172] The content of Sequence 213 is as follows. <210> 213 <211> 13 <212> PRT <213> Artificial sequence <220>[[ID=​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​ [Figure 1-1] FIG. 1 shows the transposition efficiency of several exemplary TcBuster transposon vector constructs as measured by the percentage of mCherry positive cells in cells transfected with wild-type (WT) TcBuster transposase and exemplary TcBuster transposons. [Figure 1-2] Continued from Figure 1-1. [Figure 2-1] FIG. 1 shows a nucleotide sequence comparison of exemplary TcBuster IR / DR Sequence 1 (SEQ ID NOS: 3-4, respectively, in order of appearance) and Sequence 2 (SEQ ID NOS: 5-6, respectively, in order of appearance). [Figure 2-2] Continued from Figure 2-1. [Figure 3] A shows representative brightfield and fluorescent images of HEK-293T cells two weeks after transfection with an exemplary TcBuster transposon, Tn-8 (containing a puro-mCherry cassette; shown in Figure 1 ), and wild-type or V596A mutant transposase (containing a V596A substitution). Transfected cells were plated in 6-well plates with 1 μg / mL puromycin two days after transfection and fixed and stained with crystal violet two weeks after transfection for colony quantification. B shows representative photographs of transfected cell colonies in 6-well plates two weeks after transfection. C is a graph showing colony quantification for each transfection condition two weeks after transfection. [Figure 4] 1 shows a sequence comparison of the amino acid sequence of the TcBuster transposase compared to several transposases in the AC subfamily, with only regions of amino acid conservation shown (SEQ ID NOS: 89-194, respectively, in order of appearance). [Figure 5]

[0023] Figure 1 shows a comparison of the amino acid sequence of the TcBuster transposase relative to a number of other transposase members in the Buster subfamily (SEQ ID NOS: 195-203, respectively, in order of appearance). Specific exemplary amino acid substitutions are indicated above the protein sequence, with the percentage shown above the comparison being the percentage of other Buster subfamily members that contain the amino acid that is believed to be substituted in the TcBuster sequence, and the percentage shown below being the percentage of other Buster subfamily members that contain the standard TcBuster amino acid at that position. [Figure 6] 1 shows the vector map of the exemplary expression vector pcDNA-DEST40 used to test TcBuster transposase variants. [Figure 7] 1 is a graph quantifying the transposition efficiency of exemplary TcBuster transposase variants, as measured by the percentage of mCherry-positive cells in HEK-293T cells transfected with the TcBuster transposon Tn-8 (shown in FIG. 1 ) along with the exemplary transposase variants. [Figure 8] 1 shows one exemplary fusion transposase containing a DNA sequence-specific binding domain and a TcBuster transposase sequence linked by an optional linker. [Figure 9] 1 is a graph quantifying the transposition efficiency of exemplary TcBuster transposases containing different tags, as measured by the percentage of mCherry-positive cells in HEK-293T cells transfected with the TcBuster transposon Tn-8 (shown in FIG. 1 ) along with exemplary transposases containing the tags. [Figure 10] (A) is a graph quantifying the transduction efficiency of an exemplary TcBuster transduction system in human CD3+ T cells as measured by the percentage of GFP-positive cells. (B) is a graph quantifying the viability of transfected T cells 2 and 7 days after transfection by flow cytometry. Data relate to pulse control. [Figure 11]1 shows the amino acid sequence of wild-type TcBuster transposase with specific amino acids annotated (SEQ ID NO: 1). [Figure 12] 1 shows the amino acid sequence of a mutant TcBuster transposase containing the amino acid substitutions D189A / V377T / E469K (SEQ ID NO: 78). [Figure 13] 1 shows the amino acid sequence of a mutant TcBuster transposase containing the amino acid substitutions D189A / V377T / E469K / I452K (SEQ ID NO: 79). [Figure 14] 1 shows the amino acid sequence of a mutant TcBuster transposase containing the amino acid substitutions D189A / V377T / E469K / N85S (SEQ ID NO: 80). [Figure 15] 1 shows the amino acid sequence of a mutant TcBuster transposase containing the amino acid substitutions D189A / V377T / E469K / A358K (SEQ ID NO: 81). [Figure 16] 1 shows the amino acid sequence of a mutant TcBuster transposase containing the amino acid substitutions D189A / V377T / E469K / K573E / E578L (SEQ ID NO: 13).

Claims

1. A mutant TcBuster transposase comprising an amino acid sequence having amino acid substitutions that are at least 90% identical to the full-length sequence of SEQ ID NO:1, and when numbered according to SEQ ID NO:1, having the amino acid substitutions V377T and E469K, and at least one additional amino acid substitution.

2. The mutant TcBuster transposase of claim 1, wherein the mutant TcBuster transposase has a higher transposition efficiency than a wild-type TcBuster transposase having the amino acid sequence of SEQ ID NO:

1.

3. 2. The mutant TcBuster transposase of claim 1, wherein, when numbered according to SEQ ID NO: 1, the at least one additional amino acid substitution is selected from the group consisting of N85S, D99A, N209E, T219S, V356L, C470M, A472P, K490I, C512E, or any combination thereof.

4. 2. The mutant TcBuster transposase of claim 1, wherein, when numbered according to SEQ ID NO: 1, the at least one additional amino acid substitution is selected from the group consisting of D189A, E247K, Q258T, V297K, A358K, S447E, I452F, K573E, E578L, or any combination thereof.

5. 2. The mutant TcBuster transposase of claim 1, wherein the amino acid sequence of the mutant TcBuster transposase is at least 95%, at least 98%, or at least 99% identical to the full-length sequence of SEQ ID NO:

1.

6. A fusion transposase comprising a TcBuster transposase sequence and one or more DNA sequence-specific binding domains and an additional nuclear localization signal sequence, wherein the TcBuster transposase sequence has at least 90% identity to the full-length sequence of SEQ ID NO:1, and when numbered according to SEQ ID NO:1, comprises the amino acid substitutions V377T and E469K, and when numbered according to SEQ ID NO:1, comprises at least one additional amino acid substitution selected from the group consisting of N85S, D99A, D189A, N209E, T219S, E247K, Q258T, V297K, V356L, A358K, S447E, I452F, C470M, A472P, K490I, C512E, K573E, E578L, or any combination thereof.

7. The fusion transposase of claim 6, wherein the DNA sequence-specific binding domain comprises a TALE domain, a zinc finger domain, an AAV Rep DNA-binding domain, or any combination thereof.

8. The fusion transposase of claim 6, wherein the TcBuster transposase sequence has a higher transposition efficiency than the wild-type TcBuster transposase having the amino acid sequence SEQ ID NO:

1.

9. The fusion transposase of claim 6 , wherein the TcBuster transposase sequence and one or more DNA sequence-specific binding domains are separated by a linker.

10. The fusion transposase of claim 9, wherein the linker comprises 3 to 50 amino acids and / or comprises SEQ ID NO:

9.

11. A polynucleotide encoding the fusion transposase of claim 6.

12. 2. A polynucleotide encoding the mutant TcBuster transposase of claim 1, comprising one or more of the following: a) DNA, unmodified messenger RNA (mRNA), or chemically modified mRNA encoding the mutant TcBuster transposase or the fusion transposase of claim 6; b) a nucleic acid sequence that is at least 90% identical to or complementary to the full length of SEQ ID NO: 204 or 207; c) a DNA vector; d) a minicircle plasmid; and / or e) A nucleic acid sequence encoding a transposon that can be recognized by the mutant TcBuster transposase or the fusion transposase of claim 6.

13. 13. The polynucleotide of claim 12, wherein the polynucleotide is codon-optimized for expression in a human cell.

14. A cell containing the polynucleotide of claim 12.

15. A method of genome editing comprising introducing into a cell a mutant TcBuster transposase, a polynucleotide encoding the mutant TcBuster transposase, and a transposon recognizable by the mutant TcBuster transposase, wherein the mutant TcBuster transposase comprises an amino acid sequence that is at least 90% identical to full-length SEQ ID NO: 1, and when numbered according to SEQ ID NO: 1, the amino acid positions are 10. An ex vivo method of genome editing comprising the rearrangements V377T, V377T, and E469K, and when numbered according to SEQ ID NO: 1, comprises at least one additional amino acid substitution selected from the group consisting of N85S, D99A, D189A, N209E, T219S, E247K, Q258T, V297K, V356L, A358K, S447E, I452F, C470M, A472P, K490I, C512E, K573E, E578L, or any combination thereof.

16. 16. The method of claim 15, wherein said introducing comprises transfecting cells with the aid of electroporation, microinjection, calcium phosphate precipitation, cationic polymers, dendrimers, liposomes, biolistics, FuGene®, direct sonication, cell squeezing, optical transfection, protoplast fusion, imparefection, magnetofection, nucleofection, or any combination thereof.

17. 16. The method of claim 15, wherein the cells comprise primary cells obtained from a subject.

18. 18. The method of claim 17, wherein the primary cells are immune cells.

19. A system for genome editing comprising the mutant TcBuster transposase of claim 1 and a transposon that can be recognized by the mutant TcBuster transposase.

20. 20. The system of claim 19, wherein the transposon comprises a cargo cassette positioned between two inverted repeat sequences.

21. 21. The system of claim 20, wherein the left inverted repeat of the two inverted repeats comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO:

3.

22. 21. The system of claim 20, wherein the right inverted repeat of the two inverted repeats comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO:

4.

23. 21. The system of claim 20, wherein the cargo cassette is in a reverse orientation.

Citation Information

Patent Citations

  • Enhanced hAT family transposon-mediated gene transfer and related compositions, systems and methods

    JP2020501612A

  • Enhanced hAT family transposon-mediated gene transfer and related compositions, systems, and methods

    JP2021527427A

  • Development of a transposon system for site-specific DNA integration in mammalian cells

    US20060252140A1