Chimeric transposases and their use
Chimeric transposases with C-terminal cysteine-rich domain deletions and MosI DNA-binding domains address the inefficiencies of existing transposases, enabling precise and efficient gene editing by enhancing site-specific transposition and integration.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- POSEIDA THERAPEUTICS INC
- Filing Date
- 2024-04-04
- Publication Date
- 2026-05-19
AI Technical Summary
The need for site-specific transposases for gene editing is not adequately met by existing transposases, particularly due to the asymmetric bonding of cysteine-rich domains in transposases like piggyBac, which can lead to inefficiencies in transposition and require separate expression from plasmids, leaving empty transposons.
Development of chimeric transposases, such as piggyBac:MosI transposases, with C-terminal cysteine-rich domain deletions and MosI DNA-binding domains, linked by a linker sequence, to enhance site-specific translocation and integration of transposons into genomic DNA.
The chimeric transposases facilitate efficient and targeted integration of transgenes into genomic sites, improving the precision and efficiency of gene editing processes.
Smart Images

Figure 2026515654000001 
Figure 2026515654000002 
Figure 2026515654000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 494,297, filed on April 5, 2023, which is hereby incorporated by reference in its entirety.
[0002] Reference to Electronically Submitted Sequence Listing This application includes a sequence listing submitted in XML format via Patent Center, which is hereby incorporated by reference in its entirety. The XML copy created on March 27, 2024, is named "POTH - 080_001WO_SeqList" and is 233,942 bytes in size.
[0003] The present disclosure generally relates to chimeric transposases, particularly chimeric transposases, chimeric site - specific transposase fusion proteins comprising a chimeric transposase, a chimeric transposase, and a DNA - targeting domain, and chimeric transposon inverted repeat (ITR) polynucleotides. Methods of using chimeric site - specific transposase fusion proteins for site - specific translocation are also provided.
Background Art
[0004] Transposases can be used to introduce non - endogenous DNA sequences into genomic DNA and are advantageous in many respects over other gene editing methods. However, the need for, for example, site - specific transposases for use in gene editing is not met.
[0005] In their native forms, PiggyBac transposons and related family members have a single open reading frame that encodes their respective transposases. Inverted terminal repeat (ITR) sequences are present at both ends of the transposon and act as transposase binding sites. Just outside the ITRs, a 4-nucleotide sequence TTAA acts as a cleavage site, allowing the transposon to jump from one TTAA site in the genome to another via a cut-and-paste mechanism.
[0006] When used for genome editing purposes, transposases are often removed from the transposon and expressed from a separate plasmid or mRNA. This leaves an empty transposon that can encode the desired "cargo" or "payload" sandwiched between ITR and TTAA sequences. This synthetic transposon can be delivered to the cell along with the transposase, which catalyzes the transposon's transposition to the TTAA site in the cell's genome.
[0007] PiggyBac transposases and related family members contain two distinct DNA-binding domains that interact with the inverted end (ITR) sequence of the transposon. The DNA-binding and dimerization domain (DDBD) binds to the proximal sequence of the TTAA site adjacent to the transposon, while the C-terminal cysteine-rich domain (CRD) of the protein binds to the distal repeat (ITR) sequence of the TTAA site. During transposase dimerization, the DDBD binds symmetrically to the ITR, with one monomer of the DDBD binding to the left-end (LE) ITR and the second monomer of the DDBD binding to the right-end (RE) ITR. Interestingly, the two CRDs of the dimer bind asymmetrically, both binding to the LE ITR. Based on ITR sequence identity, it has been proposed that the second dimer of piggyBac transposase may bind distally to the first dimer. Both CRDs of this second dimer have predicted binding sites within the RE ITR. The proximal dimer catalyzes rearrangement reactions, but the role of the distal dimer remains unknown.
[0008] The asymmetric bonding of the two CRD sequences in piggyBac transposase is not characteristic of all transposase families. For example, MosI, a member of the mariner / Tc1 family of transposases, interacts with its ITR as a symmetric dimer, as observed in its crystal structure (see, e.g., Morris et al., Elife 2016 May 25;5:e15537.doi:10.7554 / eLife.15537). Transposases that function as symmetric dimers may have several advantageous features compared to transposases that function as asymmetric dimer pairs (e.g., piggyBac transposase), including a shorter ITR sequence required to complete the rearrangement and fewer transposase molecules. [Overview of the project]
[0009] In one embodiment, a chimeric transposase is provided herein, comprising, in order from the N-terminus to the C-terminus: (i) a target-specific DNA-binding domain; (ii) a cleaved Super piggyBac (SPB) transposase containing a C-terminal cysteine-rich domain (CRD) deletion within amino acid residues 535-594 of the SPB sequence containing the sequence shown in SEQ ID NO: 1; and (iii) one or more MosI DNA-binding domains. In some embodiments, the cleaved PB transposase of (i) contains the sequence shown in any one of SEQ ID NOs: 68-84. In some embodiments, the cleaved PB transposase contains one or more hyperactive mutations. In some embodiments, one or more hyperactive mutations are amino acid substitutions selected from I30V, S103P, G165S, M226F, M282V, S509G, N538K, or N571S of SEQ ID NO: 1.
[0010] In some embodiments, the PB transposase containing the CRD deletion further includes an in-frame N-terminal nuclear localization sequence (NLS) containing the amino acid sequence of SEQ ID NO: 39. In some embodiments, the chimeric transposase contains two MosI DNA-binding domains. In some embodiments, the two MosI DNA-binding domains contain the amino acid sequence shown in SEQ ID NO: 6. In some embodiments, the chimeric transposase contains one MosI DNA-binding domain. In some embodiments, the one MosI DNA-binding domain contains the amino acid sequence shown in SEQ ID NO: 8.
[0011] In some embodiments, the PB transposase and one or more MosI DNA-binding domains are linked by a linker sequence. In some embodiments, the linker sequence is GGGGS (SEQ ID NO: 86). In some embodiments, the target-specific DNA-binding domain is a zinc finger domain or a TAL array.
[0012] In another embodiment, polynucleotides comprising nucleic acid sequences encoding the chimeric transposase described herein are provided herein.
[0013] In another embodiment, a vector comprising a polynucleotide provided herein is provided herein.
[0014] In another embodiment, a chimeric transposon inverted terminal repeat (ITR) polynucleotide is provided herein, comprising, in the 5' to 3' direction: (i) a polynucleotide containing nucleotides 1 to 16 of the nucleic acid sequence shown in SEQ ID NO: 12, and (ii) a polynucleotide containing two MosI DNA binding sites. In some embodiments, the polynucleotide containing the two MosI DNA binding sites is fused in reverse. In some embodiments, the polynucleotide containing the two MosI DNA binding sites contains the nucleic acid sequence of SEQ ID NO: 13 or 14. In some embodiments, the chimeric transposon ITR polynucleotide contains the nucleic acid sequence of SEQ ID NO: 15 or 16.
[0015] In some embodiments, the chimeric transposon ITR polynucleotide further comprises one or two additional nucleotides between the polynucleotide containing nucleotides 1-16 of the nucleic acid sequence shown in SEQ ID NO: 12 and the polynucleotide containing two MosI binding sites. In some embodiments, the chimeric transposon ITR polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 17. In some embodiments, the chimeric transposon ITR polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 18.
[0016] In another embodiment, a vector comprising a chimeric transposon ITR polynucleotide provided herein is provided herein.
[0017] In another embodiment, transposons comprising the chimeric transposon ITR polynucleotide described herein are provided herein.
[0018] In another embodiment, a chimeric transposon ITR polynucleotide is provided herein, comprising, in the 5' to 3' direction, (i) a polynucleotide containing nucleotides 1 to 16 of the nucleic acid sequence shown in SEQ ID NO: 65 and (ii) a polynucleotide containing two MosI DNA binding sites. In some embodiments, the polynucleotide containing the two MosI DNA binding sites is fused in the reverse direction. In some embodiments, the polynucleotide containing the two MosI DNA binding sites contains the nucleic acid sequence of SEQ ID NO: 13 or 14. In some embodiments, the chimeric transposon ITR polynucleotide contains the nucleic acid sequence of SEQ ID NO: 19 or 20. In some embodiments, the chimeric transposon ITR polynucleotide further comprises one or two additional nucleotides between the polynucleotide containing nucleotides 1 to 16 of the nucleic acid sequence shown in SEQ ID NO: 65 and the polynucleotide containing the two MosI DNA binding sites. In some embodiments, the chimeric transposon ITR polynucleotide contains the nucleic acid sequence of SEQ ID NO: 21 or 22.
[0019] In another embodiment, a vector comprising a chimeric transposon ITR polynucleotide described herein is provided herein.
[0020] In another embodiment, transposons comprising the chimeric transposon ITR polynucleotide described herein are provided herein.
[0021] In another embodiment, a chimeric site-specific transposase fusion protein is provided herein, comprising, in order from the N-terminus to the C-terminus: (i) a target-specific DNA-binding domain; (ii) an N-terminal deletion of SEQ ID NO: 1, one or more embedded deletion mutations of SEQ ID NO: 1, and a C-terminal cysteine-rich domain (CRD) deletion within residues 535-594 of the amino acid sequence of SEQ ID NO: 1; and (iii) one or more MosI DNA-binding domains. In some embodiments, the target-specific DNA-binding domain is a TAL or zinc finger motif (ZFM).
[0022] In some embodiments, the cleavage-type PB transposase contains the sequence shown in any one of SEQ ID NOs: 68-84. In some embodiments, the N-terminal deletion of the cleavage-type PB transposase is the deletion of amino acid residues 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103 of SEQ ID NO: 1. In some embodiments, the N-terminal deletion of the cleavage-type PB transposase is the deletion of amino acid residues 1-93. In some embodiments, the PB transposase contains one or more hyperactive mutations. In some embodiments, one or more hyperactive mutations are amino acid substitutions selected from G165S, M226F, M282V, and N538K (numbering begins at residue 5) of SEQ ID NO: 1. In some embodiments, one or more embedded deletion mutations of the cleavage-type PB transposase are amino acid substitutions selected from R372A, K375A, and D450N of SEQ ID NO: 1. In some embodiments, the N-terminal deletion-cleavage-type PB transposase contains the amino acid sequence of SEQ ID NO: 85.
[0023] In some embodiments, the PB transposase containing a C-terminal cysteine-rich domain (CRD) deletion further includes an in-frame N-terminal nuclear localization sequence (NLS), the NLS comprising the amino acid sequence shown in SEQ ID NO: 39. In some embodiments, the chimeric site-specific transposase fusion protein comprises two MosI DNA-binding domains. In some embodiments, the two MosI DNA-binding domains comprise the amino acid sequence shown in SEQ ID NO: 6. In some embodiments, the chimeric site-specific transposase fusion protein comprises one MosI DNA-binding domain. In some embodiments, the one MosI DNA-binding domain comprises the amino acid sequence shown in SEQ ID NO: 8.
[0024] In some embodiments, the PB transposase and one or more MosI DNA-binding domains are linked by a linker sequence. In some embodiments, the linker sequence is GGGGS (SEQ ID NO: 86).
[0025] In some embodiments, the target-specific DNA-binding domain is a zinc finger domain or a TAL array.
[0026] In another embodiment, a LINE1-targeted chimeric site-specific transposase fusion protein is provided herein, comprising (i) a left or right LINE1 TAL-targeted DNA-binding domain, (ii) a linker sequence, and (iii) a chimeric PB:MosI transposase, in the direction from the N-terminus to the C-terminus. In some embodiments, the chimeric PB:MosI transposase comprises the amino acid sequence shown in SEQ ID NO: 9 or 7. In some embodiments, the left or right LINE1 TAL-targeted DNA-binding domain comprises the amino acid sequence shown in SEQ ID NO: 66 or 67.
[0027] In another embodiment, polynucleotides comprising nucleic acid sequences encoding the chimeric site-specific transposase fusion proteins described herein are provided herein.
[0028] In another embodiment, a vector comprising a nucleic acid sequence encoding a chimeric site-specific transposase fusion protein described herein is provided herein.
[0029] In another embodiment, a transposon comprising (i) a chimeric LE PB:MosI ITR polynucleotide and (ii) a chimeric RE PB:MosI ITR polynucleotide is provided herein, wherein the LE PB:MosI ITR polynucleotide and the RE PB:MosI ITR polynucleotide contain the same nucleic acid sequence. In some embodiments, the chimeric LE PB:MosI ITR polynucleotide and the chimeric RE PB:MosI ITR polynucleotide each contain one of the nucleic acid sequences of SEQ ID NOs: 15 to 22.
[0030] In another embodiment, transposons comprising (i) a chimeric LE PB:MosI ITR polynucleotide containing the nucleic acid sequence of SEQ ID NO: 54 and (ii) a chimeric RE PB:MosI ITR polynucleotide containing the nucleic acid sequence of SEQ ID NO: 55 are provided herein.
[0031] In another embodiment, a transposon comprising the nucleic acid sequence of SEQ ID NO: 56 is provided herein.
[0032] In another embodiment, a method for incorporating a transgene into a genomic target site of a cell is provided herein, comprising introducing a chimeric transposase and transposon disclosed herein into a cell, wherein the transposon comprises, in the order 5' to 3', a 5'ITR, a transgene, and a 3'ITR, where the 5'ITR is a chimeric transposon inverted terminal repeat (ITR) described herein and the 3'UTR is a chimeric transposon inverted terminal repeat (ITR) described herein. In some embodiments, the 5'ITR chimeric transposon inverted terminal repeat sequence comprises the nucleic acid sequence shown in SEQ ID NO: 15. In some embodiments, the 3'ITR chimeric transposon inverted terminal repeat sequence comprises the nucleic acid sequence shown in SEQ ID NO: 21. In some embodiments, the transposon further comprises an exogenous promoter between the 5'ITR and the transgene.
[0033] In some embodiments, the transgene encodes a detectable marker. In some embodiments, the detectable marker is GFP.
[0034] In some embodiments, the genome target site is located in a repeating element. In some embodiments, the repeating element is a LINE element.
[0035] In another embodiment, a method for site-specific transposition of a DNA molecule into the genome of a cell is provided herein, comprising introducing into a cell a nucleic acid encoding a chimeric site-specific transposase fusion protein comprising a piggyBac transposase comprising a C-terminal cysteine-rich domain (CRD) deletion and a chimeric transposase comprising one or more MosI DNA-binding domains in the order of N-terminus to C-terminus; (the fusion protein is expressed in the cell) and (b) a DNA molecule comprising a transposon comprising a chimeric LE transposon ITR polynucleotide and a chimeric RE transposon ITR polynucleotide, wherein the expressed chimeric site-specific transposase fusion protein incorporates the transposson into the TTAA sequence of the cell genome by site-specific transposition.
[0036] In another embodiment, a method for generating cells engineered by site-directed transposition is provided herein, comprising introducing into the cells a nucleic acid encoding a chimeric site-directed transposase fusion protein comprising a piggyBac transposase having a C-terminal cysteine-rich domain (CRD) deletion and a chimeric transposase having one or more MosI DNA-binding domains in the order of N-terminus to C-terminus (the chimeric site-directed transposase fusion protein is expressed in the cells) and (b) a DNA molecule comprising transposons comprising a chimeric LE transposon ITR polynucleotide and a chimeric RE transposon ITR polynucleotide, wherein the expressed chimeric site-directed transposase fusion protein incorporates the transposons into the TTAA sequence of the cell's genome by site-directed transposition, thereby generating the engineered cells. [Modes for carrying out the invention]
[0037] This specification provides chimeric transposases and chimeric site-specific fusion proteins, in particular chimeric piggyBac:MosI (PB:MosI) transposases and chimeric PB:MosI site-specific fusion proteins, polynucleotides encoding chimeric transposases and chimeric site-specific fusion proteins, and vectors and transposons containing such polynucleotides. Also provided are methods for producing chimeric transposases and chimeric fusion proteins, cells modified using the chimeric transposases or chimeric site-specific fusion proteins provided herein, and methods for using such cells.
[0038] Chimeric transposon inverted terminal repeat (ITR) sequences are also provided herein. In some embodiments, the chimeric transposon inverted terminal repeat (ITR) sequence is the left-end inverted terminal repeat (ITR) sequence of a chimeric PB:MosI transposon. In some embodiments, the chimeric transposon inverted terminal repeat (ITR) sequence is the right-end inverted terminal repeat (ITR) sequence of a chimeric PB:MosI transposon. Transposons containing chimeric left-end and right-end ITRs, as well as methods and uses of transposons containing the chimeric transposon left-end and right-end ITR sequences provided herein in a method of transposing cells, are also provided.
[0039] Chimera piggyBac:MosI transposase In one embodiment, a piggyBac transposase comprising a C-terminal cysteine-rich domain (CRD) deletion in the order of N-terminus to C-terminus and a chimeric transposase comprising one or more MosI DNA-binding domains are provided herein.
[0040] An exemplary sequence of the wild-type PB transposase is shown in SEQ ID NO: 1. (Sequence ID 1)
[0041] In one embodiment, a chimeric transposase is provided herein that, in order from the N-terminus to the C-terminus, includes: (i) a piggyBac transposase containing a C-terminal cysteine-rich domain (CRD) deletion and (ii) one or more MosI DNA-binding domains. In some embodiments, the PB CRD consists of residues 553-594 of the 594-amino acid PB transposase protein (i.e., in some embodiments, the PB CRD includes the amino acid sequence shown in SEQ ID NO: 2). The CRD domain may be bound to the rest of the N-terminus of the transposase via a linker sequence. In some embodiments, the linker sequence extends to residues 535-552 of the PB transposase sequence (including residues 535 and 552). In some embodiments, the linker is the amino acid sequence ILPNEVPGTSDDSTEEPV (Sequence ID 171)Includes. In some embodiments, the linker is an amino acid sequence ILPKEVPGTSDDSTEEPV (Sequence ID 174) Includes.
[0042] In some embodiments, the C-terminal CRD deletion of piggyBac transposase is a deletion of C-terminal amino acid residues 545-594 of the PB transposase amino acid sequence shown in SEQ ID NO: 1. In such embodiments, the resulting cleaved piggyBac transposase has the following amino acid sequence: (Sequence ID 3)
[0043] In some embodiments, the PB transpose containing the CRD deletion described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the PB transpose containing the CRD deletion described herein includes an amino acid sequence shown in SEQ ID NO: 3 having one, two, three, or four of the five conservative amino acid substitutions. In some embodiments, the PB transpose containing the CRD deletion described herein includes an amino acid sequence shown in SEQ ID NO: 3.
[0044] The MosI DNA-binding domain used to generate this C-terminal CRD deletion-containing chimeric PB:MosI transposase can be bound to amino acid S544 of the PB transposase sequence shown in SEQ ID NO: 1. The resulting chimeric transposase is referred to herein as the "chimeric S544 PB:MosI transposase". In some embodiments, the S544 PB transposase sequence and the MosI DNA-binding domain are linked by a linker sequence. In some embodiments, the linker sequence is GGGGS (SEQ ID NO: 86).
[0045] In some embodiments, the deletion of the C-terminal cysteine-rich domain (CRD) of the PB transposase is the deletion of amino acid residues 553-594 of the PB transposase amino acid sequence shown in SEQ ID NO: 1. In such embodiments, the resulting cleaved piggyBac transposase has the following amino acid sequence: (Sequence No. 4)
[0046] In some embodiments, the PB transpose containing the CRD deletion described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 4. In some embodiments, the PB transpose containing the CRD deletion described herein includes an amino acid sequence shown in SEQ ID NO: 4 having one, two, three, or four of the five conservative amino acid substitutions. In some embodiments, the PB transpose containing the CRD deletion described herein includes an amino acid sequence shown in SEQ ID NO: 4.
[0047] Alternatively, the MosI DNA-binding domain used to generate a chimeric PB:MosI transposase containing this C-terminal CRD deletion may be bound to amino acid V552 of the piggyBac transposase sequence shown in SEQ ID NO: 1. The resulting chimeric transposase is referred to herein as the "chimeric V552 PB:MosI transposase". In some embodiments, the V552 PB transposase sequence and the MosI DNA-binding domain are linked by a linker sequence. In some embodiments, the linker sequence is GGGGS (SEQ ID NO: 86).
[0048] In some embodiments, the cleavage-type PB transposase contains one or more hyperactive mutations. Examples of hyperactive mutations include I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S of SEQ ID NO: 1. In some embodiments, the cleavage-type piggyBac transposase contains at least four hyperactive mutations selected from I30V, G165S, M226F, M282V, and N538K of SEQ ID NO: 1.
[0049] In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 1 and one of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 1 and two of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 1 and three of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 1 and four of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 1 and five of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 1 and six of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 1 and seven of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S.
[0050] In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 3 and one of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 3 and two of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 3 and three of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 3 and four of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 3 and five of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 3 and six of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 3 and seven of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S.
[0051] In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 4 and one of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 4 and two of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 4 and three of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 4 and four of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 4 and five of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 4 and six of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S. In some embodiments, the cleavage-type PB transposase described herein comprises the sequence of SEQ ID NO: 4 and seven of the following mutations: I30V, S103P, G165S, M226F, M282V, S509G, N538K, and N571S.
[0052] In some embodiments, the cleaved piggyBac transposase domain is Super piggyBac® transposase (SPB). Non-limiting examples of SPB transposases are described in detail in U.S. Patents 6,218,182; 6,962,810; 8,399,643 and PCT Publication No. WO2010 / 099296, each of which is incorporated herein by reference in whole, as examples of SPB transposases that may be used in the fusion proteins described herein.
[0053] An exemplary Super piggyBac transposase-MosI chimeric transposase is further described in Example 6.
[0054] In some embodiments, the cleaved PB transposase sequence further comprises an in-frame N-terminal nuclear localization sequence (NLS). An exemplary wild-type SPB sequence containing the NLS is shown in Sequence ID No. 5, where the NLS is italicized and the hyperactive mutation is shown in bold, with the cysteine-rich domain (CRD) deleted in the cleaved PB transposase sequence being replaced by underline The sequence numbering of the SPB transposase domain for the purpose of describing deletions and mutations begins at residue 12 of sequence number 5. MAPKKKRKVGGGGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIY DPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQ LLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPA KMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLDQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPV MKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (Sequence ID 5).
[0055] The transposases described herein can be isolated or derived from insects, vertebrates, crustaceans, or chordates, as described in detail in PCT publication numbers WO2019 / 173636 and PCT / US2019 / 049816. In preferred embodiments, SPB transposases are isolated or derived from the insects Trichoplusia ni (GenBank accession number AAA87375), Macdunnoughia crassisigna (ABZ85926.1), or Bombyx mori (GenBank accession number BAD11135).
[0056] In some embodiments, the chimeric transposase includes two MosI DNA-binding domains. In some embodiments, each of the two MosI DNA-binding domains includes amino acid residues 3-111 of the MosI transposase (i.e., the amino acid sequence shown in SEQ ID NO: 6): SFVPNKEQTRTVLIFCFHLKKTAAESHRMLVEAFGEQVPTVKTCERWFQRFKSGDFDVDDKEHGKPPKRYEDAELQALLDEDDAQTQKQLAEQLEVSQQAVSNRLREMG (SEQ ID NO: 6). In some embodiments, each of the two MosI DNA-binding domains includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 6. In some embodiments, each of the two MosI DNA-binding domains contains the amino acid sequence shown in Sequence ID No. 3, which has one, two, three, or four of the five conserved amino acid substitutions.
[0057] In some embodiments, the chimeric transposase comprises a cleaved piggyBac transposase domain containing the amino acid sequence shown in SEQ ID NO: 3, linked to two MosI DNA-binding domains, each of which contains the sequence shown in SEQ ID NO: 6. The two MosI DNA-binding domains may also be linked to a cleaved PB transposase domain via a linker sequence, resulting in a chimeric PB:MosI S544-111 transposase containing the amino acid sequence shown in SEQ ID NO: 7 (remainders 3-111 of the MosI transposase are highlighted in bold, the PB domain is underlined, and the linker sequence is highlighted in italics): MAPKKKRKVGG GGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFK LFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLL GFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKP KPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLDQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSGGGGSSFVPNKEQTRTVLIFCFHLKKTAAESHRMLVEAFGEQVPTVKTCERWFQRFKSGDFDVDDKEHGKPPKRYEDAELQALLDEDDAQTQKQLAEQLEVSQQAVSNRLREMG (Sequence ID 7)
[0058] In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 7. In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains the amino acid sequence shown in SEQ ID NO: 7 having one, two, three, or four of the five conservative amino acid substitutions. In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains the amino acid sequence shown in SEQ ID NO: 7.
[0059] In some embodiments, the chimeric transposase includes one MosI DNA-binding domain. In some embodiments, the one MosI DNA-binding domain includes amino acid residues 3-58 of the MosI transposase, i.e., the amino acid sequence shown in SEQ ID NO: 8. SFVPNKEQTRTVLIFCFHLKKTAAESHRMLVEAFGEQVPTVKTCERWFQRFKSGDF (Sequence ID 8).
[0060] In some embodiments, a chimeric transposase containing a cleaved piggyBac transposase domain having the amino acid sequence shown in SEQ ID NO: 3 is linked to a single MosI DNA-binding domain having the amino acid sequence shown in SEQ ID NO: 8. The cleaved PB transposase domain may also be linked to the MosI DNA-binding domain via a linker sequence, resulting in a chimeric PB:MosI S544-58 transposase having the amino acid sequence shown in SEQ ID NO: 9 (residues 3-58 of the MosI transposase are highlighted in bold, the PB domain is underlined, and the linker sequence is highlighted in italics): MAPKKKRKVG GGGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLDQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTS GGGGSSFVPNKEQTRTVLIFCFHLKKTAAESHRMLVEAFGEQVPTVKTCERWFQRFKSGDF (Sequence ID 9).
[0061] In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 9. In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains the amino acid sequence shown in SEQ ID NO: 9 having one, two, three, or four of the five conservative amino acid substitutions. In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains the amino acid sequence shown in SEQ ID NO: 9.
[0062] In some embodiments, a chimeric transposase containing cleaved piggyBac transposase residues with the amino acid sequence shown in SEQ ID NO: 3 is bound to two MosI DNA-binding domains with the sequence shown in SEQ ID NO: 6. The cleaved PB transposase domain may also be bound to the MosI DNA-binding domain via a linker sequence, resulting in a chimeric PB:MosI V552-111 transposase containing the amino acid sequence shown in SEQ ID NO: 10 (residues 3-11 of the Mos1 transposase are highlighted in bold, the PB domain is underlined, and the linker sequence is highlighted in italics): MAPKKKRKVGG GGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLDQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPV GGGGSSFVPNKEQTRTVLIFCFHLKKTAAESHRMLVEAFGEQVPTVKTCERWFQRFKSGDFDVDDKEHGKPPKRYEDAELQALLDEDDAQTQKQLAEQLEVSQQAVSNRLREMG (Sequence ID 10).
[0063] In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 10. In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains an amino acid sequence shown in SEQ ID NO: 10 having one, two, three, or four of the five conservative amino acid substitutions. In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains an amino acid sequence shown in SEQ ID NO: 10.
[0064] In some embodiments, a chimeric transposase containing a cleaved piggyBac transposase domain having the amino acid sequence of SEQ ID NO: 4 is linked to a single MosI DNA-binding domain having the amino acid sequence shown in SEQ ID NO: 8. The cleaved PB transposase domain may also be linked to the MosI DNA-binding domain via a linker sequence, resulting in a chimeric PB:MosI V552-58 transposase having the amino acid sequence shown in SEQ ID NO: 11 (residues 3-58 of the MosI transposase are highlighted in bold, the PB domain is underlined, and the linker sequence is highlighted in italics): MAPKKKRKVGG GGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLDQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPV GGGGSSFVPNKEQTRTVLIFCFHLKKTAAESHRMLVEAFGEQVPTVKTCERWFQRFKSGDF (Sequence ID 11).
[0065] In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 11. In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains an amino acid sequence shown in SEQ ID NO: 11 having one, two, three, or four of the five conservative amino acid substitutions. In some embodiments, the chimeric PB:MosI transpose containing the CRD deletion described herein contains an amino acid sequence shown in SEQ ID NO: 11.
[0066] Chimeric transposases may be used to transpose cells using transposons containing the chimeric transposon inverse repeat (ITR) polynucleotides described herein.
[0067] Chimeric transposon inverted terminal repeat (ITR) polynucleotides Also provided herein are chimeric transposon inverted terminal repeat (ITR) polynucleotides comprising, in the 5' to 3' direction: (i) a cleaved piggyBac left ITR sequence lacking a CRD binding site, and (ii) a MosI ITR sequence containing binding sites for two MosI DNA binding domains.
[0068] In some embodiments, the leftmost (LE) ITR sequence of piggyBac includes the sequence shown in SEQ ID NO: 12. In some embodiments, the 35 bp LE PB ITR sequence is cleaved after position 16, resulting in the deletion of a 19 bp CRD binding site. In some embodiments, the resulting cleaved PB LE ITR sequence lacking the CRD binding site contains only nucleotides 1-16 (highlighted in bold) of the full-length PB LE ITR sequence shown in SEQ ID NO: 12: CCCTAGAAAGATAGTCTGCGTAAAATTGACGCATG (SEQ ID NO: 12). In some embodiments, the PB LE ITR sequence lacking the CRD binding site includes the sequence CCCTAGAAAGATAGTC (SEQ ID NO: 172).
[0069] In some embodiments, the MosI ITR sequence containing binding sites for two MosI DNA-binding domains includes the following sequence: TGTACAAGTATGAAATGTCGTTT (Sequence ID 13).
[0070] In some embodiments, a 23 bp MosI sequence (SEQ ID NO: 13) containing binding sites for two MosI DNA binding domains is fused in reverse to a cleaved PB LE ITR (SEQ ID NO: 172), resulting in the LE PiggyBac:MosI ITR sequence shown in SEQ ID NO: 15:CCCTAGAAAGATAGTCAAACGACATTTCATACTTGTACA (SEQ ID NO: 15). In some embodiments, a single nucleotide is deleted from the 3' end of a 23 bp MosI sequence containing binding sites for two MosI DNA-binding domains (i.e., from the sequence shown in SEQ ID NO: 13), leaving only residues 1-22, i.e., the sequence shown in SEQ ID NO: 14 (TGTACAAGTATGAAATGTCGTT).
[0071] In some embodiments, a 1-22 nucleotide MosI ITR fragment is fused in reverse to a cleaved PB LE ITR sequence shown in SEQ ID NO: 12, resulting in the LE PiggyBac:MosI ITR-1 sequence shown in SEQ ID NO: 16 (CCCTAGAAAGATAGTCAACGACATTTCATACTTGTACA).
[0072] In some embodiments, the chimeric transposon ITR polynucleotide further comprises one or two additional nucleotides between the cleaved piggyBac left ITR sequence and the MosI ITR sequence. In some embodiments, one nucleotide is added to the 3' end of the MosI ITR sequence to create the LE PiggyBac:MosI ITR+1 sequence shown in SEQ ID NO: 17:CCCTAGAAAGATAGTCTAAACGACATTTCATACTTGTACA (SEQ ID NO: 17).
[0073] Similarly, two nucleotides can be added to the 3' end of a MosI ITR sequence to create the LE PiggyBac:MosI ITR+2 sequence shown in SEQ ID NO: 18:CCCTAGAAAGATAGTCTGAAACGACATTTCATACTTGTACA (SEQ ID NO: 18).
[0074] In some embodiments, the right-end (RE) ITR sequence of piggyBac includes the sequence shown in SEQ ID NO: 65. In some embodiments, the 63 bp RE PB ITR sequence is cleaved after position 16, resulting in the deletion of a 19 bp CRD binding site and the remaining ITR sequence. The cleaved PB RE ITR sequence lacking the CRD binding site contains only nucleotides 1-16 (highlighted in bold) of the full-length PB RE ITR shown in SEQ ID NO: 65: CCCTAGAAAGATAATCATATTGTGACGTACGTTAAAGATAATCATGCGTAAAATTGACGCATG (SEQ ID NO: 65). In some embodiments, the cleaved PB RE ITR sequence lacking the CRD binding site includes the sequence CCCTAGAAAGATAATC (SEQ ID NO: 173).
[0075] In some embodiments, the MosI ITR sequence containing binding sites for two MosI DNA-binding domains includes the following sequence: TGTACAAGTATGAAATGTCGTTT (SEQ ID NO: 13).
[0076] In some embodiments, a 23 bp MosI sequence (shown in SEQ ID NO: 13) containing binding sites for two MosI DNA binding domains is fused in reverse to a cleaved PB RE ITR (SEQ ID NO: 173) to create the RE PB:MosI ITR sequence shown in SEQ ID NO: 19: CCCTAGAAAGATAATCAAACGACATTTCATACTTGTACA (SEQ ID NO: 19). In some embodiments, a single nucleotide is deleted from the 3' end of a 23 bp MosI sequence (shown in SEQ ID NO: 13) containing binding sites for two MosI DNA-binding domains, resulting in the following sequence: TGTACAAGTATGAAATGTCGTT (SEQ ID NO: 14).
[0077] In some embodiments, a 1-22 nucleotide MosI ITR fragment (SEQ ID NO: 14) is fused in reverse to a cleaved PB RE ITR (SEQ ID NO: 173) to create RE PiggyBac:MosI ITR-1, shown in SEQ ID NO: 20 (CCCTAGAAAGATAATCAACGACATTTCATACTTGTACA).
[0078] In some embodiments, the polynucleotide further includes one or two additional nucleotides between the cleaved piggyBac right ITR sequence and the MosI ITR sequence. In some embodiments, one nucleotide is appended to the 3' end of the MosI ITR sequence to create the RE PB:MosI ITR+1 sequence shown in SEQ ID NO: 21:CCCTAGAAAGATAATCAAAACGACATTTCATACTTGTACA (SEQ ID NO: 21).
[0079] Similarly, two nucleotides can be added to the 3' end of a MosI ITR sequence to create the RE PiggyBac:MosI ITR+2 sequence shown in SEQ ID NO: 22: CCCTAGAAAGATAATCATAAACGACATTTCATACTTGTACA (SEQ ID NO: 22)
[0080] Chimeric transposon ITR polynucleotides may be incorporated into vectors or transposons (described below) for use with chimeric transposases, chimeric site-specific transposase fusion proteins, and in the methods described herein.
[0081] Chimeric site-specific PB:MosI transposase fusion protein Chimeric site-specific transposase fusion proteins are also provided herein, comprising, in the order of N-terminus to C-terminus: (i) a target-specific DNA-binding domain, (ii) a cleavage-type N-terminal deletion-integrated deletion piggyBac transposase including a C-terminal cysteine-rich domain (CRD) deletion, and (iii) one or more MosI DNA-binding domains. In some embodiments, the target-specific DNA-binding domain is a TAL or zinc finger motif (ZFM).
[0082] N-terminal Pb transposase deletion N-terminal deletion PB transposase sequences and embedded defective N-terminal piggyBac transposases have been previously described in International Patent Applications PCT / 2022 / 22549 and PCT / US2022 / 077549, respectively, each of which is incorporated herein by reference in whole for examples of embedded defective N-terminal deletion piggyBac transposases that may be used in the constructs and fusion proteins described herein. In some embodiments, the N-terminal deletion of the cleaved PB transposase is a deletion of amino acid residues 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103 of Sequence ID No. 1. In some embodiments, the N-terminal deletion of the cleavage-type PB transposase is the deletion of amino acid residues 1-93 of SEQ ID NO: 1, or the corresponding residue of SEQ ID NO: 3 or 4.
[0083] Embedded missing transposase In some embodiments, the chimeric site-specific transposase fusion protein includes a chimeric transposase that is an embedded-deficient. An embedded-deficient transposase domain is a transposase that can excise its corresponding transposon but incorporates the excise transposon at a lower frequency than the corresponding wild-type transposase. Examples of embedded-deficient transposases are disclosed in U.S. Patent No. 6,218,185; U.S. Patent No. 6,962,810; U.S. Patent No. 8,399,643 and International Patent Application Publication No. WO2019 / 173636, each of which is incorporated herein by reference in whole with respect to examples of embedded-deficient transposases that may be used in the constructs described herein. A list of embedded-deficient amino acid substitutions is disclosed in U.S. Patent No. 10,041,077, which is incorporated herein by reference in whole with respect to examples of embedded-deficient transposases that may be used in the constructs described herein. Wild-type SPB can be made unintegrated by introducing mutations, such as K93A, R372A, K375A, R376A, and / or D450N (for SEQ ID NOs. 1, 3, or 4). The introduction of mutations R372A, K375A, R376A, and D450N is thought to result in unintegrated transposases but retain resectable function.
[0084] Therefore, in some embodiments, the chimeric site-specific transposase fusion protein includes a chimeric transposase domain containing the sequence of Sequence ID No. 1 having one of the following mutations: R372A, K375A, R376A, and D450N. In some embodiments, the chimeric site-specific transposase fusion protein includes a chimeric transposase domain containing the sequence of Sequence ID No. 1 having two of the following mutations: R372A, K375A, R376A, and D450N. In some embodiments, the chimeric site-specific transposase fusion protein includes a chimeric transposase domain containing the sequence of Sequence ID No. 1 having three of the following mutations: R372A, K375A, R376A, and D450N. In some embodiments, the chimeric site-specific transposase fusion protein includes a chimeric transposase domain containing the sequence of Sequence ID No. 1 having one of the following mutations: R372A, K375A, R376A, and D450N.
[0085] Therefore, in some embodiments, the chimeric site-specific transposase fusion protein comprises a chimeric transposase domain containing the sequence of Sequence ID No. 3 having one of the following mutations: R372A, K375A, R376A, and D450N. In some embodiments, the chimeric site-specific transposase fusion protein comprises a chimeric transposase domain containing the sequence of Sequence ID No. 3 having two of the following mutations: R372A, K375A, R376A, and D450N. In some embodiments, the chimeric site-specific transposase fusion protein comprises a chimeric transposase domain containing the sequence of Sequence ID No. 3 having three of the following mutations: R372A, K375A, R376A, and D450N. In some embodiments, the chimeric site-specific transposase fusion protein comprises a chimeric transposase domain containing the sequence of Sequence ID No. 3 having one of the following mutations: R372A, K375A, R376A, and D450N.
[0086] Therefore, in some embodiments, the chimeric site-specific transposase fusion protein comprises a chimeric transposase domain containing the sequence of Sequence ID No. 4 having one of the following mutations: R372A, K375A, R376A, and D450N. In some embodiments, the chimeric site-specific transposase fusion protein comprises a chimeric transposase domain containing the sequence of Sequence ID No. 4 having two of the following mutations: R372A, K375A, R376A, and D450N. In some embodiments, the chimeric site-specific transposase fusion protein comprises a chimeric transposase domain containing the sequence of Sequence ID No. 4 having three of the following mutations: R372A, K375A, R376A, and D450N. In some embodiments, the chimeric site-specific transposase fusion protein comprises a chimeric transposase domain containing the sequence of Sequence ID No. 4 having one of the following mutations: R372A, K375A, R376A, and D450N.
[0087] An example sequence of a cleavage-type N-terminal deletion integration-deficient transposase, containing deletions of amino acids 1-93 at the N-terminus of the PB transposase sequence and including three integration-deficient mutations, is shown in SEQ ID NO: 85. NKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFL IRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNW FTSIPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (Sequence ID 85).
[0088] In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 85. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 85 having one, two, three, or four of the five conserved amino acid substitutions. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 85. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 89. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 89 having one, two, three, or four of the five conserved amino acid substitutions. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 89. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 97. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 97 having one, two, three, or four of the five conservative amino acid substitutions.In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes the amino acid sequence shown in SEQ ID NO: 97. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 110. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes the amino acid sequence shown in SEQ ID NO: 110 having one, two, three, or four of the five conservative amino acid substitutions. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes the amino acid sequence shown in SEQ ID NO: 110. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 118. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 118 having one, two, three, or four of the five conservative amino acid substitutions. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 118.In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 131. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 131 having one, two, three, or four of the five conservative amino acid substitutions. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 131. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 139. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 139 having one, two, three, or four of the five conserved amino acid substitutions. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 139.In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 152. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 152 having one, two, three, or four of the five conservative amino acid substitutions. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 152. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 160. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 160 having one, two, three, or four of the five conserved amino acid substitutions. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in SEQ ID NO: 160.
[0089] In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in any one of SEQ ID NOs: 87-170. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in any one of SEQ ID NOs: 87-170 having one, two, three, or four of the five conserved amino acid substitutions. In some embodiments, the cleavage-type N-terminal deletion-integration-deficient transposase described herein includes an amino acid sequence shown in any one of SEQ ID NOs: 87-170.
[0090] DNA targeting domain The chimeric site-specific transposase fusion proteins of this disclosure may further comprise one or more DNA targeting domains. The DNA targeting domains may be bound to the C-terminus or N-terminus of the chimeric site-specific transposase fusion protein. In preferred embodiments, the DNA targeting domains are bound to the N-terminus of the chimeric site-specific transposase fusion protein. While we do not wish to be bound by theory, it is thought that the addition of DNA targeting domains to chimeric site-specific transposase fusion proteins improves site-specific transposase activity by targeting the chimeric site-specific transposase fusion protein fused to the DNA targeting domains to the targeting sites. In some embodiments, insertion of DNA targeting domains improves site-specific transposase activity by at least 2-fold, at least 3-fold, at least 4-fold, or at least 5-fold compared to the same chimeric site-specific transposase without the DNA targeting domains.
[0091] Any DNA targeting domain known in the art, including but not limited to CRISPR, zinc finger motifs, TALE, and transcription factors, may be used in association with the chimeric transposases and chimeric site-specific transposase fusion proteins described herein.
[0092] Methods for manipulating zinc finger nucleases that bind to specific targets are described, for example, in Sander et al., Nat Methods. 2011 Jan;8(1):67-69. In some embodiments, the DNA targeting domain includes three zinc finger motifs. In some embodiments, the three zinc finger motifs are adjacent to a GGGGS (SEQ ID NO: 86) linker. In some embodiments, the three zinc finger motifs (ZFM) and the adjacent GGGGS (SEQ ID NO: 86) linker are the sequence shown in SEQ ID NO: 23: GGGGSERPYACPVESCDRRFSRSDELTRHIRIHTGQKPFQCRICMRNFSRSDHLTTHIRTHTGEKPFACDICGRKFARSDERKRHTKIHLRQKDGGGGS (SEQ ID NO: 23) Or it cumulatively includes sequences that have at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity with it.
[0093] In some embodiments, chimeric site-specific transposase fusion proteins comprising three ZFM domains bound to the chimeric PBx:MosI transposase of the present disclosure via a linker sequence are provided herein. An exemplary chimeric ZFM-PB:MosI transposase fusion protein comprises three ZFM domains bound to the chimeric S544:PB:MosI-58 transposase, creating a ZFM-S544:PB:MosI-58 chimeric site-specific transposase comprising the following nucleic acid sequence: (Sequence ID 24)
[0094] In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 24. In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence shown in SEQ ID NO: 24 having one, two, three, or four of the five conserved amino acid substitutions. In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence shown in SEQ ID NO: 24.
[0095] Another exemplary chimeric ZFM-PB:MosI transposase fusion protein contains three ZFM domains bound to the chimeric S544 PB:MosI-111 transposase, creating a ZFM-S544:PB:MosI-111 chimeric site-specific transposase with the following amino acid sequence: (Sequence ID 25)
[0096] In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 25. In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence shown in SEQ ID NO: 25 having one, two, three, or four of the five conserved amino acid substitutions. In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence shown in SEQ ID NO: 25.
[0097] Another exemplary chimeric ZFM-PB:MosI transposase fusion protein contains three ZFM domains bound to the chimeric V552 PB:MosI-58 transposase, creating a ZFM-PB:MosI-58 chimeric site-specific transposase containing the following nucleic acid sequence: (Sequence ID 26).
[0098] In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 26. In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence shown in SEQ ID NO: 26 having one, two, three, or four of the five conserved amino acid substitutions. In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence shown in SEQ ID NO: 26.
[0099] Another exemplary chimeric ZFM-PB:MosI transposase fusion protein contains three ZFM domains bound to the chimeric V552:PB:MosI-111 transposase, creating a ZFM-V552:PB:MosI-111 chimeric site-specific transposase containing the following nucleic acid sequence: (Base sequence 27).
[0100] In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence shown in SEQ ID NO: 27. In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence shown in SEQ ID NO: 27 having one, two, three, or four of the five conserved amino acid substitutions. In some embodiments, the chimeric ZFM-PB:MosI transposase fusion protein described herein contains an amino acid sequence shown in SEQ ID NO: 27.
[0101] In some embodiments, the DNA targeting domain is a TAL array. Xanthomonas-derived TALEs (transcriptional activator-like effectors) typically contain a 288-amino acid N-terminus, followed by an array of a variable number of repeat sequences (repeats) approximately 34 amino acids long, and then a 278-amino acid C-terminus (SEQ ID NO: 29); however, cleaved versions have been documented in the literature (see, e.g., Miller et al., Nat Biotechnol 29, 143-148 (2011)). TAL fused to FokI nucleases (TAL effector nucleases; TALENs) usually contain cleaved N-terminal and C-terminal forms. For example, the first 152 amino acids of the N-terminus may be removed (called delta-152; the sequence of the N-terminus after the removal of the first 152 amino acids is shown in SEQ ID NO: 30), or the C-terminus may be cleaved, leaving 63 amino acids (called +63; the sequence is shown in SEQ ID NO: 31).
[0102] TALs generally contain an array of 34 amino acids repeated a variable number of times. Two amino acids at positions 12 and 13 are typically altered to determine which nucleotide the TAL repeat recognizes. Thus, amino acids NG recognize T, NI recognize A, NN recognize G or A, HD recognize C, NK recognize G, and NS recognize A, C, G, or T. Other amino acids within the 34-residue repeat can also be altered. For example, position 11 is often changed to N for a repeat that recognizes G. Also, positions 4 and 32 are often altered to reduce the repeatability of the array, but not to determine binding specificity. The number of 34-amino acid repeats in the array determines the length of the DNA sequence to be recognized (one protein repeat binds to one DNA base). Furthermore, the last base is recognized by a "half-array" which is 20 amino acids instead of 34. This feature makes it possible to program TAL arrays to bind to specific DNA sequences. Those skilled in the art will be able to modify TALEN sequences to achieve desired target specificity.
[0103] Furthermore, the N-terminal domain of TAL (e.g., the N-terminal domain containing the sequence shown in SEQ ID NO: 30) recognizes and requires a T located immediately 5' of the target DNA sequence for binding. Mutations of the TAL N-terminal domain that no longer require a 5'T have been documented in the literature (see, e.g., Lamb et al., Nucleic Acids Res. 2013 Nov;41(21):9779-85). For example, the NT-G mutant containing the amino acid sequence shown in SEQ ID NO: 32 requires a 5'G instead of a 5'T, while the NT-βN mutant containing the amino acid sequence shown in SEQ ID NO: 33 does not require any specific 5' nucleotide. These mutant N-terminal domain sequences can be used to provide additional options for developing sequences of the fusion proteins described herein that can be targeted using TAL arrays.
[0104] In some embodiments, the TAL array provided herein comprises nine 34-amino acid repeats followed by a 20-amino acid "half" repeat having an adjacent BsmBI type IIS restriction site. In one embodiment, individual TAL modules containing either a 34-amino acid repeat or a 20-amino acid "half" repeat may be designed and synthesized adjacent to the BsmBI type IIS restriction site. In some embodiments, the entire TAL module set comprises four modules (i.e., 40 TAL modules for a recognized target at 10 bp) capable of recognizing any of A, C, G, or T at each 10 bp position, and one TAL half-repeat module. Exemplary TAL modules are shown in SEQ ID NOs: 34-37, where X is any amino acid: TAL module version 1: LTPDQVVAIAXXXGGKQALETVQRLLPVLCQDHG (Sequence ID 34) TAL module version 2: LTPEQVVAIAXXXGGKQALETVQRLLPVLCQAHG (Sequence ID 35) TAL module version 3: LTPDQVVAIAXXXGGKQALETVQRLLPVLCQAHG (Sequence ID 36) TAL module version 4: LTPAQVVAIAXXXGGKQALETVQRLLPVLCQDHG (Sequence ID 37).
[0105] An exemplary TAL semimodule is shown in SEQ ID NO: 38, where X is any amino acid: LTPEQVVAIAXXXGGRPALE(SEQ ID NO: 38).
[0106] Pairs of TAL arrays targeting sequences within desired genes can be designed, corresponding modules can be selected, and they can be pooled together using "Golden Gate Assembly," with each TAL array being assembled in-frame. The DNA sequences encoding the TAL arrays generated herein can be further codon-optimized using the GeneArt algorithm (Thermo Fisher).
[0107] TAL arrays can be used to target TTAA sites within the genome. When designing left and right TAL arrays, each containing an N-terminal domain that recognizes T and a TAL C-terminal domain that fuses to a chimeric site-specific transposase, one TAL array recognizes sequence 5' of TTAA, and the other TAL array recognizes sequence 3' of TTAA. Since sequence 5' of TTAA most often differs from sequence 3' of TTAA in genomic DNA targets, TAL site-specific Super piggyBac transposase fusion proteins ("TAL-ssSPB" or "TAL-PBx") are most often used as heterodimers consisting of two transposase fusion proteins, each containing a different TAL domain that recognizes two different DNA sequences. Furthermore, the sequence recognized by the TAL array is not necessarily directly adjacent to TTAA. Instead, it can be separated from TTAA by a spacer of a given bp length, e.g., a 12bp, 13bp, or 14bp spacer.
[0108] TAL arrays can target any desired DNA sequence (e.g., genomic DNA sequence). It will be apparent to those skilled in the art that any left TAL array of a given target can be combined with any right TAL array of the same target.
[0109] In some embodiments, the TAL array targets green fluorescent protein (GFP). A TAL-piggyBac transposase fusion protein comprising a GFP-targeting N-terminal deletion piggyBac transposase sequence and an embedded defect N-terminal piggyBac transposase is described in the jointly owned international patent application PCT / 2022 / 22549, which is incorporated herein in whole for an example of a TAL-piggyBac transposase fusion protein that may be used in the constructs described herein.
[0110] Exemplary sequences of the right TAL array, which can be incorporated into chimeric site-specific transposase fusion proteins targeting GFP, are shown in SEQ ID NOs: 41-44.
[0111] In some embodiments, the TAL array targets the LINE1 repeat element. A TAL-piggyBac transposase fusion protein comprising an N-terminal deletion piggyBac transposase sequence targeting the LINE1 element and an embedded defect N-terminal piggyBac transposase is described in International Patent Application PCT / 2022 / 22549, and the entire example of a TAL-piggyBac transposase fusion protein that may be used in the constructs described herein is incorporated herein.
[0112] Exemplary sequences of the left and right TAL arrays targeting LINE1 are shown in SEQ ID NOs. 66 and 67, respectively.
[0113] Nuclear localization signals In some embodiments, the chimeric transposases and chimeric site-specific transposase fusion proteins provided herein may include an in-frame nuclear localization sequence (NLS). Examples of transposases fused to nuclear localization signals are disclosed in U.S. Patent No. 6,218,185; U.S. Patent No. 6,962,810; U.S. Patent No. 8,399,643 and International Patent Application Publication No. WO2019 / 173636, each of which is incorporated herein in whole with respect to examples of transposases fused to an NLS that may be used in the constructs described herein. In some embodiments, the NLS includes the sequence PKKKRKV (SEQ ID NO: 39). In certain embodiments, the in-frame NLS is located upstream (N-terminus) of the transposase domain, which includes an N-terminal deletion. In certain embodiments, the in-frame NLS is located downstream (C-terminus) of the transposase domain, which includes an N-terminal deletion.
[0114] Generally, the NLS is preferably located at the N-terminus of a chimeric transposase or a chimeric site-specific transposase fusion protein.
[0115] In certain embodiments, the in-frame NLS is directly fused to the amino terminus of a chimeric transposase or chimeric site-specific transposase fusion protein. In some embodiments, an initiating methionine is introduced before the NLS. In some embodiments, additional adenine residues are introduced before and / or after the NLS to ensure in-frame translation. Thus, the residue numbering in SEQ ID NO: 5 begins with the 12th residue of SEQ ID NO: 5 for the purpose of identifying deleted and mutated residues.
[0116] nucleic acid Polynucleotides comprising nucleic acid sequences encoding chimeric transposases and chimeric site-specific transposase fusion proteins described herein are also provided herein. In some embodiments, the polynucleotides are isolated or purified.
[0117] The isolated polynucleotides of this disclosure can be prepared using (a) recombinant methods, (b) synthetic techniques, (c) purification techniques, and / or (d) combinations thereof, which are well known in the art. Methods for constructing nucleic acids encoding chimeric transposases and chimeric site-specific transposase fusion proteins described herein are well known in the art or described herein, for example, PCR-based mutagenesis.
[0118] The fusion products of the present invention can be produced using any suitable method known in the art or described herein.
[0119] The isolated polynucleotides of this disclosure, such as RNA, cDNA, genomic DNA, or any combination thereof, can be obtained from biological sources using any number of cloning methodologies known to those skilled in the art. In some embodiments, desired sequences in a cDNA or genomic DNA library are identified using oligonucleotide probes that selectively hybridize to the polynucleotides of this disclosure under stringent conditions.
[0120] Methods for amplifying RNA or DNA are well known in the art and can be used in accordance with this disclosure without excessive experimentation, based on the teachings and guidance presented herein. Known methods for amplifying DNA or RNA include polymerase chain reaction (PCR) and related amplification processes (e.g., Mullis et al., U.S. Patents No. 4,683,195, 4,683,202, 4,800,159, and 4,965,188; Tabor et al., U.S. Patents No. 4,795,699 and 4,921,794; Innis, U.S. Patent No. 5,142,033; Wilson et al., U.S. Patent No. 5,122,464; Innis, U.S. Patent No. 5,091,310; Gyllensten et al.) This includes, but is not limited to, U.S. Patent No. 5,066,584; No. 4,889,818 by Gelfand et al.; No. 4,994,370 by Silver et al.; No. 4,766,067 by Biswas; and No. 4,656,134 by Ringold, as well as RNA-mediated amplification using antisense RNA against a target sequence as a template for double-stranded DNA synthesis (U.S. Patent No. 5,130,238 by Malek et al., trade name NASBA), the contents of which are incorporated herein by reference.
[0121] For example, polymerase chain reaction (PCR) technology can be used to amplify the sequences of the polynucleotides and related genes of this disclosure directly from a genomic DNA or cDNA library. PCR and other in vitro amplification methods may also be useful for other purposes, for example, cloning nucleic acid sequences encoding proteins to be expressed, preparing nucleic acids to be used as probes to detect the presence of desired mRNA in a sample, for nucleic acid sequencing, or for other purposes. Examples of techniques sufficient to direct those skilled in the art through in vitro amplification methods can be found in Berger (ibid.), Sambrook (ibid.), and Ausubel (ibid.), as well as Mullis et al., U.S. Patent No. 4,683,202 (1987); and Innis, et al., PCR Protocols: A Guide to Methods and Applications, Eds., Academic Press Inc., San Diego, Calif. (1990). Commercial kits for genomic PCR amplification are known in the art. See, for example, the Advantage-GC Genomic PCR Kit (Clontech). Furthermore, for example, the T4 gene 32 protein (Boehringer Mannheim) can be used to improve the yield of long PCR products.
[0122] The polynucleotides of this disclosure can also be prepared by direct chemical synthesis using known methods (see, for example, Ausubel et al. (cited above)). Chemical synthesis generally produces single-stranded oligonucleotides, which can be converted to double-stranded DNA by hybridization with complementary sequences or polymerization with DNA polymerase using a single strand as a template. While the chemical synthesis of DNA may be limited to sequences of about 100 bases or more, those skilled in the art will recognize that longer sequences can be obtained by ligating shorter sequences.
[0123] Expression vectors and host cells This disclosure also relates to vectors containing the polynucleotides of this disclosure, host cells genetically engineered with recombinant vectors, and the production of at least one protein backbone by recombinant technology, as is well known in the art. See, for example, Sambrook et al. and Ausubel et al. (each incorporated herein by reference in their entirety).
[0124] Polynucleotides can optionally be linked to vectors containing selectable markers for replication in a host. Generally, plasmid vectors are introduced into precipitates such as calcium phosphate precipitates or into complexes with charged lipids. If the vector is a virus, it can be packaged in vitro using a suitable packaging cell line and then transduced into host cells.
[0125] In some embodiments, the DNA insert is operably linked to a suitable promoter. In some embodiments, the promoter is the EF-1α promoter containing the sequence shown in SEQ ID NO: 45. The expression construct may further contain sites for transcription start and termination, and a ribosome binding site for translation within the transcription region. The coding portion of the mature transcript expressed by the construct preferably includes a translation start codon (e.g., ATG) at the start and a appropriately placed stop codon (e.g., UAA, UGA, or UAG) at the end of the mRNA being translated, with UAA and UAG being preferred for expression in mammalian or eukaryotic cells.
[0126] The expression vector may further contain one or more selection markers. Examples of selection markers include, but are not limited to, ampicillin, zeosin (Sh bla gene), puromycin (pac gene), hygromycin B (hygB gene), G418 / genethecin (neo gene), DHFR (encoding dihydrofolate reductase and conferring resistance to methotrexate), mycophenolic acid or glutamine synthetase (GS, U.S. Patent Nos. 5,122,464 and 5,770,359; and 5,827,739), blastocydin (bsd gene), resistance genes to eukaryotic cell culture, and ampicillin, zeosin (Sh Examples include the bla gene, puromycin (pac gene), hygromycin B (hygB gene), G418 / genethecin (neo gene), kanamycin, spectinomycin, streptomycin, carbenicillin, bleomycin, erythromycin, polymyxin B, or tetracycline resistance genes for culture in Escherichia coli and other bacteria or prokaryotes. Suitable culture media and conditions for the above host cells are known in the art. Suitable vectors will be readily apparent to those skilled in the art. Introduction of vector constructs into host cells can be carried out by calcium phosphate transfection, DEAE-dextran-mediated transfection, cationic lipid-mediated transfection, electroporation, transduction, infection, or other known methods. Such methods are described in the art, for example, in Sambrook (next op.), Chapters 1-4 and 16-18; and Ausubel (next op.), Chapters 1, 9, 13, 15, and 16.
[0127] An expression vector may include at least one selectable cell surface marker for isolating cells modified by the compositions and methods of the present disclosure. The selectable cell surface markers of the present disclosure include surface proteins, glycoproteins, or groups of proteins that distinguish a cell or cell subset from another defined cell subset. In some embodiments, the selectable cell surface marker distinguishes cells modified by the compositions or methods of the present disclosure from cells not modified by the compositions or methods of the present disclosure. Examples of cell surface markers, but not limited to, include “designated cluster” or “classification determinant” proteins (often abbreviated as “CD”), e.g., cleaved or full-length forms of CD19, CD271, CD34, CD22, CD20, CD33, CD52, or any combination thereof. Cell surface markers include the suicide gene marker RQR8 (Philip B et al. Blood. 2014 Aug 21;124(8):1277-87).
[0128] The expression vector may further comprise at least one selectable drug resistance marker for isolating cells modified by the compositions and methods of the present disclosure. The selectable drug resistance markers of the present disclosure may comprise wild-type or mutant Neo, DHFR, TYMS, FRANCF, RAD51C, GCS, MDR1, ALDH1, NKX2.2, or any combination thereof.
[0129] Those skilled in the art will be familiar with the numerous expression systems available for expressing the nucleic acids encoding the proteins of the Disclosure. The nucleic acids of the Disclosure can be expressed in host cells by being (operationally) turned on in host cells containing endogenous DNA encoding the protein backbone of the Disclosure. Such methods are well known in the art, as described, for example, in U.S. Patents 5,580,734, 5,641,670, 5,733,746, and 5,733,761, each of which is incorporated herein by reference in whole.
[0130] Cell cultures useful for producing protein backbones, specific parts thereof, or variants are known in the art to be bacterial, yeast, and mammalian cells. Mammalian cell lines often exist in the form of a single layer of cells, but mammalian cell suspensions or bioreactors can also be used. Several suitable host cell lines capable of expressing intact glycosylated proteins have been developed in the art, including COS-1 (e.g., ATCC CRL 1650), COS-7 (e.g., ATCC CRL-1651), HEK293, BHK21 (e.g., ATCC CRL-10), CHO (e.g., ATCC CRL 1610), and BSC-1 (e.g., ATCC CRL-26) cell lines, Cos-7 cells, CHO cells, hepG 2 cells, P3X63Ag8.653, SP2 / 0-Ag14, 293 cells, and HeLa cells, which are readily available, for example, from the American Type Culture Collection, Manassas, Va (www.atcc.org). Preferred host cells include lymphoid cells such as myeloma cells and lymphoma cells. Particularly preferred host cells are P3X63Ag8.653 cells (ATCC accession number CRL-1580) and SP2 / 0-Ag14 cells (ATCC accession number CRL-1851). In a preferred embodiment, the recombinant cells are P3X63Ab8.653 or SP2 / 0-Ag14 cells.
[0131] These cell expression vectors may contain one or more of the following expression regulatory sequences, e.g., origins of replication, promoters (e.g., late or early SV40 promoter, CMV promoter (US Patent No. 5,168,062, 5,385,839), HSV tk promoter, pgk (phosphoglycerin kinase) promoter, EF-1 alpha promoter (US Patent No. 5,266,491), at least one human promoter, enhancer, and / or processing information sites, e.g., ribosome binding sites, RNA splice sites, polyadenylation sites (e.g., SV40 large T Ag poly-A addition site), and transcription terminator sequences. See, for example, Ausubel et al. and Sambrook et al. Other cells useful for generating the nucleic acids or proteins of this disclosure are known and available, for example, from the American Type Culture Collection Catalogue of Cell Lines and Hybridoma (www.atcc.org) or other known or commercial sources.
[0132] When eukaryotic host cells are used, polyadenylated or transcriptional terminator sequences are typically incorporated into the vector. An example of a terminator sequence is a polyadenylated sequence derived from the bovine growth hormone gene. In some embodiments, the polyA sequence is the SV40 polyA sequence, which includes the sequence shown in SEQ ID NO: 47.
[0133] Sequences for precise splicing of the transcript can also be included. An example of a splicing sequence is the VP1 intron derived from SV40 (Sprague, et al., J. Virol. 45:773-781 (1983)). Furthermore, as is known in the art, gene sequences for controlling replication in host cells can be incorporated into the vector.
[0134] Plasmid constructs described herein may be used to deliver nucleic acids encoding transposase domains or fusion proteins described herein to cells.
[0135] The transposase domains and fusion proteins described herein may be delivered to cells using mRNA constructs. Accordingly, in one embodiment, mRNA sequences encoding the transposase domains or fusion proteins described herein are provided herein. Such mRNA sequences may be delivered to cells using nanoparticles, such as lipid nanoparticles. Examples of lipid nanoparticles are described, for example, in International Patent Application Publication Nos. WO2022 / 087148, WO2022 / 182792, WO2023 / 141576, WO2024 / 035783, and WO2023 / 141576, each of which is incorporated herein by reference in whole with respect to examples of lipid nanoparticles that may be used to deliver the fusion proteins or transposase domains described herein. The mRNA constructs may also be delivered to cells by electroporation or nucleofection. The mRNA may be capped or otherwise modified.
[0136] Transposons, cells, and modified cells The chimeric transposases and chimeric site-specific transposase fusion proteins described herein may be used in conjunction with transposons to modify cells. The transposon may be a piggyBac®(PB) transposon. In some embodiments where the transposon is a PB transposon, the transposase is a piggyBac®(PB) transposon, a piggyBac-like(PBL) transposon, or a Super piggyBac®(SPB) transposon. Non-limiting examples of PB transposons are described in detail in U.S. Patents 6,218,182; 6,962,810; 8,399,643 and PCT Publication No. WO2010 / 099296, each of which is incorporated herein by reference in whole, with respect to examples of transposons that may be used in the methods disclosed herein. Transposons may include nucleic acids encoding therapeutic proteins or therapeutic agents. Examples of therapeutic proteins are disclosed in PCT publication numbers WO2019 / 173636 and WO2020 / 051374, each of which is incorporated herein by reference in its entirety as an example of therapeutic proteins that may be used in the chimeric transposases and chimeric site-specific transposase fusion proteins described herein.
[0137] In some embodiments, the transposon may be a transposon comprising one or more chimeric PB:MosI ITR polynucleotides as described herein. In some embodiments, the transposon comprises a chimeric LE PB:Mos1 ITR polynucleotide and a chimeric RE PB:MosI ITR polynucleotide. In some embodiments, the LE PB:MosI ITR polynucleotide is a PB:MosI ITR comprising the sequence shown in SEQ ID NO: 19. In some embodiments, the LE PB:MosI ITR polynucleotide is a PB:MosI ITR-1 comprising the sequence shown in SEQ ID NO: 20. In some embodiments, the LE PB:MosI ITR polynucleotide is a PB:MosI ITR+1 comprising the sequence of SEQ ID NO: 21. In some embodiments, the LE PB:MosI ITR polynucleotide is a PB:MosI ITR+2 comprising the sequence of SEQ ID NO: 22.
[0138] In some embodiments, the RE PB:MosI ITR polynucleotide is a PB:MosI ITR containing the sequence of SEQ ID NO: 15. In some embodiments, the RE PB:MosI ITR polynucleotide is a PB:MosI ITR-1 containing the sequence of SEQ ID NO: 16. In some embodiments, the RE PB:MosI ITR polynucleotide is a PB:MosI ITR+1 containing the sequence of SEQ ID NO: 17. In some embodiments, the RE PB:MosI ITR polynucleotide is a PB:MosI ITR+2 containing the sequence of SEQ ID NO: 18.
[0139] In some embodiments, the PB:MosI ITR transposon contains the following nucleotide sequence: TTAACCCTAGAAAGATAGTCAAACGACATTTCATACTTGTACAGACGCATGCATTCTTGAAATATTGCTCTCTCTTTCTAAATAGCGCGAATCCGTCGCTGTGCATTTAGGACATCTCAGTCGCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAG TACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAAATGTACAAGTATGAAATGTCGTTTGATTATCTTTCTAGGGTTAA (SEQ ID NO: 50). In some embodiments, the PB:MosI ITR transposons described herein contain a nucleic acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleic acid sequence shown in SEQ ID NO: 50.
[0140] In some embodiments, the PB:MosI ITR-1 transposon includes the following nucleotide sequence: TTAACCCTAGAAAGATAGTCAACGACATTTCATACTTGTACAGACGCATGCATTCTTGAAATATTGCTCTCTCTTTCTAAATAGCGCGAATCCGTCGCTGTGCATTTAGGACATCTCAGTCGCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAG TACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAAATGTACAAGTATGAAATGTCGTTGATTATCTTTCTAGGGTTAA (SEQ ID NO: 51). In some embodiments, the PB:MosI ITR transposons described herein contain a nucleic acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleic acid sequence shown in SEQ ID NO: 51.
[0141] In some embodiments, the PB:MosI ITR+1 transposon contains the following nucleotide sequence: TTAACCCTAGAAAGATAGTCTAAACGACATTTCATACTTGTACAGACGCATGCATTCTTGAAATATTGCTCTCTCTTTCTAAATAGCGCGAATCCGTCGCTGTGCATTTAGGACATCTCAGTCGCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAG TACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAAATGTACAAGTATGAAATGTCGTTTTGATTATCTTTCTAGGGTTAA (SEQ ID NO: 52). In some embodiments, the PB:MosI ITR transposons described herein contain a nucleic acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleic acid sequence shown in SEQ ID NO: 52.
[0142] In some embodiments, the PB:MosI ITR+2 transposon contains the following nucleotide sequence: TTAACCCTAGAAAGATAGTCTGAAACGACATTTCATACTTGTACAGACGCATGCATTCTTGAAATATTGCTCTCTCTTTCTAAATAGCGCGAATCCGTCGCTGTGCATTTAGGACATCTCAGTCGCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAG TACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAAATGTACAAGTATGAAATGTCGTTTATGATTATCTTTCTAGGGTTAA (SEQ ID NO: 53). In some embodiments, the PB:MosI ITR transposons described herein contain a nucleic acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleic acid sequence shown in SEQ ID NO: 53.
[0143] In some embodiments, the transposons disclosed herein include a truncated PB:MosI ITR polynucleotide. In some embodiments, the leftmost PB:MosI truncated ITR includes the nucleic acid sequence shown in SEQ ID NO: 54. In some embodiments, the rightmost PB:MosI truncated ITR includes the nucleic acid sequence shown in SEQ ID NO: 55. In some embodiments, the transposon includes a leftmost PB:MosI truncated ITR and a rightmost PB:MosI truncated ITR to produce a PB:MosI truncated transposon containing the nucleic acid sequence shown in SEQ ID NO: 56. In some embodiments, the PB:MosI truncated transposon contains a sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleic acid sequence shown in SEQ ID NO: 56.
[0144] In another embodiment, modified cells comprising one or more transposons and one or more chimeric transposases or chimeric site-specific transposase fusion proteins described herein are provided herein. The cells and modified cells of this disclosure may be mammalian cells. Preferably, the cells and modified cells are human cells.
[0145] Cells modified using the chimeric transposases or chimeric site-specific transposase fusion proteins described herein may be germ cells or somatic cells. The cells and modified cells of this disclosure include immune cells, such as lymphoid progenitor cells, natural killer (NK) cells, T lymphocytes (T cells), and stem memory T cells (T SCM T cells), central memory T cells (T CMModified cells may be stem cells, T cells, B lymphocytes (B cells), antigen-presenting cells (APCs), cytokine-induced killer (CIK) cells, myeloid progenitor cells, neutrophils, basophils, eosinophils, monocytes, macrophages, platelets, erythrocytes, red blood cells (RBCs), megakaryocytes, or osteoclasts. Modified cells may be differentiated, undifferentiated, or immortalized. Modified undifferentiated cells may be stem cells. Modified undifferentiated cells may be induced pluripotent stem cells. Modified cells may be hepatocytes, T cells, hematopoietic stem cells, natural killer cells, macrophages, dendritic cells, monocytes, megakaryocytes, or osteoclasts. Modified cells can be modified while the cell is in the quiescent, activated, resting, interphase, prophase, metaphase, anaphase, or telophase. Modified cells may be fresh cells, cryopreserved cells, bulk cells, cells classified into subpopulations, cells derived from whole blood, cells derived from leukocyte apheresis, or cells derived from immortalized cell lines. Detailed descriptions for isolating cells from leukocyte apheresis products or blood are disclosed in PCT publications WO2019 / 173636 and WO2020 / 051374, each of which is incorporated herein by reference in whole.
[0146] The method of the present disclosure can be used to modify or produce a population of modified T cells in which at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% or any percentage between them are stem memory T cells (T SCM ) or T SCM The cells express one or more cell surface markers; these one or more cell surface markers include CD45RA and CD62L. SCM or T SCMCell surface markers of similar cells may also include one or more of CD62L, CD45RA, CD28, CCR7, CD127, CD45RO, CD95, CD95, and IL-2Rβ. SCM or T SCM Cell surface markers for similar cells may also include one or more of the following: CD45RA, CD95, IL-2Rβ, CCR7, and CD62L.
[0147] This disclosure provides a method for expressing a CAR on the surface of a cell. The method comprises (a) obtaining a cell population, (b) contacting the cell population with a composition comprising a CAR or a CAR-coding sequence under conditions sufficient to transfer the CAR or CAR-coding sequence across the cell membrane of at least one cell in the cell population, thereby generating a modified cell population, (c) culturing the modified cell population under conditions suitable for the incorporation of the CAR-coding sequence, and (d) growing and / or selecting at least one cell from the modified cell population that expresses a CAR on its cell surface. A more detailed description of the method for expressing a CAR on the surface of a cell is disclosed in PCT publication numbers WO2019 / 049816 and WO2020 / 051374.
[0148] The Disclosure further provides cells or a population of cells comprising a composition comprising (a) an inducible transgene construct comprising a sequence encoding an inducible promoter and a sequence encoding a transgene, and (b) a receptor construct comprising a sequence encoding a constitutive promoter and a sequence encoding an exogenous receptor such as a CAR, wherein when constructs (a) and (b) are incorporated into the genomic sequence of cells, the exogenous receptor is expressed, and when the exogenous receptor binds to a ligand or antigen, it transmits an intracellular signal that directly or indirectly targets an inducible promoter that modulates the expression of the inducible transgene (a), thereby modifying gene expression.
[0149] This disclosure further provides compositions comprising modified, amplified, and selected cell populations as described herein.
[0150] The modified cells of this disclosure (e.g., CAR T cells) may be further modified to enhance their therapeutic potential. Alternatively, or in addition to this, the modified cells may be further modified to reduce their sensitivity to immunological and / or metabolic checkpoints, for example, by blocking and / or diluting specific checkpoint signals that are naturally delivered to the cell within the tumor immunosuppressive microenvironment (e.g., checkpoint inhibition).
[0151] The modified cells of this disclosure (e.g., CAR T cells) may be further modified to silence or reduce the expression of (i) one or more genes encoding receptors for inhibitory checkpoint signals; (ii) one or more genes encoding intracellular proteins involved in checkpoint signaling; (iii) one or more genes encoding transcription factors that interfere with therapeutic efficacy; (iv) one or more genes encoding receptors for cell death or apoptosis; (v) one or more genes encoding metabolic sensing proteins; (vi) one or more genes encoding proteins that confer sensitivity to cancer treatment, including monoclonal antibodies; and / or (vii) one or more genes encoding growth advantage factors. The modified cells may also be modified to silence or reduce the expression of endogenous T cell receptors. Non-exclusive examples of genes that can be modified to silence or reduce their expression or suppress their function include, but are not limited to, inhibitory checkpoint signals, intracellular proteins, transcription factors, cell death or apoptosis receptors, metabolic sensing proteins, proteins that confer sensitivity to cancer treatment, and genes encoding growth-advantage factors, the entirety of which is disclosed in PCT publication number WO2019 / 173636, which is incorporated herein by reference, for examples of genes that can be silenced in the cells of this disclosure.
[0152] The modified cells of this disclosure (e.g., CAR T cells) may be further modified to express modified / chimeric checkpoint receptors. Modified / chimeric checkpoint receptors may include null receptors, decoy receptors, or dominant-negative receptors. Exemplary null, decoy, or dominant-negative intracellular receptors / proteins include, but are not limited to, downstream signaling components of inhibitory checkpoint signals, transcription factors, cytokines or cytokine receptors, chemokines or chemokine receptors, cell death or apoptosis receptors / ligands, metabolic sensing molecules, proteins that confer sensitivity to cancer treatment, and oncogenes or tumor suppressor genes. Non-exclusive examples of cytokines, cytokine receptors, chemokines, and chemokine receptors are disclosed in PCT publication number WO2019 / 173636, which is incorporated herein by reference in its entirety with respect to examples of modified / chimeric checkpoint receptors that may be expressed in the cells described herein.
[0153] Genome modification involves introducing nucleic acid sequences, transgenes, and / or genome editing constructs into cells ex vivo, in vivo, in vitro, or in situ to stably incorporate nucleic acid sequences, transiently incorporate nucleic acid sequences, cause site-directed integration of nucleic acid sequences, or cause biased integration of nucleic acid sequences. A nucleic acid sequence can be a transgene.
[0154] Stable chromosome integration can be random, site-specific, or biased. While we do not wish to be bound by theory, the addition of DNA-binding domains to the chimeric transposases and chimeric site-specific transposase fusion proteins described herein is thought to improve the site specificity of the transposases.
[0155] Site-specific integration can occur at safe harbor sites. Genomic safe harbor sites can provide a place for the integration of new genetic material in a way that ensures the newly inserted gene element functions reliably (e.g., is expressed at therapeutically effective levels) and does not cause harmful changes in the host genome that pose a risk to the host organism. Non-limiting examples of potential genomic safe harbors include the intron sequence of the human albumin gene, adeno-associated virus site 1 (AAVS1), the naturally occurring integration site of the AAV virus on chromosome 19, the site of the chemokine (CC motif) receptor 5 (CCR5) gene, and the site of the human ortholog at the mouse Rosa26 locus.
[0156] Site-directed transgene integration can occur at sites that disrupt the expression of a target gene. This disruption can occur through site-directed integration at introns, exons, promoters, gene elements, enhancers, suppressors, start codons, stop codons, and response elements. Non-exclusive examples of target genes targeted by site-directed integration include TRAC, TRAB, PD1, any gene encoding an immunosuppressive protein, a gene encoding an endogenous T cell receptor, and a gene encoding a protein involved in allorejection.
[0157] Site-directed transgene integration can occur at sites that result in enhanced expression of the target gene. Enhanced target gene expression can occur through site-directed integration at introns, exons, promoters, gene elements, enhancers, suppressors, start codons, stop codons, and response elements.
[0158] Site-specific transgene integration sites can be unstable chromosomal insertions. Unstable integration can be transient non-chromosomal integration, semi-stable non-chromosomal integration, semi-persistent non-chromosomal insertion, or unstable chromosomal insertion. Transient non-chromosomal insertions can be epichromosomes or cytoplasm. In one embodiment, transient non-chromosomal insertions of transgenes are not integrated into chromosomes, and the modified genetic material is not replicated during cell division.
[0159] Site-specific transgene integration sites may be modified binding sites for DNA targeting domains in the chimeric transposase or chimeric site-specific transposase fusion protein described herein. For example, the TTAA target DNA integration site for SPB may be modified to insert an adjacent DNA binding site for a DNA targeting domain containing three zinc finger motifs (e.g., a DNA targeting domain containing or consisting of the sequence of SEQ ID NO: 23, or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity thereto). For example, a DNA targeting domain containing three zinc finger motifs is thought to bind to the DNA sequence GCGTGGGCG. Therefore, the introduction of two copies of the DNA sequence GCGTGGGCG adjacent to the TTAA target integration site for a chimeric site-specific transposase fusion protein is thought to improve site-specific integration of the chimeric transposase containing the DNA targeting domain containing three zinc finger motifs. In such embodiments, the two copies of the DNA sequence GCGTGGGCG are oriented in reverse (5') and complementary (3') orientations.
[0160] In some embodiments, polynucleotides are provided herein that include the reverse complement of the sequence of a target site for a DNA targeting domain, a first spacer, a TTAA target integration site for the SPB, a second spacer, and the sequence of the target site for the DNA targeting domain in 5' to 3' order. In some embodiments, the first and second spacers are of the same length. In some embodiments, the first and / or second spacers are 3 bp long. In some embodiments, the first and / or second spacers are 4 bp long. In some embodiments, the first and / or second spacers are 5 bp long. In some embodiments, the first and / or second spacers are 6 bp long. In some embodiments, the first and / or second spacers are 7 bp long. In some embodiments, the first and / or second spacers are 8 bp long. In some embodiments, the first and / or second spacers are 9 bp long. In some embodiments, the first and / or second spacers are 10 bp long.
[0161] The modified target site can be introduced into a cell or cell line to facilitate targeted genomic manipulation. For example, a cell line engineered to include a modified target site for a chimeric transposase or chimeric site-specific transposase fusion protein provided herein can be transfected with the chimeric transposase or chimeric site-specific transposase fusion protein, as well as a transposon containing donor DNA, so that donor DNA is inserted at the modified target site. In some embodiments, the cell line is a T cell line. In some embodiments, the ZFM-PB:MosI fusion protein includes one amino acid sequence of SEQ ID NOs. 24-27.
[0162] Genome modification can result in the unstable chromosomal integration of introduced genes. Integrated introduced genes may be silenced, removed, excised, or further modified.
[0163] In some embodiments, the chimeric transposases and chimeric site-specific transposase fusion proteins provided herein have better transposase efficacy than their wild-type counterparts. Transposase activity can be measured by any suitable assay known in the Art or described herein, such as the Split GFP assay. For example, the chimeric transposases and chimeric site-specific transposase fusion proteins provided herein may have on-target genomic integration activity equivalent to their wild-type counterparts, but may have reduced off-target genomic integration activity compared to their wild-type counterparts.
[0164] In some embodiments, chimeric PB:MosI transposases containing the DNA targeting domain provided herein have an on-target activity to off-target activity ratio that is at least 50-fold, at least about 100-fold, at least about 150-fold, at least about 200-fold, at least about 250-fold, at least about 300-fold, at least about 350-fold, at least about 400-fold, at least about 450-fold, at least about 500-fold, at least about 550-fold, at least about 600-fold, at least about 650-fold, at least about 700-fold, at least about 750-fold, at least about 800-fold, at least about 850-fold, at least about 900-fold, at least about 950-fold, or at least about 1000-fold increased compared to a wild-type transposase domain (e.g., PB or SPB).
[0165] In certain embodiments, the modified cells are used therapeutically in adoptive cell therapy.
[0166] Adoptive cell compositions that are “universally” safe for administration to any patient (not just the patient from which they originate) require a significant reduction or elimination of alloreactivity. For this purpose, the cells of the Disclosure (e.g., allogeneic cells) can be modified to disrupt the expression or function of classes of T cell receptors (TCRs) and / or major histocompatibility complexes (MHCs). TCRs mediate graft-versus-host (GvH) responses, and MHCs mediate host-versus-graft (HvG) responses. In a preferred embodiment, any expression and / or function of TCRs is eliminated in the modified cells disclosed herein to prevent T cell-mediated GvH that could cause death in the subject. Thus, in a preferred embodiment, the Disclosure provides a pure TCR-negative allogeneic T cell composition (e.g., each cell in the composition expresses endogenous TCRs at such low levels that they are undetectable or absent).
[0167] The expression and / or function of MHC class I (MHC-I, specifically HLA-A, HLA-B, and HLA-C) may be reduced or eliminated to prevent HvG and consequently improve cell engraftment in the target. Improved engraftment is thought to result in longer cell persistence and therefore a larger therapeutic area of the target. In some embodiments, the expression and / or function of beta-2-microglobulin (B2M), a structural element of MHC-I, is reduced or eliminated. Non-limiting examples of guide RNAs (gRNAs) for targeting and deleting MHC activators are disclosed in PCT application publication number WO2020 / 051374, which is incorporated herein by reference in whole with respect to examples of gRNAs that may be used in the methods described herein.
[0168] A detailed description of genetic modifications of endogenous sequences encoding the naturally occurring chimeric stimulatory receptors, TCR-alpha (TCR-α), TCR-beta (TCR-β), and / or beta-2-microglobulin (β2M), as well as naturally occurring polypeptides including the HLA class I histocompatibility antigen, alpha-E (HLA-E) polypeptide, is disclosed in PCT application publication number WO2020 / 051374, which is incorporated herein by reference in its entirety, with respect to the naturally occurring receptors and polypeptides, as well as examples of other genetic modifications that may be introduced into cells as disclosed herein.
[0169] Under normal conditions, complete T cell activation depends on the involvement of the TCR in conjunction with a second signal mediated by one or more co-stimulatory receptors (e.g., CD28, CD2, 4-1BBL) that enhance the immune response. However, in the absence of a TCR, stimulation with standard activation / stimulation reagents containing an agonist anti-CD3 mAb significantly reduces T cell proliferation. Therefore, this disclosure provides a chimeric stimulatory receptor (CSR) that does not exist in nature, comprising: (a) an ectodomain comprising an activating component, the activating component being isolated or induced from a first protein; (b) a transmembrane domain; and (c) an endodomain comprising at least one signaling domain, the at least one signaling domain being isolated or induced from a second protein (the first and second proteins are not identical).
[0170] The activating components of CSRs described herein may include a portion of one or more components of T cell receptors (TCRs), TCR complexes, TCR coreceptors, TCR costimulatory proteins, TCR inhibitory proteins, cytokine receptors, and chemokine receptors to which the agonist of the activating component binds. The activating component may include the extracellular domain of CD2 or a portion thereof to which the agonist binds.
[0171] The signaling domains of CSRs described herein may include one or more components of human signaling domains, TCR complexes, TCR coreceptors, TCR costimulatory proteins, TCR inhibitory proteins, cytokine receptors, and chemokine receptors. The signaling domain may include a CD3 protein or a portion thereof. The CD3 protein may include a CD3ζ domain or a portion thereof.
[0172] The endodomain of the CSR described herein may further comprise a cytoplasmic domain. The ectodomain of the CSR may further comprise a signal peptide. The CSR may further comprise a transmembrane domain. The cytoplasmic domain, signal peptide, and / or transmembrane domain may originate from the same protein or different proteins. For example, the signal peptide and transmembrane domain may originate from CD8 or CD28 or a combination thereof.
[0173] This disclosure also provides non-naturally occurring ectodomain CSRs, including modified ectodomains. Modifications may include mutations or cleavage of the amino acid sequence of the active component compared to the wild-type sequence of the active component. Mutations or cleavage of the amino acid sequence of the active component may include mutations or cleavage of the CD2 extracellular domain or a portion thereof to which the agonist binds. Mutations or cleavage of the CD2 extracellular domain may reduce or eliminate binding to naturally occurring CD58.
[0174] This disclosure also provides nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides transposons or vectors comprising nucleic acid sequences encoding any CSR disclosed herein.
[0175] This disclosure provides cells containing any CSR disclosed herein. This disclosure provides cells containing nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides cells containing vectors containing nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides cells containing transposons containing nucleic acid sequences encoding any CSR disclosed herein.
[0176] This disclosure provides compositions comprising any CSR disclosed herein. This disclosure provides compositions comprising nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides compositions comprising vectors comprising nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides compositions comprising transposons comprising nucleic acid sequences encoding any CSR disclosed herein. This disclosure provides compositions comprising modified cells disclosed herein or compositions comprising a plurality of modified cells disclosed herein.
[0177] Methods for site-specific gene insertion are also provided herein. Chimeric transposases and chimeric site-specific transposase fusion proteins provided herein may be used to deliver a transgene to a cell and to incorporate the transgene into a target site. The target site may be, for example, a genome-safe harbor, i.e., a genomic site into which the transgene can be incorporated in such a way that the transgene functions predictably and does not cause any alteration of the host genomic DNA sequence. In some embodiments, the target site is a repeating element, such as a LINE-1 or ALU sequence. The repeating element does not encode an essential gene product, reducing the likelihood that the insertion will result in a detrimental change to the cell's gene expression profile. One, two, or more target sites may be present within a single repeating element. In some embodiments, the target site is located within an intron.
[0178] Site-directed integration can be used in vitro or in vivo. One example of in vivo application is gene therapy, which involves the delivery of a transgene to the genomic DNA of a cell.
[0179] Formulation, dosage, and mode of administration This disclosure provides formulations, dosages, and methods for administering the compositions and cells described herein. In one embodiment, a pharmaceutical composition comprising a chimeric transposase or chimeric site-specific transposase fusion protein described herein and a pharmaceutically acceptable carrier is provided herein. In another embodiment, a pharmaceutical composition comprising modified cells described herein and a pharmaceutically acceptable carrier is provided herein.
[0180] The disclosed compositions and pharmaceutical compositions may contain any suitable adjuvants, for example, at least one of diluents, binders, stabilizers, buffers, salts, lipophilic solvents, preservatives, and adjuvants. Pharmaceutically acceptable adjuvants are preferred. Non-limited examples of such sterile solutions and methods for their preparation are well known in the art, for example, but not limited to Gennaro, Ed., Remington's Pharmaceutical Sciences, 18th Edition, Mack Publishing Co. (Easton, Pa.) 1990 and "Physician's Desk Reference", 52nd ed., Medical Economics (Montvale, NJ) 1998. Pharmaceutically acceptable carriers suitable for the dosage, solubility, and / or stability of protein backbone, fragment, or variant compositions can be typically selected, as are well known in the art or as described herein.
[0181] Non-limiting examples of suitable pharmaceutical excipients and additives include proteins, peptides, amino acids, lipids, and carbohydrates (e.g., sugars including monosaccharides, disaccharides, trisaccharides, tetrasaccharides, and oligosaccharides; derivatized sugars such as alditol, aldonic acid, and esterified sugars; and polysaccharides or sugar polymers), which may exist alone or in combination, and which constitute 1 to 99.99% (by weight or volume). Non-limiting examples of protein excipients include serum albumin, e.g., human serum albumin (HSA), recombinant human albumin (rHA), gelatin, and casein. Representative amino acid / protein components that can also function in buffering capacity include alanine, glycine, arginine, betaine, histidine, glutamic acid, aspartic acid, cysteine, lysine, leucine, isoleucine, valine, methionine, phenylalanine, and aspartame. One preferred amino acid is glycine.
[0182] Non-limiting examples of carbohydrate excipients suitable for use include monosaccharides such as fructose, maltose, galactose, glucose, D-mannose, and sorbose; disaccharides such as lactose, sucrose, trehalose, and cellobiose; polysaccharides such as raffinose, meletitose, maltodextrin, dextran, and starch; and algitols such as mannitol, xylitol, maltitol, lactitol, xylitol sorbitol (glucitol), and myo-inositol. Preferably, the carbohydrate excipient is mannitol, trehalose, and / or raffinose.
[0183] The composition may also contain buffers or pH adjusters. Typically, buffers are salts prepared from organic acids or bases. Representative buffers include organic acid salts, such as salts of citric acid, ascorbic acid, gluconic acid, carbonate, tartaric acid, succinic acid, acetic acid, or phthalic acid; Tris, tromethamine hydrochloride, or phosphate buffers. Preferred buffers are organic acid salts such as citrates.
[0184] Furthermore, the disclosed compositions may include polymer excipients / additives, such as polyvinylpyrrolidone, Ficol (polymer sugar), dextrose (e.g., cyclodextrin, e.g., 2-hydroxypropyl-β-cyclodextrin), polyethylene glycol, flavoring agents, antimicrobial agents, sweeteners, antioxidants, antistatic agents, surfactants (e.g., polysorbates, e.g., "TWEEN 20" and "TWEEN 80"), lipids (e.g., phospholipids, fatty acids), steroids (e.g., cholesterol), and chelating agents (e.g., EDTA).
[0185] Many known and developed methods can be used to administer a therapeutically effective amount of the composition or pharmaceutical composition disclosed herein. Non-limiting examples of administration methods include bolus, buccal, injection, intra-articular, intra-bronchial, intraperitoneal, intrasacral, intra-cartilaginous, intracavitary, intraperitoneal, intracerebellar, intraventricular, intracolonic, intracervical, intragastric, intrahepatic, intrafocal, intramuscular, intramyocardial, intranasal, intraocular, intraosseous, intraskeletal, intrapelvic, intrapericardial, intraperitoneal, intrapleural, intraprostatic, intrapulmonary, intrarectal, intrarenal, intraretinal, intraspinal cord, synovial, intrathoracic, intrauterine, intratumoral, intravenous, intrabladder, oral, parenteral, rectal, sublingual, subcutaneous, percutaneous, or vaginal means. In preferred embodiments, the composition comprising the modified cells described herein is administered intravenously, for example, by intravenous infusion.
[0186] The compositions of this disclosure can be prepared for use in parenteral administration (subcutaneous, intramuscular, or intravenous) or any other administration, particularly in the form of liquid solutions or suspensions. For parenteral administration, the compositions disclosed herein can be formulated as solutions, suspensions, emulsions, particles, powders, or lyophilized powders, either in combination with a pharmaceutically acceptable parenteral vehicle or provided separately. Formulations for parenteral administration may contain, as common excipients, sterile water or saline, polyalkylene glycols such as polyethylene glycol, plant-derived oils, hydrogenated naphthalene, and the like. Aqueous or oily suspensions for injection can be prepared by using appropriate emulsifiers or humectants and suspending agents according to known methods. Drugs for injection or infusion may be non-toxic, orally unadministerable diluents such as aqueous solutions, sterile injection solutions, or suspensions in solvents. Suitable vehicles or solvents include water, Ringer's solution, isotonic saline, and the like; sterile non-volatile oils can be used as common solvents or suspension solvents. For these purposes, any type of non-volatile oil and fatty acid, including natural, synthetic, or semi-synthetic fatty oils or fatty acids; natural, synthetic, or semi-synthetic mono-, di-, or tri-glycerides, may be used. Parenteral administration is known in the art and is not limited to conventional injection methods, including gas-pressurized needle-free injection devices such as those described in U.S. Patent No. 5,851,198, and laser puncture devices such as those described in U.S. Patent No. 5,839,446.
[0187] It may be desirable to deliver the disclosed compounds to subjects over a long period, for example, from a single dose to one week to one year. Various sustained-release, depot, or implantable dosage forms can be utilized. For example, dosage forms may include pharmaceutically acceptable non-toxic salts of compounds with low solubility in body fluids, e.g., (a) acid addition salts with polybasic acids, e.g., phosphoric acid, sulfuric acid, citrate, tartaric acid, tannic acid, pamoic acid, alginic acid, polyglutamic acid, naphthalene mono or disulfonic acid, polygalacturonic acid, etc., (b) salts with polyvalent metal cations, e.g., zinc, calcium, bismuth, barium, magnesium, aluminum, copper, cobalt, nickel, cadmium, etc., or salts with organic cations formed from, for example, N,N'-dibenzyl ethylenediamine or ethylenediamine, or (c) combinations of (a) and (b), e.g., zinc tannate salt. Furthermore, the disclosed compounds or preferably relatively insoluble salts, such as those described above, can be formulated into gels suitable for injection, such as aluminum monostearate gel containing sesame oil. Particularly preferred salts include zinc salts, zinc tannate salts, pamoate salts, etc. Another type of sustained-release depot formulation for injection contains the compound or salt dispersed for encapsulation in a slowly degradable, non-toxic, non-antigenic polymer, such as polylactic acid / polyglycolic acid polymer, as described, for example, in U.S. Patent No. 3,773,919. The compounds or preferably relatively insoluble salts, such as those described above, can also be formulated into cholesterol matrix silastic pellets, particularly for use in animals. Further sustained-release formulations, depot formulations, or implant formulations, such as gaseous or liquid liposomes, are known in the literature (U.S. Patent No. 5,770,222 and “Sustained and Controlled Release Drug Delivery Systems”, JR Robinson ed. Marcel Dekker, Inc., NY, 1978).
[0188] Treatment method In another embodiment, a method for treating a disease or disorder of interest is provided herein, comprising administering a composition comprising the modified cells described herein to the subject. The terms “subject” and “patient” are used interchangeably herein. In a preferred embodiment, the patient is human.
[0189] The modified cells may be allogeneic or autologous to the patient. In some preferred embodiments, the modified cells are allogeneic cells. In some embodiments, the modified cells are autologous T cells or modified autologous CAR T cells. In some preferred embodiments, the modified cells are allogeneic T cells or modified allogeneic CAR T cells.
[0190] In some embodiments, the disease or disorder treated according to the methods described herein is cancer. In some embodiments, the treatment methods described herein may delay the progression of cancer and / or reduce the tumor burden.
[0191] In some embodiments, the disease or disorder treated according to the methods described herein is an autoimmune disease. In some embodiments, the autoimmune diseases are autoimmune neutropenia, Guillain-Barré syndrome, epilepsy, autoimmune encephalitis, Isaac syndrome, nevus syndrome, pemphigus vulgaris, pemphigus deciduousis, bullous pemphigoid, acquired epidermolysis bullosa, pemphigoid of pregnancy, mucosal pemphigoid, antiphospholipid syndrome, autoimmune anemia, myasthenia gravis, autoimmune Graves' disease, thyroid eye disease (TED), Goodpasture syndrome, multiple sclerosis, rheumatoid arthritis, lupus, idiopathic thrombocytopenic purpura (ITP), warm autoimmune hemolytic anemia (WAIHA), chronic inflammatory demyelinating polyneuropathy (CIDP), lupus nephritis, or membranous nephropathy.
[0192] The dosage of the pharmaceutical composition administered to the subject may vary depending on known factors, such as the pharmacodynamic characteristics of a particular drug, as well as its mode and route of administration; the recipient's age, health condition, and weight; the nature and severity of the symptoms, the type and frequency of concomitant treatments, and the desired effect.
[0193] In embodiments where the composition administered to the subject of interest is the modified cell disclosed herein, from about 1x10 3 to about 1x10 4 cells; from about 1x10 4 to about 1x10 5 cells; from about 1x10 5 to about 1x10 6 cells; from about 1x10 6 to about 1x10 7 cells; from about 1x10 7 to about 1x10 8 cells; from about 1x10 8 to about 1x10 9 cells; from about 1x10 9 to about 1x10 10 cells, from about 1x10 10 to about 1x10 11 cells, from about 1x10 11 to about 1x10 12 cells, from about 1x10 12 to about 1x10 13 cells, from about 1x10 13 to about 1x10 14 cells, from about 1x10 14 to about 1x10 15 cells, from about 1x10 15 to about 1x10 16 cells, from about 1x10 16 to about 1x10 17 cells, from about 1x10 17 to about 1x10 18 cells, from about 1x10 18 to about 1x10 19 cells; or from about 1x10 19 to about 1x10 20 cells may be administered. In some embodiments, the cells are administered at a dose of from about 5x10 6 to about 25x10 6 cells.
[0194] In other embodiments, the dosage of the cells may depend on the body weight of the human, for example, from about 1x10 3 to about 1x10 4 cells; from about 1x10 4~approximately 1x10 5 Individual cells; approximately 1 x 10⁻¹⁶ 5 ~approximately 1x10 6 Individual cells; approximately 1 x 10⁻¹⁶ 6 ~approximately 1x10 7 Individual cells; approximately 1 x 10⁻¹⁶ 7 ~approximately 1x10 8 Individual cells; approximately 1 x 10⁻¹⁶ 8 ~approximately 1x10 9 Individual cells; approximately 1 x 10⁻¹⁶ 9 ~approximately 1x10 10 Each cell, approximately 1 x 10⁶ 10 ~approximately 1x10 11 Each cell, approximately 1 x 10⁶ 11 ~approximately 1x10 12 Each cell, approximately 1 x 10⁶ 12 ~approximately 1x10 13 Each cell, approximately 1 x 10⁶ 13 ~approximately 1x10 14 Each cell, approximately 1 x 10⁶ 14 ~approximately 1x10 15 Each cell, approximately 1 x 10⁶ 15 ~approximately 1x10 16 Each cell, approximately 1 x 10⁶ 16 ~approximately 1x10 17 Each cell, approximately 1 x 10⁶ 17 ~approximately 1x10 18 Each cell, approximately 1 x 10⁶ 18 ~approximately 1x10 19 A single cell; or about 1 x 10⁶ 19 ~approximately 1x10 20 A single cell can be administered per kilogram of body weight of the subject.
[0195] A more detailed description of the disclosed compositions and pharmaceutically acceptable excipients, formulations, dosages, and methods of administration of the pharmaceutical compositions is disclosed in International Publication No. 2019 / 049816.
[0196] The transposon domains and fusion proteins provided herein may be used to deliver gene therapy. Gene therapy typically involves the delivery of a transgene into the genomic DNA of a cell. Typically, the transgene replaces a gene that is mutated or otherwise not properly expressed within the cell. The fusion proteins, transposase domains, and complexes described herein may be used to deliver therapeutic transgenes to cells and to incorporate the transgenes into target sites. In some embodiments, the treatment method comprises introducing a fusion protein and transposon described in any one of claims 1 to 13 into a cell, wherein the transposon comprises a 5'ITR, a transgene, and a 3'ITR in the order of 5' to 3'.
[0197] kit In another embodiment, the foregoing provides a kit comprising a cell line engineered to contain a modified target site for an SPB or PBx provided herein within its genome, preferably within a highly expressed genomic region. The kit may further comprise a composition comprising one or more SPB or PBx transposase domains or fusion proteins described herein. In some embodiments, the cell line is a T cell line.
[0198] term As used throughout this disclosure, the singular forms “a,” “an,” and “the” include multiple referents unless the context otherwise explicitly indicates otherwise. Thus, for example, a reference to “a method” includes multiple such methods, and a reference to “a dose” includes one or more doses and their equivalents known to those skilled in the art.
[0199] The terms “about” or “approximately” can mean that a particular value is within an acceptable margin of error, as determined by those skilled in the art, and this depends to some extent on the method by which the value is measured or determined, for example, on the limitations of the measuring system. For example, “about” can mean within one or more standard deviations. Alternatively, “about” can mean a range of up to 20%, or up to 10%, or up to 5%, or up to 1% of a given value. Or, particularly with respect to biological systems or processes, the term can mean within one order of magnitude of the value, preferably up to five times, and more preferably up to two times. Where a particular value is described in this application and claims, unless otherwise specified, the term “about” should be assumed to mean within an acceptable margin of error of that particular value.
[0200] This disclosure provides isolated or substantially purified polynucleotide or protein compositions. “Isolated” or “purified” polynucleotides or proteins, or their biologically active portions, substantially or essentially contain components that would normally accompany or interact with the polynucleotide or protein as found in its naturally occurring environment. Therefore, isolated or purified polynucleotides or proteins, if produced by recombinant technology, substantially contain other cell material or culture medium, or if chemically synthesized, substantially contain chemical precursors or other chemicals. Optimally, “isolated” polynucleotides do not contain sequences naturally adjacent to the polynucleotide in the genomic DNA of the organism from which the polynucleotide originates (i.e., sequences located at the 5' and 3' ends of the polynucleotide) (optimally protein-coding sequences). For example, in various embodiments, isolated polynucleotides may contain nucleotide sequences of approximately 5kb, 4kb, 3kb, 2kb, 1kb, 0.5kb, or less than 0.1kb that are naturally adjacent to the polynucleotide in the genomic DNA of the cell from which the polynucleotide originates. Substantially cellular material-free proteins include protein preparations containing approximately 30%, 20%, 10%, 5%, or less than 1% (dry weight) of contaminated proteins. When the proteins of this disclosure or their biologically active portions are recombinantly produced, the culture medium optimally contains approximately 30%, 20%, 10%, 5%, or less than 1% (dry weight) of chemical precursors or non-target protein chemicals.
[0201] This disclosure provides disclosed DNA sequences and fragments and variants of proteins encoded by these DNA sequences. As used throughout this disclosure, the term “fragment” refers to a portion of a DNA sequence or a portion of an amino acid sequence, and therefore to the protein encoded thereby. A DNA sequence fragment consisting of a coding sequence may encode a protein fragment that retains the biological activity of the native protein and thus retains DNA recognition or binding activity to a target DNA sequence, as described herein. Alternatively, a DNA sequence fragment useful as a hybridization probe will generally not encode a protein that retains biological activity or promoter activity. Therefore, DNA sequence fragments may range from at least about 20 nucleotides, about 50 nucleotides, about 100 nucleotides, and up to the full-length polynucleotides of this disclosure.
[0202] The term “comprising” is intended to mean that a composition and method includes the enumerated elements but does not exclude other elements. “Consisting essentially of,” when used to define a composition and method, when used for its intended purpose, means excluding other elements that are essentially important to the combination. Thus, a composition consisting essentially of the elements defined herein does not exclude trace amounts of contaminants or inert carriers. “Consisting of” means excluding elements that are more than trace amounts of other components and substantial method steps. The embodiments defined by each of these transitional terms are within the scope of this disclosure.
[0203] As used herein, “expression” refers to the process by which a polynucleotide is transcribed into mRNA and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. If the polynucleotide is derived from genomic DNA, expression may include the splicing of mRNA in eukaryotic cells.
[0204] Gene expression refers to the conversion of information contained in a gene into a gene product. A gene product can be a direct transcript of a gene (e.g., mRNA, tRNA, rRNA, antisense RNA, ribozyme, shRNA, microRNA, structural RNA, or any other type of RNA) or a protein produced by the translation of mRNA. Gene products also include RNA modified by processes such as capping, polyadenylation, methylation, and editing, as well as proteins modified by processes such as methylation, acetylation, phosphorylation, ubiquitination, ADP-ribosylation, myristylation, and glycosylation.
[0205] The term “operatively linked” or its equivalent (e.g., “linked operatively”) means that two or more molecules are positioned relative to each other so that they can interact and influence the function of one or both molecules or a combination thereof. In relation to nucleic acids, a promoter may be operatively linked to a nucleotide sequence encoding a transpose domain or fusion protein as described herein, resulting in the expression of the nucleotide sequence under the control of the promoter.
[0206] Non-covalently linked components and methods for constructing and using non-covalently linked components are disclosed. Various components can take on various different forms as described herein. For example, non-covalently linked (i.e., operably linked) proteins can be used to enable transient interactions that avoid one or more problems in the art. The ability of non-covalently linked components, such as proteins, to associate and dissociate allows for functional association only under circumstances where such association is required for the desired activity, or primarily under such circumstances. The linkage may be for a duration sufficient to enable the desired effect.
[0207] A method for inducing a protein to a specific gene locus within the genome of an organism is disclosed. The method may include steps of providing a DNA localization component and providing an effector molecule, the DNA localization component and the effector molecule being operablely linked via non-covalent bonds.
[0208] A "target site" or "target sequence" is a nucleic acid sequence that defines the portion of the nucleic acid to which a binding molecule will bind, provided that sufficient conditions for binding are present.
[0209] The terms “nucleic acid,” “oligonucleotide,” or “polynucleotide” refer to at least two nucleotides covalently linked to one another. A single-stranded description also defines the sequence of the complementary strand. Thus, a nucleic acid may also encompass the complementary strand of the single-stranded molecule described. The nucleic acids of this disclosure also include substantially identical nucleic acids and their complements that retain the same structure or encode the same protein.
[0210] The nucleic acids of this disclosure may be single-stranded or double-stranded. The nucleic acids of this disclosure may contain double-stranded sequences even if the majority of the molecule is single-stranded. The nucleic acids of this disclosure may contain single-stranded sequences even if the majority of the molecule is double-stranded. The nucleic acids of this disclosure may include genomic DNA, cDNA, RNA, or hybrids thereof. The nucleic acids of this disclosure may contain combinations of deoxyribonucleotides and ribonucleotides. The nucleic acids of this disclosure may contain combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, and isoguanine. The nucleic acids of this disclosure may be synthesized to include non-natural amino acid modifications. The nucleic acids of this disclosure may be obtained by chemical synthesis or recombinant methods.
[0211] The nucleic acids of this disclosure may not exist in nature in whole or in any part thereof. The nucleic acids of this disclosure may contain one or more mutations, substitutions, deletions or insertions that do not exist in nature, making the entire nucleic acid sequence non-natural. The nucleic acids of this disclosure may contain one or more replicate sequences, inverted sequences or repeat sequences, resulting in sequences that do not exist in nature, making the entire nucleic acid sequence non-natural. The nucleic acids of this disclosure may contain modified nucleotides, artificial nucleotides or synthetic nucleotides that do not exist in nature, making the entire nucleic acid sequence non-natural.
[0212] Considering the redundancy of the genetic code, multiple nucleotide sequences can encode any particular protein. All such nucleotide sequences are assumed herein.
[0213] As used throughout this disclosure, the term “promoter” refers to a synthetic or naturally occurring molecule that can confer, activate, or enhance the expression of nucleic acids within a cell. A promoter may include one or more specific transcriptional regulatory sequences to further enhance expression and / or alter its spatial and / or temporal expression. A promoter may also include distal enhancer or repressor elements that may be located thousands of base pairs away from the transcription start site. Promoters may originate from sources including viruses, bacteria, fungi, plants, insects, and animals. Promoters can constitutively or differentially modulate the expression of gene components with respect to the cell, tissue, or organ in which expression occurs, or with respect to the developmental stage in which expression occurs, or in response to external stimuli such as physiological stress, pathogens, metal ions, or inducers. Representative examples of promoters include the bacteriophage T7 promoter, bacteriophage T3 promoter, SP6 promoter, lac operator promoter, tac promoter, SV40 late promoter, SV40 early promoter, RSV-LTR promoter, CMV IE promoter, EF-1 alpha promoter, CAG promoter, SV40 early promoter or SV40 late promoter, and CMV IE promoter.
[0214] As used throughout this disclosure, the term “vector” refers to a nucleic acid sequence containing an origin of replication. A vector may be a viral vector, a bacteriophage, a bacterial artificial chromosome, or a yeast artificial chromosome. A vector may be a DNA or RNA vector. A vector may be a self-replicating extrachromosomal vector, preferably a DNA plasmid. A vector may contain a combination of amino acids and a DNA sequence, an RNA sequence, or both a DNA sequence and an RNA sequence.
[0215] Conservative substitutions of amino acids, i.e., substitution of an amino acid with a different amino acid having similar properties (e.g., hydrophilicity, degree and distribution of charged region), are recognized in the art as typically resulting in only minor changes. These minor changes can be partially identified by considering the hydropathic index of the amino acid, as understood in the art. Kyte et al., J.Mol.Biol.157:105-132 (1982). The hydropathic index of an amino acid is based on consideration of its hydrophobicity and charge. Amino acids with similar hydropathic indexes may be substituted and still retain protein function. In one embodiment, amino acids with hydropathic indexes of ±2 are substituted. The hydrophilicity of amino acids can also be used to identify substitutions that result in proteins that retain biological function. By considering the hydrophilicity of an amino acid in relation to a peptide, it becomes possible to calculate the maximum local mean hydrophilicity of that peptide, which is a useful measure that has been reported to correlate well with antigenicity and immunogenicity. U.S. Patent No. 4,554,101, fully incorporated herein by reference.
[0216] Substitutions of amino acids with similar hydrophilicity values can result in peptides that retain biological activity, such as immunogenicity. Substitutions can be made with amino acids having hydrophilicity values within ±2 of each other. Both the hydrophobicity index and hydrophilicity of an amino acid are influenced by its specific side chain. Consistent with this observation, it is understood that amino acid substitutions suitable for biological function depend on the relative similarity of the amino acids, particularly their side chains, as revealed by their hydrophobicity, hydrophilicity, charge, size, and other properties.
[0217] As used herein, “conservative” amino acid substitutions may be defined as shown in Tables 1, 2, and 3 below. In some embodiments, fusion polypeptides and / or nucleic acids encoding such fusion polypeptides include conservative substitutions introduced by modifications of the polynucleotides encoding the polypeptides of this disclosure. Amino acids can be classified according to their physical properties and their contributions to secondary and tertiary protein structures. A conservative substitution is the substitution of one amino acid for another amino acid having similar properties. Exemplary conservative substitutions are shown in Table 1. [Table 1]
[0218] Alternatively, conserved amino acids can be classified as shown in Table 2, as presented in Lehninger (Biochemistry, Second Edition; Worth Publishers, Inc. NY, NY (1975), pp. 71-77). [Table 2]
[0219] Alternatively, exemplary conservative substitutions are shown in Table 3. [Table 3]
[0220] The polypeptides and proteins of this disclosure may not exist in nature in whole or in any part thereof. The polypeptides and proteins of this disclosure may contain one or more mutations, substitutions, deletions or insertions that do not exist in nature, resulting in an entire amino acid sequence that does not exist in nature. The polypeptides and proteins of this disclosure may contain one or more duplicate sequences, inverted sequences or repeat sequences, resulting in a sequence that does not exist in nature, resulting in an entire amino acid sequence that does not exist in nature. The polypeptides and proteins of this disclosure may contain modified amino acids, artificial amino acids, or synthetic amino acids that do not exist in nature, resulting in an entire amino acid sequence that does not exist in nature.
[0221] When used throughout this disclosure, the identity between two sequences may be determined by using a standalone executable BLAST engine program (bl2seq) for blasting two sequences, which is available from the National Center for Biotechnology Information (NCBI) ftp site using default parameters (Tatusova and Madden, FEMS Microbiol Lett., 1999, 174, 247-250; the entire program is incorporated herein by reference). When used in the context of two or more nucleic acid or polypeptide sequences, the terms “identical” or “identity” refer to a specific percentage of residues that are the same across a particular region of each sequence. In some embodiments, sequence identification is determined across the entire length of the sequences. The percentage can be calculated by optimally aligning the two sequences, comparing the two sequences across a given region, determining the number of positions where identical residues occur in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the given region, and multiplying the result by 100 to obtain the percentage of sequence identity. If the two sequences are of different lengths, or if the alignment produces one or more staggered ends, and a particular comparison region contains only a single sequence, the residues of the single sequence are included in the denominator but not in the numerator of the calculation. When comparing DNA and RNA, thymine (T) and uracil (U) can be considered equivalent. Identity can be determined manually or by using computer sequencing algorithms such as BLAST or BLAST 2.0.
[0222] In certain embodiments, if a sequence has a specific sequence identity with a specific sequence number (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%), the sequence and the sequence of that sequence number have the same length. In certain embodiments, if a sequence has a specific sequence identity with a specific sequence number (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%), the sequence and the sequence of that sequence number differ only for the sake of conservation amino acid substitutions.
[0223] As used throughout this disclosure, the term “endogenous” refers to a nucleic acid or protein sequence that is naturally associated with the target gene or the host cell into which it is introduced.
[0224] As used throughout this disclosure, the term “exogenous” refers to a nucleic acid or protein sequence that is not naturally associated with the target gene or the host cell into which it is introduced, including multiple copies of a naturally occurring nucleic acid (e.g., a DNA sequence) that are not naturally occurring, or a naturally occurring nucleic acid sequence located at a genomic location that is not naturally occurring.
[0225] This disclosure provides a method for introducing a polynucleotide construct containing a DNA sequence into a host cell. “Introducing” means presenting the polynucleotide construct to the cell so that the construct can access the interior of the host cell. The method of this disclosure does not depend on a specific method for introducing the polynucleotide construct into a host cell, and only the polynucleotide construct can access the interior of a single host cell. Methods for introducing polynucleotide constructs into bacteria, plants, fungi, and animals are known in the art and include, but are not limited to, stable transformation methods, transient transformation methods, and virus-mediated methods. [Examples]
[0226] The examples in this section are provided for illustrative purposes only and are not intended to limit the invention.
[0227] Example 1: Construction of a chimeric PiggyBac ("PB")-MosI transposase Chimeric PB:MosI transposases were created by deleting the CRD domain from PB transposase and replacing it with the DNA-binding domain of MosI transposase. The PB CRD consists of residues 553–594 of the 594-amino acid PB transposase protein (see SEQ ID NO: 2). The CRD domain was bound to the rest of the transposase via a linker sequence extending from residues 535–552 of the PB transposase sequence. The PB CRD domain may also be cleaved from the C-terminus of PB transposase somewhere within the linker sequence. Two fusion sites within the linker sequence were selected as exemplary sites. In the first fusion design, the C-terminus of PB transposase was cleaved from residues 545–594 (SEQ ID NO: 3) to obtain the S544 fusion site (S544 PB) and construct the S544 PB-MosI chimera. In the second fusion design, the C-terminus of the PB transposase was cleaved from residues 553-594 (SEQ ID NO: 4) to obtain the V552 fusion site (V552 PB), and the chimeric V552 PB-MosI transposase was constructed.
[0228] The C-terminally cleaved S544PB and V552 PB sequences were fused to the DNA-binding domain of MosI. Unlike PB, which binds to its N-terminus proximal to the transposon end, MosI binds to its C-terminus proximal to the transposon end. The N-terminus of MosI contains two DNA-binding domains fused to the catalytic domain of the transposase. The first DNA-binding domain extends over approximately the first 58 residues of the protein (SEQ ID NO: 8), while the two DNA-binding domains extend over approximately the first 111 residues of the protein (SEQ ID NO: 6).
[0229] As described in International Patent Application Publication No. PCT / US2022 / 77549, we produced a TAL-piggyBac transposase fusion protein containing a GFP-targeting N-terminal deletion piggyBac transposase sequence and an embedded defect N-terminal piggyBac transposase. Briefly, we designed two pairs of TAL arrays targeting sequences in the GFP coding sequence and synthesized each TAL array containing nine 34-amino acid repeats followed by 20-amino acid "half" repeats, adjacent to the BsmBI type IIS restriction site. This allowed us to clone each TAL array in frame with the remainder of the open reading frame containing the N-terminal deletion piggyBac transposase sequence in the expression plasmid to generate GFP1R Right TAL-PBx (SEQ ID NO: 40). By replacing the PBx transposase sequence with a chimeric PB:MosI transposase sequence, we constructed a chimeric site-specific transposase fusion protein that specifically targets GFP.
[0230] A chimeric PB:MosI transposase sequence was constructed by binding residues 3-58 of MosI to the S544 PBx transposase using the GGGGS linker sequence (SEQ ID NO: 86), and then inserting this into the PBx transposase sequence to create GFP1R TAL-PB:S544-MosI-58 (SEQ ID NO: 41). Similarly, residues 3-58 of MosI were fused to the V552 PBx transposase sequence using the GGGGS linker (SEQ ID NO: 86), and then inserting this into the PBx transposase sequence to create GFP1R TAL-PB:V552-MosI-58 (SEQ ID NO: 42). Furthermore, residues 3-111 of MosI were fused to S544 PB using the GGGGS linker (SEQ ID NO: 86), and then inserting this into the PBx transposase sequence to create GFP1R TAL-PB:S544-MosI-111 (SEQ ID NO: 43). Finally, using the GGGGS linker (SEQ ID NO: 86), residues 3-111 of MosI were fused to PiggyBac with V552 PBx and inserted in place of the PBx transposase sequence to create GFP1R TAL-ssSPB:V552-MosI-111 (SEQ ID NO: 44).
[0231] Example 2: Construction of Chimeric PB:MosI Inverted Terminal Repeat (ITR) This example demonstrates an exemplary method for constructing a chimeric PB:MosI ITR sequence.
[0232] Each of the chimeric PB:MosI transposases prepared in Example 1 contains a PB DNA-binding and dimerization domain (DDBD) fused to a MosI-derived DNA-binding domain(s). Therefore, a chimeric PB:MosI ITR sequence was constructed that included the binding site of the PB DDBD fused to the binding site of the MosI DNA-binding domain.
[0233] The 35bp LE PB ITR sequence (SEQ ID NO: 12) was cleaved after position 16, resulting in a deletion of a 19bp CRD binding site. For this purpose, a 23bp MosI sequence (SEQ ID NO: 13) containing the binding sites of two MosI DNA binding domains was fused in reverse to create the LE PB:MosI ITR (SEQ ID NO: 15). In addition, alternative versions of the chimeric ITR, including those with various spacer lengths between the PB domain and the MosI ITR domain in the chimeric protein, were designed to test the effect of linker length. For example, a 1bp MosI binding site was deleted adjacent to the DDBD binding site to create LE PB:MosI ITR-1 (SEQ ID NO: 16). Furthermore, 1 or 2bp of DNA was added between the MosI binding site and the DDBD binding site to create LE PB:MosI ITR+1 (SEQ ID NO: 17) and LE PB:MosI ITR+2 (SEQ ID NO: 18). By similarly modifying the 63bp RE ITR (sequence number 65), we created RE PB:MosI ITR (sequence number 19), RE PB:MosI ITR-1 (sequence number 20), RE PB:MosI ITR+1 (sequence number 21), and RE PB:MosI ITR+2 (sequence number 22).
[0234] Chimeric PB:MosI transposases and chimeric ITR sequences were analyzed for their relative resective activity.
[0235] Example 3: Method for measuring the cleavage activity of chimeric PB-MosI transposase using chimeric PB:MosI ITR This example describes an assay designed to measure the cleavage activity of chimeric PB-MosI transposases and chimeric ITRs. In this assay, each chimeric PB-MosI transposase prepared in Example 1 is co-administered to cells with a reporter transposon construct, the transposon comprising a DNA nucleotide sequence encoding non-functional GFP whose coding sequence is disrupted by intervening DNA fragments flanked by TTAA sequences, and a pair of inverted terminal repeats (ITRs) prepared in Example 2. The TTAA sequences and ITRs function as recognition sites for the chimeric PB:MosI transposase, and if the chimeric transposase has cleavage activity, the intervening DNA is cleaved, restoring the intact full-length coding sequence of the GFP gene. Therefore, chimeric transposases that have transposase activity against constructs containing congeneral chimeric ITR sequences can produce GFP-positive cells in this assay, which can then be identified and quantified by FACS.
[0236] In short, the GFP reporter system includes an EF1a promoter (SEQ ID NO: 45) that drives the expression of the GFP reporter (SEQ ID NO: 46), followed by an SV40 polyadenylated sequence (SEQ ID NO: 47). The GFP reporter is disrupted at the TTAA sequence by a transposon, which degrades it into a first GFP moiety (SEQ ID NO: 48) and a second GFP moiety (SEQ ID NO: 49). Disrupted GFP reporters were constructed from each chimeric ITR from Example 2, including the PB:MosI ITR transposon (SEQ ID NO: 50), the PiggyBac:MosI ITR-1 transposon (SEQ ID NO: 51), the PiggyBac:MosI ITR+1 transposon (SEQ ID NO: 52), or the PiggyBac:MosI ITR+2 transposon (SEQ ID NO: 53).
[0237] Using GFP reporter systems, the cleavage activity of the GFPR1 TAL-PB:S544-MosI-58, GFP1R TAL-PB:V552-MosI-58, GFPR1 TAL-PB:S544-MosI-111, and GFP1R TAL-ssSPB:V552-MosI-111 transposases described in Example 1 was determined using each chimeric PB:MosI ITR prepared in Example 2. On day 0, each reporter was co-transfected with each chimeric transposase into HEK293T cells. Briefly, 120,000 cells were seeded in 24-well plates in 500 μL of DMEM medium supplemented with 10% v / v FBS one day prior to transfection. On day 1, 50 ng of transposase expression vector was combined with 450 ng of reporter and transfected using 1 μL of JetPrime transfection reagent (Polyplus Transfection) according to the manufacturer's instructions. One day after transfection, the percentage of GFP-positive cells was measured by flow cytometry. The results are shown in Table 4. [Table 4]
[0238] As shown in Table 4, chimeric transposases with MosI inserted at the S544 fusion site exhibited higher transposon cleavage activity than those with MosI inserted at the V552 fusion site. Furthermore, reporters containing PB:MosI ITRs, rather than ITRs with -1, +1, or +2 intervals, showed higher transposon cleavage activity.
[0239] The chimeric ITRs tested in Table 4 contain a 23 bp sequence that includes binding sites for both MosI DNA-binding domains. Since the first 58 residues of MosI contain a single DNA-binding domain, a series of cleaved ITR sequences were constructed in which the MosI portion of the ITR contained binding sites for only a single domain.
[0240] In the second experiment, the MosI binding site was shortened from 23 bp to 14 bp to create left-end LE PB:MosI ITR shortenings (SEQ ID NO: 54) and right-end RE PB:MosI ITR shortenings (SEQ ID NO: 55). Using these shorter ITRs, a disrupted GFP reporter was created using a PB:MosI ITR shortening transposon (SEQ ID NO: 56), as described in Example 1. Using the above GFP reporter system, the shortened and original chimeric ITR reporters were tested with the S544 fusion site chimeric transposases, S544-MosI-58 and S544-MosI-111. The results are shown in Table 5. [Table 5]
[0241] As shown in Table 5, the original ITR and the shortened ITR each yielded functional transposon excision activity; however, the original 23bp version yielded higher activity against chimeric transposases containing either MosI fusions at residues 3–58 or 3–111, as observed by the percentage of GFP-positive cells.
[0242] Example 4: Construction and analysis of a chimeric TAL-PB:MosI transposase designed for site-specific transposition of low molecular weight transposons in the LINE-1 element. This example illustrates the construction of a chimeric TAL-PB:MosI transposase fusion protein composition useful in achieving site-directed transposition at a specific target gene locus, LINE1.
[0243] Previously, a specific TAL DNA binding domain was ligated to the N-terminal deleted piggyBac transposase sequence to construct a LINE1 TAL-ssSPB fusion protein (e.g., left; SEQ ID NO: 57; right SEQ ID NO: 58). Similar to the construction of the chimeric TAL-PB:MosI transposase fusion protein targeting GFP, the PB transposase sequence was replaced with PB:MosI transposase as described above in Example 1, converting the LINE1 left and right TAL-ssSPB fusion proteins into chimeric site-specific PB:MosI transposases. Two versions of the left LINE1 L2 TAL-PB:MosI transposase were created using either the S544 fusion point and the 3-58 residue MosI fragment (SEQ ID NO: 59) or the 3-111 residue MosI fragment (SEQ ID NO: 60). Similarly, two versions of the right LINE1 R2.2 TAL-PB:MosI transposase were created using either the S544 fusion point and the 3-58 residue MosI fragment (SEQ ID NO: 61) or the 3-111 residue MosI fragment (SEQ ID NO: 62). A small 365 bp transposon containing the PB minimal ITR (SEQ ID NO: 63) and a small 353 bp transposon containing the chimeric PB:MosI (23 bp) ITR (SEQ ID NO: 50) from Example 2 were cloned into a 4.5 kb donor vector. On day 0, 450 ng of each transposon donor was co-transfected into 120,000 HEK293T cells (seeded 1 day prior) along with a total of 50 ng of the LINE1 TAL-PB pair using 1 μL of JetPrime transfection reagent (Polyplus Transfection) according to the manufacturer's instructions. On day 3, genomic DNA was harvested from the transposed cells, and site-specific integration of the transposon into the forward target site was quantified by digital droplet PCR (ddPCR) to determine the number of integrations per haploid genome. The results are shown in Table 6.
Table 6
[0244] As shown in Table 6, the chimeric MosI 3-111 residue version of the chimeric site-specific PB:MosI transposase resulted in a higher editing rate than the chimeric MosI 3-58 residue site-specific PB:MosI version. Each of the two chimeric transposases used the chimeric PB:MosI ITR to catalyze higher site-specific transposition than a transposon containing the non-chimeric PB minimal ITR. The non-chimeric TAL-ssSPB transposase used the non-chimeric PB minimal ITR to catalyze higher site-specific transposition than a transposon containing the chimeric PB:MosI ITR. These data suggest that chimeric and non-chimeric transposases can be used as orthogonal site-specific transposition systems to simultaneously deliver transposons to two distinct loci, one targeting the chimeric PB:MosI ITR target and the other targeting the wild-type PB minimal ITR.
[0245] Example 5: Site-Specific Transposition of a Large Cargo Transposon at a Specific Locus Using Chimeric PB:MosI Transposase This example illustrates site-specific transposition of a transposon containing a large DNA cargo at a specific target locus, LINE-1.
[0246] The chimeric TAL-PB:MosI transposase and the control transposase TAL-ssSPB constructed in Example 4 were analyzed for their ability to site-specifically transpose a transposon containing a large DNA cargo to the LINE1 element.
[0247] 5' to 3' direction: A transposon donor nanoplasmid containing a PB transposon consisting of a 309 bp fragment containing the first TTAA sequence, portions of the PB 5'ITR and UTR, an EF1a promoter, the (SEQ ID NO: 45) puromycin resistance gene, a 2A peptide, and a GFP reporter, followed by a 238 bp fragment containing portions of the PB 3'ITR and UTR, and a second TTAA sequence (SEQ ID NO: 64). This donor plasmid was used to analyze wild-type PB control TAL-ssSPB and negative control PB transposases containing mutations that render transposase integration defective.
[0248] For the analysis of chimeric TAL-PB:MosI transposases, the transposon donor nanoplasmid was modified by replacing the 309 bp LE ITR / UTR and 238 bp RE ITR / UTR with the chimeric PB:MosI(23 bp)ITR of SEQ ID NO: 1 and SEQ ID NO: 19. The results are shown in Table 7. [Table 7]
[0249] As shown in Table 7, the two chimeric and non-chimeric transposases each functioned as an orthogonal system exhibiting high specificity for their respective ITRs. The chimeric TAL-PB:MosI 3-111 residue version of the chimeric PB:MosI transposase yielded higher editing rates than the chimeric TAL-PB:MosI 3-58 residue version and the non-chimeric version TAL-ssSPB. As a negative control, the integration-defective PBx transposase lacking the TAL DNA targeting domain yielded only background integration levels.
[0250] Example 6: Construction of a chimeric PB:MosI transposase containing a PB hyperactive mutation ("SPB") that increases transposition frequency. This example illustrates a method for constructing a chimeric SPB:MosI transposase and the analysis of its integration and cleavage activity. In this example, an SPB (SEQ ID NO: 5) having an N-terminal NLS was modified as described above to create an SPB:MosI chimeric transposase.
[0251] More specifically, residues 3-58 of MosI (SEQ ID NO: 8) were fused to SPB at S544 to create SPB:MosI-S544-58 (SEQ ID NO: 9), and residues 3-111 of MosI (SEQ ID NO: 6) were fused to SPB at S544 to create SPB:MosI-S544-111 (SEQ ID NO: 7). The chimeric SPB:MosI transposases were tested using disruptive GFP reporter systems containing PB:MosI ITR transposons (SEQ ID NO: 50), PB:MosI ITR truncated transposons (SEQ ID NO: 56), or PB minimal ITR transposons (SEQ ID NO: 63). On day 0, each reporter construct was co-transfected into HEK293T cells with each chimeric PB:MosI transposase. In short, 120,000 cells were seeded one day before transfection in a 24-well plate in 500 μL of DMEM supplemented with 10% v / v FBS. On day 0, 10 ng of transposase expression vector was combined with 490 ng of reporter and transfected using 1 μL of JetPrime transfection reagent (PolyPlus Transfection). On day 1, the percentage of GFP-positive cells was measured by flow cytometry. The results are shown in Table 8. [Table 8]
[0252] As shown in Table 8, chimeric SPB:MosI transposase catalyzed transposon excision for transposons containing chimeric PB:MosI ITR, as measured by the percentage of GFP-positive cells, but did not catalyze transposons containing non-chimeric minimal PB ITR. Non-chimeric SPB transposase catalyzed transposon excision for non-chimeric minimal PB ITR, but did not catalyze transposons containing chimeric PB:MosI ITR, demonstrating the specificity of the chimeric PB:MosI transposase / PB:MosI ITR transposon combination.
Claims
1. A chimeric transposase comprising, in order from the N-terminus to the C-terminus: (i) a target-specific DNA-binding domain, (ii) a cleaved Super piggyBac (SPB) transposase containing a C-terminal cysteine-rich domain (CRD) deletion within amino acid residues 535-594 of the SPB sequence including the sequence shown in Sequence ID No. 1, and (iii) one or more MosI DNA-binding domains.
2. The chimeric transposase according to claim 1, wherein the cleavage-type SPB transposase of (ii) comprises the sequence shown in any one of sequence numbers 68 to 84.
3. The chimeric transposase according to claim 1 or 2, wherein the cleavage-type PB transposase comprises one or more hyperactive mutations selected from I30V, S103P, G165S, M226F, M282V, S509G, N538K, or N571S of SEQ ID NO:
1.
4. The chimeric transposase according to any one of claims 1 to 3, wherein the PB transposase containing the CRD deletion further comprises an in-frame N-terminal nuclear localization sequence (NLS) containing the amino acid sequence of SEQ ID NO:
39.
5. The chimeric transposase according to any one of claims 1 to 4, wherein the chimeric transposase comprises two MosI DNA-binding domains having the amino acid sequence shown in SEQ ID NO:
6.
6. The chimeric transposase according to any one of claims 1 to 4, wherein the chimeric transposase comprises one MosI DNA-binding domain having the amino acid sequence shown in SEQ ID NO:
8.
7. The chimeric transposon according to any one of claims 1 to 6, wherein the target-specific DNA-binding domain is a zing finger domain or a TAL array.
8. A polynucleotide comprising a nucleic acid sequence encoding a chimeric transposase according to any one of claims 1 to 7.
9. A vector comprising the polynucleotide described in claim 13.
10. In the direction from 5' to 3', a chimeric transposon inverted terminal repeat (ITR) polynucleotide comprising (i) nucleotides 1-16 of the nucleic acid sequence shown in SEQ ID NO: 12 and (ii) a polynucleotide comprising two MosI DNA binding sites.
11. The chimeric transposon ITR polynucleotide according to claim 10, wherein the polynucleotide containing two MosI DNA binding sites contains the nucleic acid sequence of SEQ ID NO: 13 or 14.
12. A chimeric transposon ITR polynucleotide according to claim 10 or 11, comprising the nucleic acid sequence of sequence number 15 or 16.
13. A chimeric transposon ITR polynucleotide according to any one of claims 10 to 12, comprising the nucleic acid sequence of sequence number 17 or 18.
14. A vector comprising a chimeric transposon ITR polynucleotide according to any one of claims 10 to 13.
15. A chimeric transposon ITR polynucleotide comprising, in the direction from 5' to 3', (i) a polynucleotide containing nucleotides 1 to 16 of the nucleic acid sequence shown in SEQ ID NO: 65, and (ii) a polynucleotide containing two MosI DNA binding sites.
16. The chimeric transposon ITR polynucleotide according to claim 15, wherein the polynucleotide containing two MosI DNA binding sites contains the nucleic acid sequence of SEQ ID NO: 13 or 14.
17. The chimeric transposon ITR polynucleotide according to claim 15, comprising one nucleic acid sequence from sequence numbers 19 to 22.
18. A vector comprising a chimeric transposon ITR polynucleotide according to any one of claims 15 to 17.
19. A LINE1-targeted chimeric site-specific transposase fusion protein comprising (i) a left or right LINE1 TAL-targeted DNA-binding domain, (ii) a linker sequence, and (iii) a chimeric PB:MosI transposase, in the direction from the N-terminus to the C-terminus.
20. The LINE1-targeted chimeric site-specific transposase fusion protein according to claim 19, wherein the chimeric PB:MosI transposase comprises the amino acid sequence shown in SEQ ID NO: 9 or 7.
21. The LINE1-targeted chimeric site-specific transposase fusion protein according to claim 20, wherein the left or right LINE1 TAL-targeted DNA-binding domain comprises the amino acid sequence shown in SEQ ID NO: 66 or 67.
22. A transposon comprising (i) a chimeric LE PB:MosI ITR polynucleotide and (ii) a chimeric RE PB:MosI ITR polynucleotide, wherein the LE PB:MosI ITR polynucleotide and the RE PB:MosI ITR polynucleotide contain the same nucleic acid sequence.
23. The transposon according to claim 22, wherein each of the chimeric LE PB:MosI ITR polynucleotides and the chimeric RE PB:MosI ITR polynucleotides each contain one nucleic acid sequence from sequence numbers 15 to 22.
24. A transposon comprising (i) a chimeric LE PB:MosI ITR polynucleotide containing the nucleic acid sequence of SEQ ID NO: 54 and (ii) a chimeric RE PB:MosI ITR polynucleotide containing the nucleic acid sequence of SEQ ID NO:
55.
25. A transposon containing the nucleic acid sequence of sequence number 56.
26. A method for incorporating a transgene into a genomic target site of a cell, comprising introducing the chimeric transposase and transposon described in any one of claims 1 to 7 into the cell, wherein the transposon comprises a 5'ITR, the transgene, and a 3'ITR in the order of 5' to 3', the 5'ITR being a chimeric transposon inverted terminal repeat (ITR) described in any one of claims 10 to 13, and the 3'UTR being a chimeric transposon inverted terminal repeat (ITR) described in any one of claims 15 to 17.
27. The method according to claim 26, wherein the 5'-ITR chimeric transposon inverted terminal repeat sequence comprises the nucleic acid sequence shown in SEQ ID NO: 15, and / or the 3'-ITR chimeric transposon inverted terminal repeat sequence comprises the nucleic acid sequence shown in SEQ ID NO:
21.
28. The method according to claim 26 or 27, wherein the genome target site is located within a repeating element.
29. The method according to claim 28, wherein the repeating element is a LINE element.
30. A method for site-specific transposition of DNA molecules into the genome of a cell, wherein the cell a) Nucleic acids encoding a chimeric site-specific transposase fusion protein comprising a DNA-binding domain and, in the order of N-terminus to C-terminus: a piggyBac transposase containing a C-terminal cysteine-rich domain (CRD) deletion and a chimeric transposase containing one or more MosI DNA-binding domains; b) The fusion protein is expressed in the cell, and c) A method comprising introducing a DNA molecule containing a transposon comprising a chimeric LE transposon ITR polynucleotide and a chimeric RE transposon ITR polynucleotide, wherein the expressed chimeric site-specific transposase fusion protein incorporates the transposon into the TTAA sequence of the cell genome by site-specific transposition.
31. A method for generating cells manipulated by site-specific transposition, wherein the cells are: a) Nucleic acids encoding a chimeric site-specific transposase fusion protein comprising a DNA-binding domain and, in the order of N-terminus to C-terminus: a piggyBac transposase containing a C-terminal cysteine-rich domain (CRD) deletion and a chimeric transposase containing one or more MosI DNA-binding domains; b) (The chimeric site-specific transposase fusion protein is expressed in the cells) and c) A method comprising introducing a DNA molecule containing a transposon comprising a chimeric LE transposon ITR polynucleotide and a chimeric RE transposon ITR polynucleotide, wherein the expressed chimeric site-specific transposase fusion protein incorporates the transposon into the TTAA sequence of the cell's genome by site-specific transposition, thereby generating the manipulated cell.